Find the Broken Step Before Your Users Do

For teams that own a docs quickstart. FirstRun runs every command, not only the links, and proves each fix in a fresh container.

Paste a docs URL. A partial URL or a tool name also works, and your agent finds the page. Nothing leaves your browser.

Try

How Doc-test works. Broken steps go in. Real ones FirstRun found: jafgen --years 6 failed with No such option in the dbt guide. entire: command not found after the Entire install. A migration guide link that returns 404. A dbt guide that never states it needs Python 3.12. No module named typer. curl: command not found. AccuKnox docs that say 0.9.0 while the installer fetches 0.9.58. KeyError NEATLOGS_API_KEY.

How Doc-audit works, on dbt-labs/dbt-core. The project flags page was last edited 453 days ago, with 29 code changes since. The infer-configs page is 709 days old. 15 of 1,405 pages are stale. 72 shipped features have no docs. AccuKnox/help has 87 stale pages, and its oldest page is 1,100 days old.

Portrait of Atharva ShahAtharva ShahBuilt for neatHack, the neatlogs hackathon

Doc-test

FirstRun Runs Your Quickstart Like a Brand-New User

You give it

  • A public docs URL, like docs.getdbt.com/guides/duckdb
  • A tool name. Your coding agent finds its quickstart page.
  • A GitHub repo, for Doc-audit

It plans, runs and proves

  • One model call turns the page into at most 10 steps, one command each.
  • Each step runs in a clean Docker container.
  • A broken step gets a fix. A fresh container reruns every step with no model calls.

Not yet

  • Pages behind a login
  • API specs and bug reports as input
  • Long-running servers. It skips them and says why, like dbt docs serve below.

Every step runs on your machine. Only the docs page and the output of a failed step go to the model. Read the known limits

  1. 1A step breaks

    jafgen --years 6Error: No such option: --yearsThe dbt guide never worked on a fresh machine.
  2. 2FirstRun fixes the docs

    jafgen --years 6jafgen 6The tool now takes the years as a plain number.
  3. 3A second clean run proves it

    9 of 9 steps passed

The Plan FirstRun Made From the dbt Page

10 steps, one command each

  1. 1git clone https://github.com/dbt-labs/jaffle_shop_duckdb.gitThe container had no git, so FirstRun installs it first.Fixed
  2. 2cd jaffle_shop_duckdbPassed
  3. 3pip install -r requirements.txtA docs bug. The guide never says it needs Python 3.12, which networkx 3.7 needs. The fix pins 3.6.1.Fixed
  4. 4dbt seedPassed
  5. 5dbt buildPassed
  6. 6dbt docs generatePassed
  7. 7dbt docs serveIt starts a web server, which a headless container cannot check.Skipped
  8. 8ls -lah *.duckdbPassed
  9. 9python -m pip install jafgenPassed
  10. 10jafgen --years 6A docs bug. The tool now takes jafgen 6.Fixed

Fresh container: 9 of 9 runnable steps passed, with no model calls.

A real run on the dbt + DuckDB quickstart. Replay this run

FirstRun Already Sent Fixes and Bug Reports to dbt, Entire, Polars, Chroma and HTTPX

Test your docs

Doc-audit

Each Fix Goes to the Person Who Knows the Code

AltimateAI/altimate-code14 people, 619 commits in 12 months

People

Code · share of each part

    Top 3 Tasks on This Board

      Built from public commits in the last 12 months. Suggestions, not assignments.

      Doc-audit

      Find the Pages Your Code Left Behind

      1. 1Paste a repo

        dbt-labs/dbt-core
        Any public GitHub repo with docs in it.
      2. 2FirstRun reads the history

        Page edited Jul 202529 code changes since
        fix(flags): normalize project flag keys to uppercaseJul 2, 2026
      3. 3You get the list

        15stale pages
        72features with no docs
      The Doc-audit board for dbt-labs/dbt-core: 15 stale pages and 72 features with no docs, in columns to ship next

      Open the dbt-core audit

      Audit a repo

      Receipt

      Every Run Ends With a Receipt

      FirstRunOct 10, 2026

      Docs page help.accuknox.com/knoxctl
      Broke at step 2. Missing curl, gnupg and sudo.

      1Model cost$0.0453
      2Team plan cap$0.0469796.4% of the cap used
      3Clean rerun5 of 5
      Verified
      1. PricedYou see what the run cost.
      2. CappedYour plan sets the most a run may spend.
      3. ProvenA clean rerun passed every step.

      Watch FirstRun Test a Guide and Audit a Repo

      Doc-test · 2:59

      Your Docs Lie. FirstRun DocTest Proves It

      A clean container runs the dbt guide, hits the step that breaks, fixes the page and passes a second run.

      Doc-audit · 1:56

      Find the Docs Your Code Left Behind

      A real audit of dbt-core: one repo in, a run on a laptop, the board, an owner for each fix, and the receipt.

      Teaser · 0:36

      The dbt Guide That Never Worked

      Doc-audit · 0:57

      Docs Rot Quietly. FirstRun DocAudit Finds It

      Powered by neatlogs, Entire CLI and cfo.ai

      neatlogs

      neatlogs traces every run, so a failed step comes with its cause. On the same 4 quickstarts, agent v2 finished 3 where v1 finished 2, and the mean cost fell from $0.0262 to $0.0200. See v1 against v2

      1. neatlogs incident: 25 Execution Failed detections against a baseline of 0.5, a 52.8x spike, investigated with a root cause at 66% confidence

        Investigate traced a 52.8x spike in failed steps to one root cause.

        Read the neatlogs evidence
      2. neatlogs trace of a FirstRun run: step 3 fails, the recover.3 agent span classifies the failure and retries, and step 3 runs again

        Every recovery is a span. In this Dagster run, step 3 failed twice and recover.3 retried it.

        Read another recovered run in full
      3. neatlogs trace of a Docs Bisect run: 5 probes across jafgen releases, last good 0.3.1, first bad 0.4.6

        A Bisect trace tested 5 jafgen releases and found the first bad one, 0.4.6.

        Read the finding on dbt PR #10149

      Entire

      Entire records the session behind each commit and reviews code before it merges.

      1. GitHub commit 3a0dcad whose message ends with an Entire-Checkpoint ID

        91 of 132 commits link to the Entire session that wrote them.

        Open a commit's session on Entire
      2. Entire Trail 3 for Layer 1: merge blocked, no approvals and 1 unresolved high finding

        Trail 3 blocked the merge while 1 high finding stayed open.

        Read the Entire evidence
      3. Entire Trail Review findings: a high finding in recovery.py, a retry loop on a missing secret, and a medium finding in config.py, a Docker default that only works on the author's machine

        Trail reviews caught 3 real bugs before merge, and each one was fixed.

        See the fix commit

      cfo.ai

      Ari, the cfo.ai agent, built the 18-month model that sets each plan's cost cap.

      1. Cash on hand in the cfo.ai model $10k$0-$10k Nov 2026Apr 2028 Month 14: cash runs out $16,685Base -$12,212Old cap

        Under the old $0.50 cap, the model runs out of cash in month 14.

        Open the public cfo.ai model
      2. cfo.ai drivers table: the mean, p50, p90, p95 and max cost per run are marked Measured yes, with drivers.csv as the source

        Its cost per run comes from 88 neatlogs traces, and each row names its source.

        See drivers.csv
      3. cfo.ai model notes: caps written to the FirstRun CLI, Scale $0.0279, Team $0.0470 and Free $0.0135 per run

        The plan caps a Scale run at $0.0279 and a Team run at $0.0470.

        See how each cap is set

      See all the proof for judges

      Compare FirstRun With Claude Code, GitHub Actions and GitHub Agents

      FirstRun does one job: it runs the quickstart, fixes the docs, proves the fix and opens the PR. A link checker or a docs linter only reads the page.

      Yes Partly or with setup No
      FeatureFirstRunClaude CodeGitHub ActionsGitHub AgentsCustom workflow
      Runs the quickstart as a new developerEvery step runs in a fresh Docker containerRuns on your machine. An optional sandbox limits files and networkA fresh runner per job, but you script every stepWorks in its own VM on your repo, not a new-user setupOnly as clean as the environment you build
      Plans steps from a docs URLWrites one step per command on the pageWhen you prompt it with the pageYou write each stepStarts from an issue or task, not a docs pageYou write the parser
      Tries again when a step failsFinds the cause, patches, retries up to 3 times per stepIterates inside one sessionThe job fails and stopsIterates and pushes commits to its draft PROnly what you code
      Proves the fixA second clean container must pass the whole quickstartOnly if you ask for itOnly if you script a second runYour repo's own tests, not a clean-user rerunOnly if you code it
      Shows the cost of each runneatlogs traces every run. The mean run costs $0.018Cost and token metrics, after you set them upLogs and run history, no agent tracesA session log, with no cost trace you ownYou build it
      Opens the docs PRAfter you approve onceWhen you askOnly if you script itOpens a draft PR and asks you for reviewOnly if you code it
      Easy to startPaste a docs URL, then paste the prompt into your coding agentOne prompt, but you define the test and the clean setupA YAML file per quickstartAssign an issue to CopilotHighest effort
      Proof in public repos4 upstream docs PRs, all openNot specific to docsNot specific to docsNot specific to docsDepends on your build

      The Build Happened in Public

      1. Sat Oct 10, 10:45
      2. Sat Oct 10, 11:00
      3. Sat Oct 10, 12:32
      4. Sat Oct 10, 13:14
      5. Sat Oct 10, 13:26
      6. Sat Oct 10, 15:15
      7. Sat Oct 10, 18:15
      8. Sat Oct 10, 18:23
      9. Sat Oct 10, 21:15
      10. Sun Oct 11, 07:32
      11. Sun Oct 11, 10:32
      Portrait of Atharva Shah

      Atharva Shah

      Atharva is a DevRel engineer. He owns developer docs at AccuKnox and works with Altimate AI. He built FirstRun alone, in public.