Make green actually mean green.
Free & open · no signup · read-only — every stage ends by asking · nothing leaves your machine
Earn a green suite you can trust: find the coverage gaps, strengthen the tests that assert nothing, kill the flakes, and rebalance a top-heavy pyramid.
The conductor fetches each brief in turn and writes its report before moving on — later briefs can build on earlier findings. Or copy any single stage to run it alone.
4 reports in reports/, plus the run's own INDEX.md: TESTING.md, TESTQUALITY.md, FLAKY.md, PYRAMID.md. Feed them to the optional Studio to turn findings into commits, or run 28 · Roadmap Synthesis to merge them into one plan.
Copy the conductor into your agent inside the repo you want checked. It runs each Goal Prompt in sequence, honoring each one's ask-first rule.
Map what is tested against what is riskiest, and produce a test-writing plan that buys the most safety per hour.
Whether the tests actually protect the code — asserting meaningful behavior, or merely executing lines and passing no matter what.
The tests that fail intermittently — the ones eroding trust in the suite until a red build stops meaning anything.
Count and time the suite by layer, draw the shape it actually makes, and compute confidence-per-second — the number that says which tests earn their runtime.