Skip to main content
Version: Next

QA and red before green

Every ticket an agent implements is tested before its branch leaves the machine. QA runs between the implementer's last commit and the push, so the tests it writes, and any fix they force, travel with the change they are about. Where this sits in a ticket's run: inside one worker.

The loop​

  1. QA derives test cases from the ticket and the change. When there are enough groups of cases to share out, several agents write them at once; the QA session that follows still sees every case. On the watch board this shows as qa 2/5.
  2. QA writes the tests, runs the repository's own suite (scripts/test.sh), and commits the tests on the ticket's branch.
  3. Findings go back to the implementer, which fixes them on the same branch. QA then checks again.
  4. The loop ends when QA finds nothing, or when it reaches qa.max_rounds (default 3). At the ceiling the ticket is still pushed, and what is still open is reported for a person.

QA never blocks a run​

QA cannot stop a run, since a verifier that could block would have to be right every time. At most it spends its rounds and reports what is still open. A QA problem, such as an agent that cannot be reached or a suite that times out, is a warning and never fails the ticket, because the change is already committed.

Why the ceiling is 3, and what raising it costs, is on Tuning the guardrails.

Red before green​

A test written after the code already works can pass without ever exercising what it claims to check. So after QA, Orion takes each test file QA added or changed, lays it onto the commit the ticket's branch started from, in a throwaway worktree, and runs the suite there. A test that still passes without the change proves nothing about it, and is reported as such.

It works on whole test files, because the repository's test command is the one contract Orion has for running tests. Like the rest of QA, it reports and does not block the run. The decision is ADR 0037.

Fixing a bug: fix mode​

For a bug fix, the failing test has to stay failing until the code is fixed.

orion fix start # test files become read-only to the agent
orion fix end # back to normal

In fix mode, the shield hook refuses any edit to a test file, so the test that shows the bug cannot be weakened to make it pass. Its message tells the agent to fix the code so the test passes as written, and, if the test itself is wrong, to stop and say so: a person decides that. What counts as a test file is paths.test_globs in orion.json. The switch is gates.protect_tests_during_fix.

Settings​

SettingDefaultEffect
qa.max_rounds3QA fix rounds before the ticket is pushed with findings open
gates.protect_tests_during_fixonFix mode makes test files read-only to the agent

Not documented yet​

  • How QA runs tests in a repository that has no scripts/test.sh.