Skip to main content
Version: Next

0037: QA's tests are run against the pre-change commit, and fix mode keeps a failing test from being weakened

  • Status: Accepted (backfilled 2026-10-06: records a decision already built)
  • Date: 2026-10-06
  • Related: 0002 (adopted natively rather than through the superpowers plugin)

Context​

QA writes its tests after the implementer has committed. A test written against code that already behaves correctly can pass without ever exercising the behaviour it claims to check, and then a later regression passes straight through it. The only way to know a test would have caught something is to watch it fail first.

A related failure happens during a bug fix: the quickest way to make a failing test pass is to weaken the test.

Decision​

Red before green, checked mechanically (internal/work/redgreen.go). Orion knows the commit each ticket's branch started from, because it cut the branch there. After QA, each test file QA added or changed is laid onto that pre-change commit in a throwaway worktree, and the repository's own suite is run. A test that still passes without the change proves nothing about it, and is reported as such.

It works on whole test files, because the repository's test command (scripts/test.sh) is the one contract Orion has for running tests, and there is no language-general way to ask it for one test's verdict.

It reports and does not block, for the same reason QA does not (internal/work/qa.go): a verifier that could stop a run would have to be right every time.

Fix mode protects the failing test. orion fix start makes test files read-only to the agent until orion fix end, enforced by the shield hook (gates.protect_tests_during_fix). The test that shows the bug cannot be edited into passing.

Consequences​

  • A QA test that never exercised the change is named, not trusted.
  • The check costs one suite run per pass on the pre-change commit.
  • Granularity is the test file, so one vacuous test in a good file is not singled out.

Alternatives rejected​

  • Ask the agent whether its test would have failed. Unverifiable.
  • Block the run on a vacuous test. Gives QA an authority it cannot carry.
  • Per-test selection. Needs a test selector per language and framework; Orion has one contract for running tests, not one per stack.