Skip to main content
Version: Next

Recovery runbook

Most failures clear themselves: the watch retries a failed ticket, releases a held one when the environment recovers, and reconciles claims nobody is holding. When it cannot, find what you see in the first column below.

You seeWhat it meansDo
A ticket under NEEDS YOUOut of automatic retriesLook, steer, requeue
orion-failed, retries leftWaiting for develop to moveWait, or requeue now
A failed ticket whose comment holds a questionThe agent asked what the advisor could not answerAnswer it
Tickets listed as heldThe environment stopped them before any workRelease the hold
A ticket held "evicted N times; a person should look"Tripped the breaker, ran out of fix rounds, or was stranded, repeatedlyLook before requeueing
A tripped breaker on a runA loop, repeated failure or budget stopReview, then reset
A batch that will not landRed together but every part green, or develop itself is redChange the set, or fix develop
Collect cannot rebase a ticket's branchSomething in its worktree is blocking itSettle it
orion-working on a ticket nothing is runningA leftover claimLet the watch clear it
The watch stopped: "no progress"An hour of sweeps changed nothingRead what it waited on
Planning stopped on open questionsIntent, spec or plan needs your answersorion answer

Everything here works without tracker access beyond the ticket itself, except where a step says to comment on the ticket.

Finding out what happened​

CommandShows
orion queueEvery ticket the watch can see, grouped by state
orion logs KEYWhat Orion and the agents did on a ticket; -f follows it live
orion logs KEY --transcriptThe raw agent output
orion status <id>A workspace's stage, breaker state and last run
orion aiops KEYA finished run's log, read for what deserves your attention
orion report --since 7dFailures, workspaces, budget and usage over a period

The ticket's own comments say what Orion found at each failure, and the watch's log file is kept under ~/.orion/logs/.

Out of retries​

The watch requeues a failed ticket at most twice, each time after develop has moved past where it failed. After the second retry the ticket is named under NEEDS YOU and the watch stops trying.

  1. Read why: the failure comments on the ticket and orion logs KEY.
  2. Fix the cause if it is outside the ticket: a missing prerequisite, a red develop, a credential.
  3. Steer the next run by commenting on the ticket. The most recent comments a person wrote are given to its agents.
  4. Requeue: orion queue add KEY --reset. This removes orion-failed, adds the queue label and returns the ticket to To Do. Relabelling by hand leaves the status wrong.

Failed with retries left​

Nothing to do. Most of these failures are a sibling ticket that had not landed yet, and the watch requeues the ticket once develop moves. Requeue it now with orion queue add KEY --reset only if you have fixed the cause yourself.

Blocked on a question​

The agent produced nothing and stopped to ask, and the advisor could not answer from the confirmed project artifacts. The ticket is orion-failed and the question is on it.

Answer by commenting on the ticket, or by amending the artifact the advisor reads so the next ticket does not ask again. Then run orion queue add KEY --reset. Advisors and decisions explains when the advisor answers.

Held tickets​

A hold means the environment stopped the run before any work: the tracker or a credential was unreachable, or a quota wall was hit. Nothing was spent, no retry was used, and the ticket went back to the queue.

The watch re-checks the cause every sweep and releases the hold by itself when the check passes; a quota hold lifts at its reset time. To force a re-check now:

orion reset --held # every cause
orion reset --held <fault> # one cause

If the check still fails, orion doctor says why.

Evicted repeatedly​

The planner evicts a ticket that has tripped the breaker too often, run out of fix rounds, or been stranded too often. Twice is recorded; on the third time it is held for a person.

Read the ticket's history before requeueing. The usual cause is a ticket too large for one run, or a dependency it does not declare. Split it or add the dependency, then requeue. The ceilings are limits.max_breaker_trips and limits.max_stranded.

A tripped breaker​

The breaker stopped a run for repeating itself, failing the same way, or exceeding its budget. What each trip means, and what an agent may still do after one, is on Circuit breakers. After you have read why it fired:

orion reset --session <id>
orion queue add KEY --reset

If a run's work was preserved as a wip: commit, it is on the ticket's branch for you to read, resume or drop.

A stuck batch​

Two cases stop the integration loop from spending more CI on the same set:

  • Red together, green apart: isolation proved every part of the batch green and the whole red, so the fault needs two tickets together. Orion records the batch as stuck and will not test the same set again. Change the set: push a fix to one of the branches, or take one ticket off the shelf by removing its orion-ready label. The record stops applying as soon as the set or develop changes.
  • develop itself is red, so every member looked guilty. Fix develop; the next sweep assembles again.

A blocked worktree​

When collect cannot rebase a ticket's branch because something in its worktree is in the way:

orion settle KEY --dry-run # report what is there
orion settle KEY # commit it so the rebase can go ahead

A leftover claim​

orion-working is the lock a running agent holds. The watch checks every claim each sweep: one whose process has gone is released, one on a ticket closed by hand loses its lock, and finished work whose label fell behind is moved to orion-ready.

If the watch is not running and you are certain no orion process is working the ticket, remove orion-working and requeue it. Never remove it from a ticket an agent is still working: that lets a second run start on the same ticket.

The watch stopped itself​

After an hour of sweeps that changed nothing, the no-progress breaker stops the watch and names what it was waiting on: usually a stuck batch, CI that never finished, or tickets that are all held. Deal with that, then start the watch again. The hour is limits.no_progress_minutes.

The budget cannot catch this case: a loop with no model in it spends no tokens, but a batch that re-assembles, re-tests and lands nothing still spends CI minutes.

Planning stopped on open questions​

Intent, spec and plan are not done while they list open questions. orion answer <id> walks you through them and writes your answers into the artifact; then run orion plan KEY again. See Inception.