Recovery runbook
Most failures clear themselves: the watch retries a failed ticket, releases a held one when the environment recovers, and reconciles claims nobody is holding. When it cannot, find what you see in the first column below.
| You see | What it means | Do |
|---|---|---|
| A ticket under NEEDS YOU | Out of automatic retries | Look, steer, requeue |
orion-failed, retries left | Waiting for develop to move | Wait, or requeue now |
| A failed ticket whose comment holds a question | The agent asked what the advisor could not answer | Answer it |
| Tickets listed as held | The environment stopped them before any work | Release the hold |
| A ticket held "evicted N times; a person should look" | Tripped the breaker, ran out of fix rounds, or was stranded, repeatedly | Look before requeueing |
| A tripped breaker on a run | A loop, repeated failure or budget stop | Review, then reset |
| A batch that will not land | Red together but every part green, or develop itself is red | Change the set, or fix develop |
| Collect cannot rebase a ticket's branch | Something in its worktree is blocking it | Settle it |
orion-working on a ticket nothing is running | A leftover claim | Let the watch clear it |
| The watch stopped: "no progress" | An hour of sweeps changed nothing | Read what it waited on |
| Planning stopped on open questions | Intent, spec or plan needs your answers | orion answer |
Everything here works without tracker access beyond the ticket itself, except where a step says to comment on the ticket.
Finding out what happenedβ
| Command | Shows |
|---|---|
orion queue | Every ticket the watch can see, grouped by state |
orion logs KEY | What Orion and the agents did on a ticket; -f follows it live |
orion logs KEY --transcript | The raw agent output |
orion status <id> | A workspace's stage, breaker state and last run |
orion aiops KEY | A finished run's log, read for what deserves your attention |
orion report --since 7d | Failures, workspaces, budget and usage over a period |
The ticket's own comments say what Orion found at each failure, and the
watch's log file is kept under ~/.orion/logs/.
Out of retriesβ
The watch requeues a failed ticket at most twice, each time after develop
has moved past where it failed. After the second retry the ticket is named
under NEEDS YOU and the watch stops trying.
- Read why: the failure comments on the ticket and
orion logs KEY. - Fix the cause if it is outside the ticket: a missing prerequisite, a red
develop, a credential. - Steer the next run by commenting on the ticket. The most recent comments a person wrote are given to its agents.
- Requeue:
orion queue add KEY --reset. This removesorion-failed, adds the queue label and returns the ticket to To Do. Relabelling by hand leaves the status wrong.
Failed with retries leftβ
Nothing to do. Most of these failures are a sibling ticket that had not
landed yet, and the watch requeues the ticket once develop moves. Requeue
it now with orion queue add KEY --reset only if you have fixed the cause
yourself.
Blocked on a questionβ
The agent produced nothing and stopped to ask, and the advisor could not
answer from the confirmed project artifacts. The ticket is orion-failed
and the question is on it.
Answer by commenting on the ticket, or by amending the artifact the advisor
reads so the next ticket does not ask again. Then run
orion queue add KEY --reset. Advisors and decisions
explains when the advisor answers.
Held ticketsβ
A hold means the environment stopped the run before any work: the tracker or a credential was unreachable, or a quota wall was hit. Nothing was spent, no retry was used, and the ticket went back to the queue.
The watch re-checks the cause every sweep and releases the hold by itself when the check passes; a quota hold lifts at its reset time. To force a re-check now:
orion reset --held # every cause
orion reset --held <fault> # one cause
If the check still fails, orion doctor says why.
Evicted repeatedlyβ
The planner evicts a ticket that has tripped the breaker too often, run out of fix rounds, or been stranded too often. Twice is recorded; on the third time it is held for a person.
Read the ticket's history before requeueing. The usual cause is a ticket too
large for one run, or a dependency it does not declare. Split it or add the
dependency, then requeue. The ceilings are limits.max_breaker_trips and
limits.max_stranded.
A tripped breakerβ
The breaker stopped a run for repeating itself, failing the same way, or exceeding its budget. What each trip means, and what an agent may still do after one, is on Circuit breakers. After you have read why it fired:
orion reset --session <id>
orion queue add KEY --reset
If a run's work was preserved as a wip: commit, it is on the ticket's branch for you
to read, resume or drop.
A stuck batchβ
Two cases stop the integration loop from spending more CI on the same set:
- Red together, green apart: isolation proved every part of the batch green
and the whole red, so the fault needs two tickets together. Orion
records the batch as stuck and will not test the same set again. Change the
set: push a fix to one of the branches, or take one ticket off the shelf by
removing its
orion-readylabel. The record stops applying as soon as the set ordevelopchanges. developitself is red, so every member looked guilty. Fixdevelop; the next sweep assembles again.
A blocked worktreeβ
When collect cannot rebase a ticket's branch because something in its worktree is in the way:
orion settle KEY --dry-run # report what is there
orion settle KEY # commit it so the rebase can go ahead
A leftover claimβ
orion-working is the lock a running agent holds. The watch checks every
claim each sweep: one whose process has gone is released, one on a ticket
closed by hand loses its lock, and finished work whose label fell behind is
moved to orion-ready.
If the watch is not running and you are certain no orion process is working
the ticket, remove orion-working and requeue it. Never remove it from a
ticket an agent is still working: that lets a second run start on the same
ticket.
The watch stopped itselfβ
After an hour of sweeps that changed nothing, the no-progress breaker stops
the watch and names what it was waiting on: usually a stuck batch, CI that
never finished, or tickets that are all held. Deal with that, then start the
watch again. The hour is limits.no_progress_minutes.
The budget cannot catch this case: a loop with no model in it spends no tokens, but a batch that re-assembles, re-tests and lands nothing still spends CI minutes.
Planning stopped on open questionsβ
Intent, spec and plan are not done while they list open questions.
orion answer <id> walks you through them and writes your answers into the
artifact; then run orion plan KEY again. See
Inception.