Packages
What each package under internal/ is for, in its own words. How they group into layers is in the README's Layout.
| Package | What it is for |
|---|---|
actors | Package actors is the one place that knows what an actor is CALLED. |
adopt | Package adopt wires Orion into an existing repository. |
advise | Package advise answers a question the implementer could not resolve. |
agentcfg | Package agentcfg decides what a supervised run is allowed to be. |
aiops | Package aiops is the post-run triage pass: it reads one ticket's event log AFTER the run has finished and reports what is worth a person's attention, as a report plus draft tickets that nobody has filed. |
answers | Package answers is the inbox a person's answer travels through on its way from the web page to a ticket. |
budget | Package budget accounts for what Orion spends and stops it at thresholds. |
changelog | Package changelog collates per-ticket changelog fragments into CHANGELOG.md. |
ciscaffold | Package ciscaffold gives a repository one standard way to run its tests. |
claim | Package claim records which process is working a ticket, so a claim left behind by a dead one can be told apart from a claim that is doing work. |
collect | Package collect finishes what work started. |
config | Package config loads orion.json from the project root and supplies defaults for anything absent. |
conflict | Package conflict verifies that a merge resolution dropped nothing. |
conform | Package conform answers the third question about a finished change: is this the thing we agreed to build? |
cost | Package cost answers "what did this ticket cost" from the event log. |
creds | Package creds resolves Orion's credentials from the environment or from Orion's own config file. |
dashboard | Package dashboard answers one question: is the coding queue outrunning the integration queue? |
dba | Package dba decides, deterministically and for free, whether a change touches the database. |
dbaplan | Package dbaplan is the database architect's part in PLANNING. |
decide | Package decide keeps a RECOMMENDATION apart from a DECISION. |
decompose | Package decompose turns a decomposed task list into tracker items. |
discovery | Package discovery decides whether an idea is ready to be designed from, and gates the chain until it is. |
doctor | Package doctor is the precheck. |
done | Package done answers one question about a finished run, at the moment before a person is asked to approve it: is this genuinely done, or does it only look done? |
events | Package events is Orion's live orchestration log. |
export | Package export ships Orion's events to the observability backend an operator already runs, so a watch's history can be searched and charted instead of living in a file per project. |
fanout | Package fanout decides whether an implementation run may be split across concurrent subagents, and refuses when it may not. |
hook | Package hook implements the Claude Code hook protocol. |
lessons | Package lessons is Orion's cross-project memory. |
localauth | Package localauth authenticates Orion's local web surface. |
match | Package match implements the glob subset Orion needs, including the "**" segment that path.Match does not support. |
newq | Package newq holds the questions orion new asks about an idea, in one place so the terminal interview and the web form cannot ask different things. |
notify | Package notify tells the user something happened while they were not watching. |
procsafe | Package procsafe holds the two operations that have to stay correct when SEVERAL Orion processes share one ORION_HOME. |
promote | Package promote decides whether a milestone is safe to release. |
provision | Package provision creates the remote repository and its branch model. |
queue | Package queue decides which queued tickets the watch may claim on a pass. |
quota | Package quota detects model quota and rate-limit exhaustion, works out when the limit resets, and decides whether to wait. |
reconcile | Package reconcile compares what the repository says against what the tracker says, and reports where they disagree. |
registry | Package registry maps a tracker project key to the repository it belongs to, so Orion can act on a ticket without being told where the code lives. |
report | Package report summarises what Orion has been doing. |
session | Package session tracks which orion processes are actually alive. |
sessions | Package sessions enumerates every workspace Orion could be running in. |
settings | Package settings reads and writes the per-project switches and limits in orion.json: the circuit breakers orion config limits sets and the landing switches orion config collect sets. |
slack | Package slack talks to the Slack Web API with a bot token. |
state | Package state persists per-session counters across hook invocations. |
suite | Package suite runs a repository's own test suite as a process Orion owns, rather than as an instruction in an agent's prompt. |
supervisor | Package supervisor runs claude -p inside a workspace and enforces the limits a hook structurally cannot. |
toolkit | Package toolkit locates, validates and provisions the toolkit a project delegates to. |
tracker | Package tracker provisions and binds project-management projects. |
ui | Package ui renders Orion's terminal output. |
update | Package update tells a person when the Orion they are running is old. |
watch | Package watch is the loop that removes the last manual step. |
web | The snapshot endpoint: what GET /api/snapshot returns. |
work | Package work runs one tracker issue end to end. |
workspace | Package workspace provisions isolated project directories. |
Test helpers, not part of the running system: fakebin, testproc.
actorsβ
Package actors is the one place that knows what an actor is CALLED.
A ticket is worked by several agents on several models -- opus implements,
haiku routes the question it stops on, sonnet answers it and writes the
pull request description -- and the output used to present all of them as
one anonymous voice. "created a pull request description" says nothing
about who created it or what it cost.
So every rendered line names the actor. Name AND job title, together, on
every line: the name alone ("Navjyot escalated to Priya") is opaque to
anyone who has not memorised the cast, and naming on first mention only
breaks the moment a line is scrolled back to, grepped, or forwarded to
somebody who did not see the start of the run.
The identifiers themselves are NOT here. They live in internal/events and
are persisted into the append-only log, so they never change. This package
is a presentation layer applied at render time, which is what buys three
things at once: logs written before the names existed render with them,
renaming an agent later migrates nothing, and user configuration is a
change to this map rather than a change to the system.
The constraint that follows, and it is the important one: A NAME MUST
NEVER APPEAR ANYWHERE BUT THIS FILE. Not in a prompt, not in a Slack
template, not in a test fixture asserting on output. The moment "Ravi" is
written into a prompt, a team that renames the developer gets an agent
addressed by a name that appears nowhere in their config and nobody will
work out why. Prompts refer to roles; only the renderer knows names.
TestNoDefaultNameAppearsOutsideTheRegistry enforces it.
adoptβ
Package adopt wires Orion into an existing repository.
It replaces a four-step manual recipe (copy a config, make directories,
hand-edit .claude/settings.json, run doctor) with one idempotent command.
The hand-edited step was the dangerous one: settings.json usually already
contains a team's own hooks, permissions and MCP servers, and a copy-paste
that replaces the file destroys them silently.
Everything here is therefore additive and re-runnable. Adopting twice
changes nothing the second time, and nothing Orion did not put there is
ever removed.
adviseβ
Package advise answers a question the implementer could not resolve.
The loop this automates, described by the person who was doing it by hand:
Claude Code works a ticket, hits an ambiguity, stops and asks; the question
is carried to the model that designed the project; that model says which
option to take; the answer is carried back. Orion does the carrying.
The answerer is not the supervisor. A process controller has no knowledge
the implementer lacks. The answerer is a second agent that holds the
DESIGN -- and the design is not a chat history, it is intent.md, spec.md
and plan.md, committed in the repository. That is what makes this
automatable at all: the context the human was supplying from a ChatGPT
conversation is already on disk.
Three roles, because three different bodies of committed truth decide
three different questions:
architect reads spec.md and plan.md "how should this be built"
pm reads intent.md "what are we building, and why"
dba reads the schema and the migrations "how should this data be shaped"
One rule binds all three: DERIVE FROM THE ARTIFACT, CITE IT, OR REFUSE.
Refusal is a success. If an advisor invents an answer, the invention now
carries a citation-shaped wrapper and reads as authoritative -- two agents
in sequence launder a guess into something a reviewer will not question.
The honest outcome when the artifact is silent is that the artifact is
INCOMPLETE, which is a human's decision and then an amendment to the
document, so the next ticket does not re-ask.
agentcfgβ
Package agentcfg decides what a supervised run is allowed to be.
A `claude -p` child started from an operator's shell inherits that
operator's ENTIRE Claude Code configuration: their plugins, their MCP
servers, their subagents, their slash commands and whatever system-prompt
text a third-party plugin injects. Orion chose none of it, and it decides
what the agent can do, how it writes, and what a run costs. Measured on a
real run: 179 tools, 148 of them MCP tools against the operator's live,
authenticated SaaS accounts -- createJiraIssue and editJiraIssue among
them. Orion's breaker, sandbox and approval path govern the filesystem,
git and the network. They do not govern any of those.
Three problems, and the order matters because the fix is the same one:
blast radius a write handle to the tracker that decides what gets
worked, held by an agent working one ticket
reproducibility the same ticket on two machines is two different runs,
so an eval baseline measures the operator's plugin
folder as much as the change under test
cost tool definitions are re-sent every turn, and the
implementer runs 120 to 600 turns
THE FIX IS NOT ISOLATION. Orion DEPENDS on inherited configuration:
/pre-push-review, /pm-plan, /pr-describe and the rest of nj-agents arrive
by exactly this mechanism, and `orion doctor` grades a missing nj-agents
as FAIL. Blanking the config would break the stages this project delegates.
So the child gets its OWN directory, Orion-managed, populated with the
toolkit Orion actually depends on and nothing else -- a decision rather
than an inheritance.
Two levers, because one does not reach far enough:
CLAUDE_CONFIG_DIR what the directory holds: skills, agents, commands
and plugins. Curated here.
--strict-mcp-config MCP servers, which are also configured per account
and would survive a directory swap. With no
--mcp-config alongside it, the run gets none.
See docs/decisions/0014-supervised-runs-get-a-curated-config-directory.md,
including why the tracker's own MCP is not among them.
aiopsβ
Package aiops is the post-run triage pass: it reads one ticket's event log
AFTER the run has finished and reports what is worth a person's attention,
as a report plus draft tickets that nobody has filed.
It exists because nobody reads the logs during an unattended run. Every
defect Orion has found in itself so far -- OR-162's false exhaustion,
OR-167's swallowed QA findings -- was found by a human pasting a log to
somebody who read it, and that does not happen overnight.
Four design choices govern everything here, and each of them is load
bearing.
IT READS THE EVENT LOG, NOT THE TRANSCRIPT. internal/events is structured,
small, append-only and already carries actor, kind and key. The transcript
is tens of thousands of tokens of tool output that would have to be re-read
on every pass. Scan below takes a slice of typed events and nothing else,
so there is no path by which a transcript could get in.
MOST DETECTION NEEDS NO AGENT. Every rule in this file is a pure function
over typed events. A rule cannot hallucinate a breaker trip that did not
happen, cannot cost anything, and can be tested against a fixture. The
agent is reserved for the far smaller and more honest job of judging
whether a pattern NOTHING here recognises is worth reporting -- which is
what Concerning exists to decide, and why it can answer "nothing".
IT PROPOSES; IT DOES NOT FILE. Not a convention: the Open interface below
carries a search method and nothing else, so this package could not create
an issue if some later change asked it to. Three reasons the rule is
absolute. The backlog is already hard to scan and an autonomous filer makes
that worse fastest; a human decides what gets created, everywhere else in
this codebase; and, most of all --
ORION DEGRADES ON PURPOSE. A lock timeout proceeds unlocked and says so. A
QA failure is a warning, not a block. An absent gh is fine. So a pass that
matched on the word "error" would file tickets for behaviour that is
working exactly as designed, and after the third such ticket nobody would
read the fourth. EVERY rule below therefore carries a Why that says what
separates it from the designed degradation it most resembles, and the ones
deliberately NOT written are listed at notRules.
answersβ
Package answers is the inbox a person's answer travels through on its way
from the web page to a ticket.
The web process holds no tracker credential, by design: a page that could
write to Jira would be one more place a token lives. So the page writes an
answer HERE, as a file under ORION_HOME, and the watcher -- which already
holds the credential -- posts it as a comment and requeues the ticket. The
cost is latency of up to one watcher pass; the gain is that the one process
facing a browser can write nothing outside ORION_HOME.
One file per answer, never a shared one: two answers submitted in the same
moment must not be able to interleave into a mixture.
budgetβ
Package budget accounts for what Orion spends and stops it at thresholds.
WHAT THIS IS NOT, stated first because the distinction matters: this is not
your Anthropic subscription's weekly limit. There is no interface for
reading that. `claude --help` has no usage command, and a run's JSON result
reports what THAT run consumed, never what remains on the plan. Any
percentage Orion showed against the provider's real quota would be invented.
So this tracks a budget YOU set, over a rolling seven days, from figures
each run genuinely reports:
usage.input_tokens, usage.output_tokens, cache read/creation
total_cost_usd
modelUsage[model].contextWindow
Crossing a threshold stops the run and waits for a human. The thresholds
exist because the failure being prevented is unattended spend, and an
unattended process cannot be trusted to judge its own budget.
changelogβ
Package changelog collates per-ticket changelog fragments into CHANGELOG.md.
Every ticket used to append its entry to the same `## Unreleased` section of
the same file, so any two branches in flight conflicted there regardless of
what code they touched -- three tickets partitioned the code cleanly across
three packages and still blocked each other on CHANGELOG.md alone. With
strict branch protection the cost compounds: each merge invalidates every
other open pull request, and every one of those rebases stops on that file.
So an implementer writes `.changelog.d/<KEY>.md` instead. A new file per
ticket means two tickets never touch the same path and the conflict cannot
occur -- prevention rather than resolution, which matters because there is
direct evidence that resolving these by hand goes wrong: one hand-merged
conflict kept both sides (correct-looking, both were additive) and shipped
two `### Changed` sections carrying the same bullet to main. The merge
actually needed was "keep both, then merge same-named sections, then
deduplicate" -- three rules, applied consistently, on a file nobody reads
carefully during a rebase.
Nothing changes for a reader of CHANGELOG.md. Collation emits the same file
the old process produced, with the section order fixed rather than left to
whoever wrote the entry.
ciscaffoldβ
Package ciscaffold gives a repository one standard way to run its tests.
The contract is a single file: scripts/test.sh. CI runs it, a person runs
it, and the agent is told to run it before finishing. That one indirection
is the whole design -- when the three can diverge, they do, and the way it
shows up is a pull request that is green in CI and broken on the machine
that has to ship it.
It also makes the CI verdict mean something. Without a workflow, GitHub
reports no checks at all, and Orion has no honest choice but to treat an
unverified branch as passing, because the alternative is every ticket
waiting forever for a verdict nobody will ever produce. Scaffolding CI at
adoption is what turns "no checks configured" from the normal case into a
deliberate one.
Nothing here overwrites. A repository that already has a test script or a
workflow has one for reasons this package cannot see.
claimβ
Package claim records which process is working a ticket, so a claim left
behind by a dead one can be told apart from a claim that is doing work.
THE LABEL IS THE LOCK and it lives on the ticket, which is what lets a
restarted watcher see the same answer a second watcher sees. What the label
cannot say is whether anyone is still holding it: `orion-working` on a
ticket whose agent was killed looks exactly like `orion-working` on a ticket
mid-run, and the queue excludes both. So an interrupted ticket was stranded
until somebody removed the label by hand.
The tracker cannot answer this. internal/watch's existing staleness check
clears a claim only where the tracker categorises the ticket's status as
Done, which most workflows -- including this project's own -- do not.
A PID and a heartbeat can. Both are needed, and neither alone is enough:
- A PID alone is wrong across a reboot, where the number is reused by an
unrelated process and the claim reads as live forever.
- A heartbeat alone is wrong for a legitimately long run. An agent that
works for fifty-eight minutes is not stalled, and a watcher that stole
its ticket at thirty would be the more expensive bug.
Together they are decisive: a claim is dead when the process that wrote it
is gone, and only then. The heartbeat is the tiebreaker for the reboot case,
where the PID may have been reused.
collectβ
Package collect finishes what work started.
`orion work` deliberately stops at the pull request. It pushes, opens the
PR, marks the ticket orion-ci-wait and releases the job slot, because
holding an agent's process open for the twenty minutes CI takes would mean
one ticket at a time and a laptop that cannot be closed.
That leaves a gap nothing crossed: the CI verdict arrives, the merge
happens on GitHub, and Orion never hears about either. Tickets sat in
orion-ci-wait forever, the user's own checkout stayed behind develop, and
job worktrees accumulated for branches that had long since merged.
This is the other half. It reads the state that lives OUTSIDE Orion -- the
PR's checks, whether it merged -- and reconciles the tracker, the working
copy and the sandbox to it.
It is a poll, not a webhook, and deliberately so: a webhook needs a public
endpoint and a secret to receive an event that this reconstructs from
scratch in one API call. Polling is also idempotent, which matters more --
running it twice must be indistinguishable from running it once, because
a person will run it twice.
configβ
Package config loads orion.json from the project root and supplies
defaults for anything absent. Every value here is a hard limit enforced
by a hook, so an absent or malformed config must never silently widen
a control: parse failures fall back to defaults and are reported.
conflictβ
Package conflict verifies that a merge resolution dropped nothing.
WHY A GREEN BUILD IS NOT ENOUGH, with the evidence that made this a ticket.
On 2026-08-29 a three-way conflict between OR-171, OR-174 and OR-176 was
hand-resolved. The first resolution BUILT, VETTED and PASSED the package's
own tests -- and was still wrong: it had reverted OR-171's actor routing
back to a hardcoded implementer, and separately violated a property OR-176
had just introduced. It was caught only because OR-176 happened to have
added a test that failed.
So the suite passing is a FLOOR, not the proof. The failure mode of a
resolution is not "it breaks", it is "one side of a hunk quietly vanishes",
and a build cannot tell a line deliberately removed from one that was lost.
What this package proves mechanically:
- no conflict markers survived into the tree
- no .orig or .rej litter was left behind
- no file that BOTH sides changed came out byte-identical to one of them
That last one is the interesting check. If both branches edited a file and
the resolution matches one of them exactly, the other side's edit is not in
the result. That is occasionally correct -- one change genuinely superseded
the other -- but it is indistinguishable from a dropped change, so it is
reported rather than passed. Nothing here decides; it hands whoever must
answer for the resolution a specific file and a specific claim to check.
conformβ
Package conform answers the third question about a finished change: is
this the thing we agreed to build?
Three readers already look at a change and none of them asks it.
the review class reads the DIFF: is this code good
QA reads the ACCEPTANCE CRITERIA: does it do what the
ticket said
done triage reads the same criteria against the diff: is it
genuinely finished, or does it only look finished
All three can pass on a change that is well written, tested, satisfies its
ticket, and quietly builds something other than what was agreed during
planning. That is the expensive failure, because it is the one nobody is
looking for: every status is green and the divergence is only visible to
somebody who reads the plan and the diff side by side, which is exactly
what nobody does at approval time.
THE PLAN ARTIFACT IS THE ONLY THING IT COMPARES AGAINST. Not the ticket --
QA and done triage own that ground, and a third pass re-deriving the same
verdict would be two-thirds cost for no new information. Not style, not
tests, not whether the code is good. One question, one source of truth.
AND ONLY A CONFIRMED ONE. This is why the pass depends on internal/decide:
before that package there was no artifact in this repository that a later
stage could treat as agreed, because a recommendation nobody answered and a
decision somebody made were the same bytes in the same directory. Checking
conformance against an unconfirmed proposal would enforce a plan nobody
approved, which is worse than not checking at all.
IT REPORTS. It never blocks, never merges, never hands a ticket back, and
never edits anything -- the same posture QA has (OR-126: findings go to a
person, the pull request still opens). That is not timidity, it is the
point: a divergence is frequently an IMPROVEMENT the implementer found
while building, and a gate that stopped the pipeline for one would be
switched off within a week. What must not happen is that it lands
unremarked. So the verdict below has no boolean a caller could gate on;
there is a report, and a person reads it.
NOTHING TO CONFORM TO IS A STATED RESULT, NOT A PASS. A ticket with no
confirmed plan artifact is not one that matches its plan; it is one with no
plan, and the two are different facts. Reporting the first as the second
would make the pass look like it ran on every ticket while it silently ran
on almost none.
costβ
Package cost answers "what did this ticket cost" from the event log.
The only cost signal Orion used to produce was a line in a log file,
discovered after the fact. The ticket -- the artifact a person actually
revisits -- said nothing about what its automation spent, so deciding
whether to keep handing a class of work to agents meant reconstructing a
number by hand from run logs.
So every supervised run writes what it consumed into the event log, keyed
by the actor id, and this aggregates those lines per ticket. The event log
is the right home for two reasons: it is already per-workspace, append-only
and rotated, and the actor ids in it are the PERSISTED ones -- a renamed
agent still attributes correctly, because the display name is applied here
at render time rather than stored.
Three rules the aggregation exists to keep:
- EVERY run counts. The implementation run, each fix-loop re-entry, and
any later reviewer or exchange run. A report that quietly counts only
the last run is a report that says a ticket was cheap.
- A run that DIED still spent tokens. Max-turns, breaker, quota: the
usage is part of the ticket's true cost and is counted, and marked.
- Usage that is MISSING is said out loud. An older event log, or a run
that crashed before its result JSON, means the total is a floor. A
lowball number presented as complete is worse than an honest gap.
credsβ
Package creds resolves Orion's credentials from the environment or from
Orion's own config file.
Why a file at all, when environment variables already worked: a shell
profile is read by INTERACTIVE shells only. Not by cron, not by launchd,
not by a GUI-launched app. So credentials that work perfectly in a terminal
vanish the moment `orion report --notify` runs on a schedule, and the
failure looks like "Slack is not configured" rather than "your cron has a
different environment". Orion reading its own file removes the class.
Precedence is environment first, file second. That is the least surprising
order: an explicitly exported variable is a deliberate override for one
invocation, and a stored value should never win against it.
The file holds secrets, so it is created 0600 at open time rather than
chmod'ed afterwards. Writing world-readable and then narrowing leaves a
window where anyone on the machine can read the token.
dashboardβ
Package dashboard answers one question: is the coding queue outrunning the
integration queue?
Agents scale horizontally. Integration does not -- it is one operation at a
time by construction, roughly a CI run plus a human. So the number
that matters is not how fast agents code, it is whether coding is producing
work faster than integration can absorb it. Nothing reported that.
EVERYTHING HERE IS DERIVED FROM events.jsonl, and nothing is estimated. The
log already records every transition this needs: a ticket claimed, a stage
crossed, a branch pushed, a batch tested, a merge landed. A dashboard that
invented a second source for any of it would eventually disagree with the
log, and the log is the thing people trust when they are debugging at
midnight.
dbaβ
Package dba decides, deterministically and for free, whether a change
touches the database.
DELIBERATELY NOT A MODEL CALL, for the reason internal/work/route.go is not
one: this runs on every ticket, before anything has been spent, and its
whole purpose is that a ticket which touches no data does not pay for the
database stage. A classifier that costs a model call to decide whether to
spend has already spent, on every ticket, forever -- and it would be
non-deterministic about a gate whose value is that a reader can predict it.
Two inputs, because the change and the ticket each know something the other
does not. The DIFF is the ground truth for what was actually touched, and it
is unavailable until the implementer has finished; the TICKET's labels,
components and type are what a planner wrote down in advance, and they are
the only signal a ticket that has not run yet has. Either one alone misses a
real case: a "performance" ticket that ends up adding an index has no data
label, and a ticket labelled `database` whose fix turned out to be in the
cache layer touched no schema file.
A miss in either direction is survivable and they are not symmetric. A false
negative skips a stage that reports rather than blocks, so the change goes
to review as it does today. A false positive spends one agent run on a
change with no schema in it, which the report then says. So the rules below
are written to be READABLE rather than exhaustive: a signal nobody can
predict is worse than one that misses.
dbaplanβ
Package dbaplan is the database architect's part in PLANNING.
Three steps, in this order, and the order is the whole design:
recommend the database, with the reasoning, from the intent and the spec
confirm a person says yes in Slack, and only then is it a decision
design the initial schema, on the database that was agreed
NEITHER THE CHOICE NOR THE SCHEMA IS A PREMISE UNTIL IT IS CONFIRMED. That
is internal/decide's invariant and this is its highest-stakes instance: a
schema is inherited by everything written against it and is expensive to
reverse once there is data in it, which is why the database architect is a
separate actor at all rather than something the implementer decides in
passing. Nothing here writes into the confirmed directory itself; it calls
decide.Recommend, which writes the proposal to the pending directory that
no later stage is allowed to read.
THE SCHEMA IS NOT DESIGNED WHILE THE CHOICE IS UNCONFIRMED, and that is a
spend decision as much as a correctness one. A schema drawn against a
database nobody has agreed to is thrown away the moment they say no, and
until then it is a second unconfirmed document arguing for the first.
THE STATE IS READ OFF THE DISK, FOR FREE. Which of the three steps a run is
at is decided by which records exist and where they sit -- the same two
directories that carry the meaning -- so the stage is idempotent and the
operator drives it the way the flow actually works: run it, answer in
Slack, run it again. There is no state file, because a second record of
which step we are at is a record that can disagree with the records
themselves.
decideβ
Package decide keeps a RECOMMENDATION apart from a DECISION.
An agent that recommends something is proposing it. Every stage downstream
reads the committed artifacts as truth -- the implementer's prompt names
them, and internal/advise scopes an advisor to them and forbids it to
answer from anything else -- so a recommendation that lands in that set
unmarked stops being a proposal the moment it is written. Three stages
later it is a premise, cited back with a file and a clause, and nobody can
tell it apart from something a person actually agreed to. That is how a
project ends up built on a decision nobody made.
So the two states are STRUCTURAL, in two independent ways, because a
convention that lives only in prose is a convention some later change
quietly stops honouring:
the file says so -- "- Status: unconfirmed" versus "- Status: confirmed"
the directory says so -- PendingDir versus ConfirmedDir
and only ConfirmedDir is in scope for a role (see advise.Artifacts) or in
the implementer's prompt. An unconfirmed recommendation is therefore not
merely LABELLED as unsettled, it is outside the set of documents any later
stage is allowed to reason from. Deleting the label would not be enough to
launder it; the file would still have to be moved.
CONFIRMATION IS THE SLACK REACTION THAT ALREADY EXISTS. Not a second
approval mechanism: collect.ReadDecision, the merge_approvers allowlist,
the bot's own reactions excluded, a rejection beating every approval. A
second path would be a second place to rediscover that the bot approved
its own request -- and it would need its own allowlist, which is the part
nobody would keep in step.
The audit record is the confirmation APPENDED to the record, naming the
person and pointing at the Slack message it was given in. Pointing at,
rather than replacing: Slack holds who reacted, when, and everything said
around it, and a record that paraphrased all that would be a second
account of the same event that can disagree with the first.
decomposeβ
Package decompose turns a decomposed task list into tracker items.
Orion delegates decomposition to a skill today: the decompose stage prompt
tells the agent to run /pm-plan, which does two jobs -- break a plan into a
tree, then create that tree. Adopting spec-kit's /speckit.tasks replaces
the first job with a stronger artifact than a skill invents on the spot:
phased, with [P] parallel markers, story groups and exact file paths. What
is left is turning that artifact into tracker items, which is a
tracker-client problem rather than a skill-shaped one -- and Orion already
owns every other piece of the contract (the client, the transitions, the
label state machine, and the routing table `orion routes` publishes).
WHAT THIS IS NOT: it is not a replacement for the delegated path. The
decompose stage prompt is untouched, so a project with no spec-kit output
decomposes exactly as it did before. This is the native route for the
projects that DO have a tasks.md: `orion plan` runs it as the decompose
step when <FeatureDir>/tasks.md exists, stamping the queue label at the
levels Queue documents, and falls back to the supervised stage otherwise;
`orion decompose` runs the same code by hand -- see docs/decisions/0001,
which is also why nothing here decides whether a later stage runs.
SCOPE LIMIT, STATED RATHER THAN HIDDEN: the tracker seam OR-303 describes
-- an interface over the Jira client with Linear, Notion and GitHub
backends behind it -- has not landed. Everything above Backend below is
written against that interface and knows nothing about Jira; the only
implementation shipped is jira.go. Per OR-302's sequencing clause that
makes this a Jira-only capability, which is why it is opt-in and why
/pm-plan remains the default and fully available path for a project on any
other tracker.
discoveryβ
Package discovery decides whether an idea is ready to be designed from,
and gates the chain until it is.
The problem it fixes: the intent stage's prompt tells the agent to
"interrogate the idea the way an analyst would", but that stage runs
through `claude -p`, which is non-interactive. There was nobody to
interrogate, so an ambiguous sentence propagated unchallenged into spec,
plan, scaffold and a tracker tree. Each stage carries a token floor of
roughly 30k, so a wrong premise costs nine floors plus the rework, against
a discovery conversation costing about one.
The gate deliberately mirrors require_plan_before_edit rather than
inventing a second pattern: an artifact must be complete before the next
stage may read it.
doctorβ
Package doctor is the precheck.
Orion supervises tools it does not own: the Claude CLI, git, gh, a
sandbox provided by the OS. Every one of those can be missing,
unauthenticated or misconfigured, and each failure looks different at
the point of use. Diagnosing them once, up front, in plain language is
worth more than a stack trace forty minutes into a run.
Checks are graded: FAIL blocks, WARN degrades, OK passes. The exit code
is non-zero only for FAIL, so `orion doctor` is usable in CI as a gate.
doneβ
Package done answers one question about a finished run, at the moment
before a person is asked to approve it: is this genuinely done, or does it
only look done?
A green pull request is not evidence that a ticket is done. On 2026-08-30 it
was evidence of nothing three times, and each time a person caught it by
reading the run against the diff rather than by reading the status:
OR-217 green and mergeable. The test that caught its off-by-one was
written by QA and never committed, so CI tested a commit that did
not contain it. Committing the test by hand turned it red at once.
OR-116 green with three passing checks. QA had crashed (claude exited 1),
so the change went to review UNVERIFIED, and one of its test files
had been destroyed rather than stranded.
OR-229 green. Its own test passed at -count=1 and failed at -count=2:
under concurrency the fan paired answers with the wrong questions,
and one run happened to schedule the goroutines the way the
assertions expected.
THE MECHANICAL CHECKS DECIDE MOST OF IT, AND THEY GO FIRST. All three cases
above are visible in evidence that already exists -- an event QA emitted, a
file sitting in a worktree, a flag on a test command -- so Check below is a
pure function over that evidence and cannot hallucinate any of it. The
model is asked ONLY when the rules are clean, which is both the cheap path
and the honest one: paying a model to agree with a rule that already
answered buys nothing.
WHERE THE MODEL EARNS ITS PLACE is the one question no rule expresses --
whether the change matches its intent. The interesting failures on
2026-08-30 were each novel, and the value of this pass is catching the class
nobody has written a rule for yet.
IT REPORTS. It never merges, never approves, never edits the change. NOT
DONE hands the ticket back rather than holding a queue open, because a gate
that stalls is a gate somebody switches off.
NEVER A SCORE. Two values and no third. A number invites a threshold, and a
threshold turns "is this done" into "is this done enough", which is the
question nobody wanted asked.
eventsβ
Package events is Orion's live orchestration log.
Separate from the agent transcript on purpose. The transcript is what the
model said and did -- tens of thousands of tokens per run, already written
to .orion/logs/*.log. This is what ORION did: claimed a ticket, cut a
branch, asked the architect, escalated, opened a pull request. A dozen
lines where the transcript has thousands.
Merging them would be the obvious choice and the wrong one. The line that
matters ("architect answered: by issuer, per spec.md Β§4") becomes
invisible inside a wall of tool output, and the whole point of a live log
is that a person can glance at it and know where things stand.
JSON Lines rather than prose so the same file serves three readers: a
human watching `orion tail`, a Slack relay picking out the events worth
posting, and whatever dashboard comes later. Prose cannot be filtered.
exportβ
Package export ships Orion's events to the observability backend an
operator already runs, so a watch's history can be searched and charted
instead of living in a file per project.
ONE STANDARD, NOT ONE PLUGIN PER VENDOR. Grafana Cloud, Datadog,
Dynatrace, New Relic and Honeycomb all accept OTLP logs over HTTP, so a
vendor is a preset -- which header carries the key -- rather than code.
Loki's own push API is kept for a self-hosted Loki older than 3.x, which
has no OTLP endpoint.
NEVER IN THE WAY. Emit hands events over without blocking: they wait in a
bounded buffer, go out in batches from one goroutine, and are dropped --
with one warning -- when the buffer is full or the backend refuses them.
Observability that can stall a watch is a new way for the watch to fail.
fakebinβ
Package fakebin puts a shell-script stand-in for a real binary on PATH,
portably.
The tests in this repository fake the `claude` CLI (and occasionally other
tools) by writing a small shell script onto PATH. On unix that is the whole
job: the script IS an executable. Windows dispatches on the file extension
rather than the shebang, so the script alone is invisible to exec.LookPath
-- and the obvious shim, a .bat that hands the script to bash, corrupts its
arguments: cmd.exe's batch processing is LINE-oriented, so an argument
containing a newline -- which every multi-line prompt does -- is truncated
at the first one. That was measured, not theorised: a prompt of several
lines arrived as its first line and nothing else.
So on Windows the fake is a real PE executable with no cmd.exe anywhere in
the chain: a COPY OF THE RUNNING TEST BINARY. Main, called from the
package's TestMain, notices when the current process is such a copy -- the
script sits beside the executable under the same name -- and runs the
script with bash, forwarding argv unchanged. CreateProcess passes one flat
string and bash's own C-runtime parsing keeps quoted newlines intact, so
the arguments survive.
Used only from _test.go files; nothing in the product imports it.
fanoutβ
Package fanout decides whether an implementation run may be split across
concurrent subagents, and refuses when it may not.
The implementer works files serially, which is wasted wall time when a
ticket touches genuinely independent code. Splitting by FILE looks like the
obvious fix and is not: per-file ownership solves the write race and
neither of the two failures that actually bite.
builds are not isolated the compiler compiles the PACKAGE. A subagent
running tests to check its own work builds
against its peers' half-written files and sees
failures that are not its own.
signature coupling file A defines a function, file B calls it. Agent
B writes the call site against a signature agent A
is still changing -- the Idea-redeclared and
duplicate-CreateProject collisions seen at merge
time, moved to write time, where there is no
conflict marker to warn anyone.
So the unit is the Go package: it is the compilation unit, therefore the
real isolation boundary. Two files in one package are coupled by
construction; two files in different packages each build on their own, and
the remaining coupling is the import edge, which is visible and enumerable.
The implementer PROPOSES and this package DISPOSES. Nothing here asks a
model anything: the checks are set membership and a dependency graph, so a
wrong split fails visibly rather than corrupting a tree silently. Any
failure means serial, with no negotiation and no retry with a better
argument -- the same shape as the require_plan_before_edit gate.
See docs/decisions/0016-fan-implementation-by-go-package.md.
hookβ
Package hook implements the Claude Code hook protocol.
Contract: the harness writes a JSON object to stdin and reads the
process exit code.
exit 0 allow, stdout is shown to the user in transcript mode
exit 2 BLOCK, stderr is fed back to the model as the reason
other non-blocking error, stderr shown to the user, action proceeds
The distinction matters. A hook that crashes with exit 1 does not stop
anything, so every guardrail here is written to fail closed on the
decisions that matter and to exit 2 with a message a stranger can act
on. Unparseable input is the one exception: it exits 0, because a
malformed payload from a future harness version must not brick the
user's session. That tradeoff is deliberate and is called out in the
README under "Known gaps".
lessonsβ
Package lessons is Orion's cross-project memory.
The playbook's two-strike rule says a mistake seen twice becomes a line
in CLAUDE.md. That rule is per-repo, so the same mistake gets re-learned
in every new project. This package lifts the store to $ORION_HOME and
shares it across every workspace.
Two failure modes govern the design, and both are easy to walk into:
1. Over-generalization. "Money is BigDecimal here" is true of one repo
and wrong advice everywhere else. So a lesson starts project-scoped
and is promoted only when it actually recurs in a DIFFERENT project.
Promotion is evidence, never inference.
2. Context bloat. The agent reads CLAUDE.md in full every session, so
an unbounded lesson list degrades every session slightly and
invisibly. The rendered block is capped, ranked and expiring.
The store is append-only JSONL. Nothing is ever edited in place, so the
provenance of a rule the agent is following can always be reconstructed:
what happened, in which project, on which date.
localauthβ
Package localauth authenticates Orion's local web surface.
The surface is not read-only any more: it approves human gates, edits agent
config and starts runs by shelling out to the CLI. Binding to 127.0.0.1 is
not a permission boundary -- every other process running as the operator can
reach it, and so can any website the operator visits, by CSRF or by DNS
rebinding.
Three layers, all mandatory, applied as ONE middleware over the whole mux.
They are not alternatives; each covers a threat the others miss, and the
table of what each one does not cover is in
docs/decisions/0024-local-surface-authentication.md.
- A per-process token in a custom request header. The only layer that
stops another LOCAL process, because a local process is not a browser
and sets its own Origin. A custom header cannot be set by a cross-origin
HTML form, so it doubles as the CSRF defence.
- Origin and Sec-Fetch-Site, with the EMPTY header rejected. A check
shaped `if origin != "" && !loopback(origin)` fails open on exactly the
request an HTML-form CSRF sends.
- An exact-string Host allowlist, never a suffix match and never a DNS
lookup on the attacker-controlled value.
The mux is guarded by default: only paths written into the read-only
allowlist are exempt, and only for GET and HEAD. A write endpoint added
later -- or a GET that mutates -- is protected without anyone remembering
to protect it.
matchβ
Package match implements the glob subset Orion needs, including the
"**" segment that path.Match does not support. Kept dependency-free so
the binary builds offline with an empty go.sum.
newqβ
Package newq holds the questions `orion new` asks about an idea, in one place
so the terminal interview and the web form cannot ask different things.
notifyβ
Package notify tells the user something happened while they were not
watching.
Orion runs long and unattended. A quota wall at minute forty of an
unattended run is exactly the moment a person needs to know, and
exactly the moment nobody is looking at the terminal.
Every channel is best-effort and non-fatal: a notification that fails
must never take down the run it was reporting on. Failures are returned
so the caller can log them, never so the caller can abort.
procsafeβ
Package procsafe holds the two operations that have to stay correct when
SEVERAL Orion processes share one ORION_HOME.
That is the supported workflow: `orion watch A` and `orion watch B` in two
terminals, one per project repository. Ticket claiming is already safe
across processes -- the orion-working label is the lock and it lives on the
Jira ticket -- and workspaces and worktrees are per project and per ticket.
What is shared is the small amount of mutable state under ORION_HOME:
usage.json and repos.json.
Both were load-modify-save with an atomic rename but no lock, and both used
a FIXED temp path. Two processes therefore raced twice over: on the temp
file itself, and on the read-modify-write, where last-rename-wins silently
dropped one process's update.
The lock is a directory rather than flock(2). os.Mkdir is atomic on every
platform Orion ships to; flock is not available on Windows without cgo or
syscall build tags, and the release Makefile cross-compiles six targets
with CGO_ENABLED=0. internal/state has used exactly this approach for the
hook path since the beginning; this package is that implementation lifted
out so budget and registry can share it rather than grow a second one.
promoteβ
Package promote decides whether a milestone is safe to release.
"Thorough verification" is a word rather than a gate unless the checks are
named, so they are named here and each exists because of a specific way a
release has gone wrong or could:
1. Every ticket in the version is Done.
2. Fragments and the version reconcile in both directions.
3. The integration branch is green ON THE EXACT SHA being promoted.
4. No open pull requests target the integration branch.
5. Every commit in the promotion range is attributable to some ticket.
CHECK 5 WILL FIRE ON REAL WORK, and that is intended rather than a defect
to tune away. On 2026-08-29 a workflow fix and two changelog assemblies
were pushed by hand, outside any ticket. An unattributed commit is a thing
to LOOK AT, not necessarily a thing to stop for.
The split between blocking and warning is the whole design. A gate that
refuses for everything trains an operator to bypass it, and a gate that
warns about everything is not a gate. So: anything that would publish
something WRONG blocks; anything merely worth knowing warns.
provisionβ
Package provision creates the remote repository and its branch model.
Branch model: two long-lived branches. `main` is the release branch and is
protected. `develop` is the integration branch, the repository default,
and the base for every pull request. Feature branches are cut from develop
and merge back into it.
Both are push-protected, and that is not belt-and-braces. Protecting only
main would leave the PR into develop optional, and a review gate you can
skip is not a gate.
Everything here is idempotent: re-running against an already-provisioned
workspace reports what exists rather than failing or duplicating.
queueβ
Package queue decides which queued tickets the watch may claim on a pass.
Every queued ticket gets one verdict: admit it to a free worker, hold it
with a named reason, or evict it with a named reason and a record. Plan
is a pure function over facts it is handed, so the same queue gets the same
answer every time; it never reorders what the tracker's priority and rank
already decided.
Beside the planner:
prose dependencies "blocked by" written in a ticket's description,
which no issue link records
evictions ledger what was evicted, why and how often, so a ticket
evicted twice goes to a person instead of round again
scope ledger the files planning predicted a ticket would touch,
beside what the work actually touched
Deciding what enters the queue is sequencing across stages, which ADR 0001
makes Orion's job rather than an agent's.
quotaβ
Package quota detects model quota and rate-limit exhaustion, works out
when the limit resets, and decides whether to wait.
A caveat worth stating plainly: the wording of these errors is not a
stable contract. Providers change the text, and Claude Code wraps it
differently in different versions. So this package is built to fail
safely rather than cleverly:
- Detection is a list of patterns, not a single regex. A miss means
Orion treats a quota error as an ordinary failure, which is
annoying but correct.
- Parsing tries several shapes and falls back to capped exponential
backoff when none match, rather than inventing a reset time.
- Anything detected as exhaustion but unparseable gets its raw text
logged verbatim, so a new pattern can be added from evidence.
Never silently sleep for hours. Every wait is announced, logged, capped,
and recorded in task.json so a killed process can be resumed.
reconcileβ
Package reconcile compares what the repository says against what the
tracker says, and reports where they disagree.
The tracker is a record of intent, not evidence of what happened, and the
two drift apart silently. On 2026-08-30 OR-211 read In Progress for hours
while its finished work sat committed-but-never-pushed in a local worktree:
develop never received it, the milestone counted it as included, and
nothing anywhere noticed. Had a release been cut on the tracker's word, it
would have shipped a changelog claiming a fix that was not in the binary.
That is the failure this exists to catch, and it is the reason OR-116
(promote automatically once a milestone is complete) blocks on it: an
automated promotion is only as trustworthy as the signal it triggers on.
What it deliberately does NOT do: change the tracker. Every finding is
reported with the evidence that produced it, and acting on one is a
separate decision. A reconciler that silently rewrites status is a second
source of drift rather than a cure for the first.
registryβ
Package registry maps a tracker project key to the repository it belongs
to, so Orion can act on a ticket without being told where the code lives.
Why this exists. Every command so far resolved the repository from the
current directory, which works for a person standing in a checkout and not
at all for a daemon: `orion watch` has no meaningful cwd, and a ticket key
is the only thing it has to go on. The key already names the project, so
the mapping is the missing half.
It also catches a mistake that is otherwise silent. Two repositories both
claiming project KEY -- a clone, a rename, a copy made for an experiment
-- would each believe the queue is theirs, and work would land in whichever
one happened to run last. The registry refuses the second binding and says
which repository already holds it.
reportβ
Package report summarises what Orion has been doing.
It exists because per-run logs answer "what happened in this run" and
nothing answered "what has been happening". A supervisor you have to
interrogate one workspace at a time is one you stop checking.
The output is deliberately plain text with no colour or box drawing, so
the same bytes are readable in a terminal, in a cron mail, and in a Slack
message. A report that needs a renderer is a report that does not travel.
sessionβ
Package session tracks which orion processes are actually alive.
A SESSION IS A PROCESS, not a project and not a run. `orion watch` with no
arguments watches every project, so a session cannot be per-workspace; the
thing a person needs distinguished is the two terminals they started.
Runs appear inside a session -- a watcher works many tickets in sequence.
LIVENESS IS A HEARTBEAT FILE, not a last-event timestamp and not a PID
check. Neither alternative works:
- Last-event timestamp cannot separate quiet from gone. An agent between
turns and a watcher killed an hour ago both emit nothing.
- PID + signal-0 is meaningless on Windows, where os.FindProcess always
succeeds, and PIDs are reused on every OS, so a dead watcher's pid can
be occupied by something unrelated and read as alive.
A goroutine touches the file every ~15s, independently of whatever tick
the caller is on, so it keeps beating through a long agent run. A reader
treats the record as stale past a few multiples of that interval -- see
StaleAfter.
STOPPED SESSIONS VANISH. The file is removed on clean exit (Stop) and
swept when stale (Sweep, called at start, never by a reader -- see
Enumerate's own doc comment for why). History is not this package's job;
events.jsonl and task.json already hold it.
sessionsβ
Package sessions enumerates every workspace Orion could be running in.
A run view has to start by answering "what is there", and the answer lives
in two places that neither one alone gets right. The registry
(internal/registry) maps a tracker project key to the workspace that owns
it -- that is where the KEY comes from, and a workspace with no key is
unlabelled on a board. But `orion new` and `orion plan` create workspaces
under ~/.orion/projects before anything binds them, and a run that failed
before adoption never gets bound at all, so a scan of the registry alone
silently omits exactly the workspaces somebody is most likely looking for.
Reading the projects directory alone loses the key. So: both, unioned.
NOTHING HERE READS AN EVENT LOG. A Session says where a run is and what it
is called; what happened inside it comes off events.jsonl on its own
ticket. Splitting it that way keeps this pure -- given a home directory it
returns the same list every time -- and keeps the expensive part (parsing
a log per workspace) out of the cheap part (deciding which logs exist).
A MISSING SOURCE IS NOT AN ERROR. No registry and no projects directory is
what a fresh install looks like, and a board that refuses to draw on a
machine where nothing has run yet reports a normal state as a fault.
settingsβ
Package settings reads and writes the per-project switches and limits in
orion.json: the circuit breakers `orion config limits` sets and the landing
switches `orion config collect` sets.
It exists so the terminal and the web page share ONE validator and ONE
writer (docs/decisions/0026). The CLI keeps what is a conversation -- its
"are you sure" prompts -- and the page keeps what is a page: neither owns
the rules for which names exist, what a value may be, which block a field
lives in, or how the file is patched.
The patch is textual (adopt.SetOrAddBlockField), not a JSON round trip,
because re-marshalling reorders keys and scatters every "_comment_*" away
from the setting it explains.
slackβ
Package slack talks to the Slack Web API with a bot token.
Why this exists alongside the webhook in internal/notify: an incoming
webhook CANNOT create a channel. It is bound at creation to exactly one
channel and has no other API surface. A channel per project therefore
needs a real Slack app with a bot token, which is a different and heavier
integration, and saying so up front is kinder than discovering it after
setting one up.
Two behaviours of the Web API shape everything here:
It returns HTTP 200 on failure. An error arrives as {"ok":false,"error":
"name_taken"} with a 200 status, so any client that checks only the status
code reports success for every failure. Every call here checks `ok`.
Channel names are constrained: lowercase, no spaces or periods, at most 80
characters, and only letters, digits, hyphens and underscores. Slack
rejects anything else rather than normalising it, so normalisation happens
before the call.
stateβ
Package state persists per-session counters across hook invocations.
Hooks are separate processes: every fire is a cold start with no memory
of the last one. Loop detection and budget enforcement are therefore
only possible against durable state. Parallel worktree sessions may
share a state directory, so every read-modify-write is serialized by a
lock that works on every platform (os.Mkdir is atomic everywhere;
flock is not available on Windows without cgo or syscall build tags).
suiteβ
Package suite runs a repository's own test suite as a process Orion owns,
rather than as an instruction in an agent's prompt.
WHY THIS EXISTS. Until now Orion never invoked a test runner. It described
one in prose -- supervisor.prompts.go's testEnv() tells the agent which
command to use -- and left the agent to run it. That gives the agent three
decisions it should not have: WHAT to run, WHETHER to run it, and HOW to
report the result. A stage could go green because an agent ran a narrower
subset than it claimed, or narrated a pass it never saw. A process can do
none of those things: it exits 0 or it does not.
It is also the cheaper arrangement. A test run carries one bit of
interesting information plus text, so spending a model call on it buys
nothing, and the runner parallelises better than agents would -- `go test`
already runs packages concurrently, which is a thing no fan of subagents
can improve on.
WHAT THIS DELIBERATELY DOES NOT DO. It does not detect every ecosystem.
nj-agents' /review-tests-build already does that across Node, Python, Go,
Rust, JVM, Make and just, and reimplementing it here would trade that reach
for ownership -- the same bad trade this project declined for /pm-plan. So
detection here covers the shapes Orion can be certain about, and anything
else returns NotFound so the caller keeps the delegated path. Degrading to
the old behaviour is always allowed; degrading SILENTLY is not.
supervisorβ
Package supervisor runs `claude -p` inside a workspace and enforces
the limits a hook structurally cannot.
A hook only fires when the agent calls a tool. That leaves three real
failure modes uncovered:
wall clock an agent thinking in circles without calling tools
wedged proc a child that stops responding entirely
quota a provider limit, which is not the agent's fault at all
A parent process sees all three. This is the argument for Orion being a
binary rather than a set of hook scripts: the supervisor is the only
layer that can kill a run, and the only one that can wait out a quota
reset and resume.
testprocβ
Package testproc starts a subprocess from a test so that killing the test
kills everything it started.
WHY THIS EXISTS. On 2026-09-03 a working day was lost to 619 orphaned
.test binaries and a load average of 256. They were cleared four times and
came back each time. The cause was not Orion: its agent path has put every
child in its own process group since OR-195, and KillAll signals the group.
It was the TESTS. Around twenty of them spawn a long-lived binary -- the
built orion binary, `go test`, `go build` -- with a plain exec.Command, so
when the parent `go test` is killed (a harness timeout, a ctrl-c, a
cancelled CI job) the child is reparented to init and runs forever.
Nothing reaps it, and every measurement taken afterwards is worthless:
packages that take twenty seconds take ten minutes, and tests with a clock
in them fail for reasons that have nothing to do with the code.
So a test that spawns uses Command or Start from here. The child leads its
own process group, and t.Cleanup kills that group when the test ends --
including when it fails, panics or is killed, because Cleanup runs on all
three.
toolkitβ
Package toolkit locates, validates and provisions the toolkit a project
delegates to.
Orion delegates review, secret scanning, test/build verification, PR
authoring and PM decomposition to whichever toolkit the project configures.
Having one is a hard dependency, not a nicety: those stages have no
fallback, and faking them with a thinner substitute would be worse than not
running them. Which one it is, is the project's choice.
nj-agents is the shipped default, not the only one. WHAT a toolkit must
ship comes from the stages the project configures (orion.json
toolkit.stages), so a project delegating to its own skill repository is
validated against the skills it actually invokes rather than against
nj-agents' catalogue.
Everything specific to nj-agents -- CONVENTIONS.md, install.sh -- is
required only of nj-agents.
Two things about the real installation shape drove this design.
First, skills are installed as SYMLINKS into a runner's config directory,
pointing back at a clone. So the presence of ~/.claude/skills/<name> says
nothing about whether the toolkit behind it is intact. The link must be
resolved to find the clone root.
Second, the shared contract every review skill reads (CONVENTIONS.md and
its siblings) lives at that clone ROOT, two levels above the skill. A
check that only looks in the skills directory passes happily while the
file the skills depend on is missing.
trackerβ
Package tracker provisions and binds project-management projects.
Scope split, and the reason for it: the binary talks to the tracker over
REST, while the agent does decomposition through an MCP connector. Project
creation is infrastructure, and infrastructure wants to be deterministic,
idempotent and verifiable. Asking an agent to create a project gives you
none of those. Asking an agent to break a plan into stories plays to what
it is actually good at.
A warning that is built into the code rather than left in a doc: creating
one project per idea is what this is configured to do, and it accumulates.
Jira project keys are globally unique per instance, capped at 10 uppercase
characters, and a non-admin cannot delete a project once made. Key
collisions across many ideas are certain, not hypothetical, so derivation
resolves them deterministically rather than failing.
uiβ
Package ui renders Orion's terminal output.
Colour is an accessibility feature here, not decoration: `orion init` now
reports a dozen actions across git, Jira, Slack and dun, and a wall of
uniform text makes the one line that failed as quiet as the eleven that
worked. Colour is what makes a failure findable at a glance.
It is therefore applied to STATUS only -- did this work, did it not --
never to carry meaning colour is the sole conveyor of. Every line still
begins with a word (created, failed, WARNING) that says the same thing, so
the output reads identically when piped, logged, or seen by someone who
cannot distinguish the colours.
updateβ
Package update tells a person when the Orion they are running is old.
Orion is left installed and run for months, so a version gap is the
default state rather than an edge case -- and it used to be silent. A fix
that had landed and been released still behaved like the old binary,
because the old binary was what was installed, and nothing on screen said
so.
Three properties matter more than the notice itself:
- It never delays a command. The cache is read synchronously and
refreshed by a detached child process, so the command in front of the
user prints what was already known and never waits on a network call.
A goroutine cannot do this job: `orion status` exits in milliseconds
and would take the unfinished request with it, so the cache would
never refresh at all.
- Every failure is silence. Offline, rate limited, DNS, GitHub down --
none of that is something wrong with the user's machine, so none of it
produces a warning, an error, or a non-zero exit.
- It never fires in hook mode. That is enforced by the caller (see
showsUpdateNotice in cmd/orion), because hooks run on every matching
tool call and a notice there would be printed hundreds of times a day
into someone's editor.
watchβ
Package watch is the loop that removes the last manual step.
Everything else in Orion is already automatic between the moment a ticket
is claimed and the moment its branch merges. What remained was TRIGGERING:
a person typing `orion work KEY-7`, then later `orion collect`. This turns
that into a label on a ticket.
The design in one line: a tick does the cheap reconciling first, then tops
the running set back up to the cap, then sleeps.
Cheap first, because collect costs one API call per waiting ticket and can
finish work already paid for -- closing a merged ticket, pushing a CI fix,
asking for an approval. Starting new work before finishing old work would
mean paying to begin something while something else sat done-but-unclosed.
SEVERAL AT ONCE, up to limits.max_concurrent_tickets. This used to be
strictly one: collect, start one ticket, block until it finished. The
reasoning was that two agents in one repository fight over git, and that is
true of a shared checkout and not of per-job worktrees -- and agent
execution was never the bottleneck anyway. On 2026-08-29 a ticket's agent
work took 14m23s while approvals and merge serialisation cost hours, so a
queue that ran one at a time was idle for most of the day it was watching.
The cap defaults to two, not because two is timid but because everything
concurrency breaks is invisible at one and obvious at two: git against the
one shared clone (workspace/gitlock.go), a budget checkpoint sailed past by
runs already in flight (budget/admit.go), one rate limit reached by N
sessions at once (limitPause below), and N tickets picked that all edit the
same files (pick below). Prove it at two, then raise it; five is the
ceiling, and it is a ceiling rather than advice.
What concurrency does NOT fix is the approval bottleneck: N tickets
finishing means N approvals waiting on one human, so without auto-merge the
queue finishes faster and lands no faster. That is worth knowing before it
is mistaken for a regression.
The state is on the tickets, not in this process. A watcher killed
mid-flight loses nothing that matters: the labels say what was claimed,
what is awaiting CI, and what failed, so a restart resumes rather than
repeats. That is what makes it safe to run this on a laptop that closes.
webβ
The snapshot endpoint: what `GET /api/snapshot` returns.
SCANNED PER REQUEST, never cached. A cached snapshot is a second copy of
the truth the event log already holds, and the two disagree the moment one
is stale -- the same argument model.go makes for why nothing here reads a
clock. The cost of re-reading every workspace's log on every request is the
price of a page that never lies about what just happened.
sessions.Scan enumerates the WORKSPACES (registered or not); web.Scan turns
one workspace's events into CARDS. Two different things named Scan in two
packages, which is the collision flagged when this file was written -- see
the session's own notes. Left as-is here rather than renamed in this
change: renaming either one is a decision with its own blast radius across
callers this ticket does not own.
The config endpoint: what GET /api/config returns.
THE ROSTER ONLY, DELIBERATELY. Roster (agents.go) already resolves
actors.Roster against the global agents.json -- one file per machine, the
same call `orion config agents --list` makes, so the page and the terminal
listing cannot disagree. This file adds only the route and envelope.
LIMITS FROM orion.json ARE NOT SERVED HERE. config.Load(root) reads one
PROJECT's orion.json, but this surface spans every workspace under
ORION_HOME with no single project selected -- there is no "the" orion.json
to read the way there is one agents.json. Serving limits would mean either
guessing a project (wrong the moment more than one exists) or inventing a
project-selection query parameter no other endpoint in this package has.
Left for the ticket that actually threads a project id through requests,
rather than guessed at here.
The detail endpoint: what GET /api/detail?key=..&run=.. returns.
SCANNED PER REQUEST, same rule api.go states for the snapshot: a click
into one ticket re-reads that workspace's log rather than trusting
whatever the last snapshot poll happened to hold, so opening the detail
panel always shows what actually happened, not a stale copy.
The history endpoint: what GET /api/history returns.
PAST RUNS FROM task.json's OWN Runs FIELD, never re-derived from the
event log. workspace.RunRec already carries exactly what the ticket asks
for -- stage, start time, duration, exit code, stop reason -- because the
supervisor writes one every time a run ends (internal/supervisor). A
second derivation from events.jsonl would be the same duplication api.go
and agents.go both refuse elsewhere in this package: two sources for one
question, disagreeing the moment either drifts.
Package web is the model behind `orion web`.
It holds what a browser needs to draw a run and nothing else: no server, no
event reader, no clock. The types come first, and deliberately, because
every other ticket in the epic -- the reader that fills them, the handler
that serves them, the page that draws them -- names these fields, and a
field renamed after three consumers exist is renamed in four places.
THE MOCKUPS ARE THE CONSUMER. docs/design/web/01-run-view.html draws a card
grid: one card per ticket, each carrying who is on it, on what model, how
far in, what it is doing right now, and how long it has taken. Each field
below exists because that page shows it; a field the page does not show is
not here yet.
EVERY STRING BELOW IS UNTRUSTED. Title is whatever a tracker
accepted, Activity and Gate come off logs agents and other tooling write,
and the board is the control plane once the write endpoints exist. So they
are plain strings and stay plain strings: rendered through html/template's
contextual escaping, never converted to template.HTML/JS/URL, which say
"already escaped" about the one kind of value that never is. escaping_test.go
holds both halves of that.
NOTHING HERE READS A CLOCK OR A FILE. A snapshot is derived from the event
log the same way internal/dashboard derives its view: every
timestamp came out of an event, so the same log always produces the same
snapshot, and a test can state an expected duration rather than tolerate
one. The moment Elapsed called time.Now, a paused run would keep ticking
and two readers of one log would disagree.
The live stream endpoint: what `GET /api/stream` pushes.
SERVER-SENT EVENTS, not a WebSocket. The traffic is one-directional --
Orion tells the page what happened, the page never talks back -- and SSE
is plain HTTP: no upgrade handshake, no extra dependency, and it survives
the loopback-only posture server.go already commits to without
adding a second protocol to that surface's threat model.
EVERY WORKSPACE, FANNED IN. The snapshot endpoint draws the whole
machine's cards from every workspace sessions.Scan finds; the stream
endpoint owes the same promise for the log panel underneath them. One
goroutine follows one workspace's log each, all writing into one channel
the handler drains -- the shape a blocking, per-file events.Follow forces
on anything that wants to watch more than one file from a single request.
workβ
Package work runs one tracker issue end to end.
The order of operations is the design. Every step that touches something
outside Orion -- the tracker, the remote, the user's checkout -- happens
only after the cheap, local, reversible steps have succeeded, so a run that
is going to fail fails before it has changed anything anyone else can see.
resolve which repository owns this key (registry, free)
preflight budget, sandbox, clean base (local, free)
merged? has this ticket's PR already landed (forge, free)
claim ORION -> orion-working, To Do -> Progress (tracker, reversible)
worktree a branch nothing else can touch (local, reversible)
run the supervised agent (COSTS MONEY)
verify commits exist (local)
push the branch reaches the remote (public)
pr open it (public)
ci-wait release the job slot (tracker)
Claiming before running rather than after is deliberate: two runs must not
pick up the same ticket, and the label is the lock. Pushing only after
commits exist is the other half -- an agent that stopped to ask a question
exits 0 having produced nothing, and pushing an empty branch would open a
pull request describing no change.
The merged check sits above the claim because the lock is only as good as
the label, and a label survives its ticket: see noop.go.
workspaceβ
Package workspace provisions isolated project directories.
Every task Orion is given gets its own directory tree, its own git
repository, its own generated settings and its own state. Nothing an
agent does in one workspace can reach another, and nothing it does
reaches the user's other work at all.
Isolation here is layered, and the layers are worth naming honestly:
directory a dedicated tree under $ORION_HOME, never the user's cwd
git a fresh repo, so a bad commit cannot touch existing history
settings generated per workspace: permission denies plus the OS
sandbox (Seatbelt on macOS, namespaces on Linux)
The OS sandbox stops credential reads and network egress. It is not a
VM. For genuinely untrusted code, run Orion itself inside a VM or
container. Task.Container is a leftover of a --container mode that nothing
sets any more; whether to remove it is a separate decision.