Safety and limits
Orion refuses some actions outright and enforces most of those refusals with hooks. It also has known gaps, listed at the end.
What Orion will not doβ
- Cut a release unattended.
orion watchhas no path to any release command, and promotingdeveloptomainis always a person's decision (Releasing a milestone). - Merge agent work into the release branch.
vcs.work_branchmust differ fromvcs.default_branch; config load,orion initandorion doctorrefuse them being equal unless you set the named opt-invcs.allow_release_branch_merges. - Let an agent push to
mainordevelop, force-push, or hard-reset onto a remote ref. Thegatehook refuses all three. - Let an agent deploy to production without a named authorization. The
gatehook refuses a production deploy unlessORION_RELEASE_APPROVALis set to an approval reference, and Orion never grants one. - Let an agent edit its own controls. The
shieldhook refuses edits to protected paths: CI configuration (.github/workflows/**), managed settings andorion.jsonitself. - Let an agent weaken the test that defines a fix. In fix mode, test files are read-only to the agent.
- Overwrite a ticket description on an agent's say-so. The agent posts the
before and after, and you approve with a label
(
orion-desc-*). - Treat an agent's recommendation as a decision. A recommendation stays out of every later stage's reach until a named approver confirms it (Advisors and decisions).
- Create a tracker tree without showing it first. The whole tree is
previewed and one yes creates it. Only
orion plan --yesrun off a terminal answers for you, which the web dashboard's Plan button uses after you confirm there. - Accept commands from Slack. Orion runs no Slack listener and reads back only approvals on its own request messages (Slack and approvals).
The guardrail hooksβ
The hooks check by literal matching, with no model judgement involved.
| Hook | Stops |
|---|---|
| breaker | identical repeated calls, repeated failure of one command, consecutive failures, polling with no progress, the tool-call budget, edits without verification, blast radius, wall clock |
| gate | production deploys with no named authorization, pushes to main or develop, force pushes, a hard reset onto a remote ref |
| shield | edits to protected paths, edits to test files during a fix, implementation before a plan exists (when the plan gate is on) |
- The breaker runs in supervised runs only;
gateandshieldrun everywhere, including your own Claude Code sessions in an adopted repository (Circuit breakers). - The breaker is wired to both PreToolUse and PostToolUse: one counts what happened, the other refuses the next call.
- Every block message names what to do instead.
Isolationβ
Every agent runs in its own worktree, inside an OS sandbox (Seatbelt or bubblewrap) that refuses to start rather than run unsandboxed, with a network allowlist and credential files and variables denied. The sandbox is not a VM: it stops credential reads and egress but not a determined exploit, so run Orion itself in a VM or container for code you do not trust. On macOS, runs inherit your own Claude Code configuration, MCP servers included, because a curated configuration cannot sign in there. The sandbox has the detail.
What leaves the machineβ
| Destination | What, and when |
|---|---|
Anthropic (api.anthropic.com) | every agent run, through the claude CLI |
| GitHub | repository creation, pushes, pull requests and CI, through git and gh |
| Your tracker | REST calls made by the Orion binary, not by agents: projects, tickets, labels, comments |
| Slack | posts to the project's channel; reading reactions and replies on Orion's own approval requests |
| Your webhook or command | only if you set ORION_NOTIFY_WEBHOOK or ORION_NOTIFY_COMMAND |
| Your observability backend | only if ~/.orion/observability.json enables it; credentials from the environment, credential-shaped values scrubbed first |
| GitHub releases API | a daily check for a newer Orion release; off with ORION_NO_UPDATE_CHECK=1 |
| GitHub, for toolkits | cloning nj-agents when orion doctor --fix fetches it, and toolkits Orion vendors |
| Package registries | whatever the agent's build fetches, within the sandbox allowlist |
Commits carry an AI-Attribution trailer when attribution is on (the
default). It names the agent and model, and travels wherever the repository
is pushed.
What stays localβ
~/.orionholds workspaces, state, watch logs, the spend ledger, and a full agent transcript per run.orion doctorwarns if the directory is readable by other users.~/.orion/config.envholds your tokens at mode 0600.orion config showmasks them, and the binary reads them and never passes them to an agent.- The web dashboard listens on
127.0.0.1only, requires a per-run token on every write, and holds no tracker credential (The web dashboard). - Nothing is pruned. Logs, transcripts and
events.jsonlaccumulate, and there is no retention policy yet.
Known limitsβ
Read these before relying on Orion.
- Unparseable hook input exits 0 and allows the call, so a malformed payload
from a future harness version cannot brick a session. A harness change
could therefore silently disable enforcement, and
orion doctordoes not detect it. - Hooks are not given token counts, so the token budget is approximated by tool-call count. Real token accounting is available only after the fact.
- Quota detection matches a list of provider error patterns, and that wording can change. A miss treats the quota wall as an ordinary failure.
- State locking is best effort, with a 3-second timeout, after which it proceeds unserialized and says so, so a wedged lock cannot block every tool call.
- Branch protection fails on free-plan private GitHub repositories: GitHub
returns 403, and provisioning reports it. The
gatehook still refuses agent pushes tomainanddevelop, but a person with a terminal is unconstrained until server-side protection exists. - A green batch lands without a person when
slack.merge_approversis empty, and green then means only as much as the project's tests and evals. The guidance is to name approvers untilevals/holds at least 20 real cases; nothing in code enforces that. - Jira projects and Slack channels accumulate, since a non-admin cannot
delete a Jira project and a bot cannot delete a channel. Bind to one
existing project with
tracker.project_keyif you run many small ideas. - The audit log is local and editable, and not tamper-evident.
- There is no OS sandbox on Windows. Workspaces refuse to start; use WSL2 or a Linux VM.
- The weekly budget is Orion's own. Orion cannot see how much of your provider plan's limit remains.
The Roadmap lists what is planned and what is deliberately not.