Skip to main content
Version: Next

Safety and limits

Orion refuses some actions outright and enforces most of those refusals with hooks. It also has known gaps, listed at the end.

What Orion will not do​

  • Cut a release unattended. orion watch has no path to any release command, and promoting develop to main is always a person's decision (Releasing a milestone).
  • Merge agent work into the release branch. vcs.work_branch must differ from vcs.default_branch; config load, orion init and orion doctor refuse them being equal unless you set the named opt-in vcs.allow_release_branch_merges.
  • Let an agent push to main or develop, force-push, or hard-reset onto a remote ref. The gate hook refuses all three.
  • Let an agent deploy to production without a named authorization. The gate hook refuses a production deploy unless ORION_RELEASE_APPROVAL is set to an approval reference, and Orion never grants one.
  • Let an agent edit its own controls. The shield hook refuses edits to protected paths: CI configuration (.github/workflows/**), managed settings and orion.json itself.
  • Let an agent weaken the test that defines a fix. In fix mode, test files are read-only to the agent.
  • Overwrite a ticket description on an agent's say-so. The agent posts the before and after, and you approve with a label (orion-desc-*).
  • Treat an agent's recommendation as a decision. A recommendation stays out of every later stage's reach until a named approver confirms it (Advisors and decisions).
  • Create a tracker tree without showing it first. The whole tree is previewed and one yes creates it. Only orion plan --yes run off a terminal answers for you, which the web dashboard's Plan button uses after you confirm there.
  • Accept commands from Slack. Orion runs no Slack listener and reads back only approvals on its own request messages (Slack and approvals).

The guardrail hooks​

The hooks check by literal matching, with no model judgement involved.

HookStops
breakeridentical repeated calls, repeated failure of one command, consecutive failures, polling with no progress, the tool-call budget, edits without verification, blast radius, wall clock
gateproduction deploys with no named authorization, pushes to main or develop, force pushes, a hard reset onto a remote ref
shieldedits to protected paths, edits to test files during a fix, implementation before a plan exists (when the plan gate is on)
  • The breaker runs in supervised runs only; gate and shield run everywhere, including your own Claude Code sessions in an adopted repository (Circuit breakers).
  • The breaker is wired to both PreToolUse and PostToolUse: one counts what happened, the other refuses the next call.
  • Every block message names what to do instead.

Isolation​

Every agent runs in its own worktree, inside an OS sandbox (Seatbelt or bubblewrap) that refuses to start rather than run unsandboxed, with a network allowlist and credential files and variables denied. The sandbox is not a VM: it stops credential reads and egress but not a determined exploit, so run Orion itself in a VM or container for code you do not trust. On macOS, runs inherit your own Claude Code configuration, MCP servers included, because a curated configuration cannot sign in there. The sandbox has the detail.

What leaves the machine​

DestinationWhat, and when
Anthropic (api.anthropic.com)every agent run, through the claude CLI
GitHubrepository creation, pushes, pull requests and CI, through git and gh
Your trackerREST calls made by the Orion binary, not by agents: projects, tickets, labels, comments
Slackposts to the project's channel; reading reactions and replies on Orion's own approval requests
Your webhook or commandonly if you set ORION_NOTIFY_WEBHOOK or ORION_NOTIFY_COMMAND
Your observability backendonly if ~/.orion/observability.json enables it; credentials from the environment, credential-shaped values scrubbed first
GitHub releases APIa daily check for a newer Orion release; off with ORION_NO_UPDATE_CHECK=1
GitHub, for toolkitscloning nj-agents when orion doctor --fix fetches it, and toolkits Orion vendors
Package registrieswhatever the agent's build fetches, within the sandbox allowlist

Commits carry an AI-Attribution trailer when attribution is on (the default). It names the agent and model, and travels wherever the repository is pushed.

What stays local​

  • ~/.orion holds workspaces, state, watch logs, the spend ledger, and a full agent transcript per run. orion doctor warns if the directory is readable by other users.
  • ~/.orion/config.env holds your tokens at mode 0600. orion config show masks them, and the binary reads them and never passes them to an agent.
  • The web dashboard listens on 127.0.0.1 only, requires a per-run token on every write, and holds no tracker credential (The web dashboard).
  • Nothing is pruned. Logs, transcripts and events.jsonl accumulate, and there is no retention policy yet.

Known limits​

Read these before relying on Orion.

  • Unparseable hook input exits 0 and allows the call, so a malformed payload from a future harness version cannot brick a session. A harness change could therefore silently disable enforcement, and orion doctor does not detect it.
  • Hooks are not given token counts, so the token budget is approximated by tool-call count. Real token accounting is available only after the fact.
  • Quota detection matches a list of provider error patterns, and that wording can change. A miss treats the quota wall as an ordinary failure.
  • State locking is best effort, with a 3-second timeout, after which it proceeds unserialized and says so, so a wedged lock cannot block every tool call.
  • Branch protection fails on free-plan private GitHub repositories: GitHub returns 403, and provisioning reports it. The gate hook still refuses agent pushes to main and develop, but a person with a terminal is unconstrained until server-side protection exists.
  • A green batch lands without a person when slack.merge_approvers is empty, and green then means only as much as the project's tests and evals. The guidance is to name approvers until evals/ holds at least 20 real cases; nothing in code enforces that.
  • Jira projects and Slack channels accumulate, since a non-admin cannot delete a Jira project and a bot cannot delete a channel. Bind to one existing project with tracker.project_key if you run many small ideas.
  • The audit log is local and editable, and not tamper-evident.
  • There is no OS sandbox on Windows. Workspaces refuse to start; use WSL2 or a Linux VM.
  • The weekly budget is Orion's own. Orion cannot see how much of your provider plan's limit remains.

The Roadmap lists what is planned and what is deliberately not.