Budget and monitoring
Orion spends money every time an agent runs. You can cap that spend, read what happened without watching, and send Orion's events to your own observability backend.
Weekly budget checkpointsβ
Orion accounts for what it spends over a rolling seven days and stops for confirmation at 50%, 75%, 90% and 95% of your limit.
orion budget status # spend, tokens, next checkpoint, recent runs
orion budget ack # confirm the current checkpoint and continue
Set the limit in orion.json:
"budget": { "weekly_usd": 40, "weekly_tokens": 0, "pause_at_percent": [50, 75, 90, 95] }
- Zero means unlimited here. That is the opposite of the circuit-breaker convention, where zero restores a default.
- Acknowledging one checkpoint covers only that one; the next threshold stops again.
- A crossed checkpoint stops
orion planbefore it dispatches anything, and the web dashboard refuses to start work while the budget is spent.
Your budget and your plan's limitβ
This budget is yours, separate from your Anthropic plan's weekly limit. Orion
cannot read the plan's limit: claude has no usage command, and a run's
result reports what that run consumed, never what remains on the plan. So
Orion shows no percentage against the provider's quota.
What a run costsβ
Every claude -p invocation carries a floor of roughly 30k input tokens.
Measured, claude -p "say ok" reports about 34k input tokens and about $0.19, almost all of it system prompt and cached
context. That is paid per invocation, so nine small stages cost nine floors
before any work happens. Prefer fewer, larger stages over many trivial ones.
Orion cannot trigger compaction. It rarely needs to, since every stage is a
separate claude -p run that reads committed artifacts, so context resets at
each stage boundary. Within one long stage context can still climb: when a
run's input passes 70% of the model's context window, Orion says so and
suggests splitting the stage.
The digestβ
orion report # failures, workspaces, budget, usage
orion report --since 24h # a narrower window (7d, 24h, 90m)
orion report KEY # one project
orion report --notify # also send it to your webhook
The digest leads with what is actionable: quota-parked workspaces, an unacknowledged budget checkpoint, failed runs with their log paths.
orion report exits 1 when something needs a person and 0 otherwise,
so cron stays quiet unless there is a problem:
0 9 * * * /opt/homebrew/bin/orion report --notify
Orion reads its credentials from ~/.orion/config.env itself, so this works
under cron and launchd, where your shell profile is not read.
Logs and post-run triageβ
| Command | Shows |
|---|---|
orion logs KEY | what Orion is doing on a ticket; -f follows it live |
orion logs KEY --actor implementer | only that agent's lines |
orion logs KEY --transcript | the raw agent output instead |
orion aiops KEY | reads a finished run's event log and reports what is worth filing, with draft tickets. It proposes only and never creates anything; --no-agent uses rules alone |
orion dashboard | whether coding is outrunning integration: queue depth, batch cost, CI runs saved |
orion status <id> | a workspace's stage, breaker state and last run, including a pause on quota and when it resumes |
The watch's own log is kept at ~/.orion/logs/watch-*.log, and each
workspace keeps a full transcript per run and its events.jsonl under
.orion/.
Quota wallsβ
On a provider limit, Orion parses the reset time, waits, retries, and tells you. For a wait longer than 90 minutes, Orion records a resume time and hands back instead of holding a process open. An estimated wait is always labelled as an estimate. In a watch, a ticket that hits a quota wall is held, not failed, and costs no retry.
Detection is a list of known error wordings. A wording it does not recognise is treated as an ordinary failure.
Exporting events to an observability backendβ
Off by default. Create ~/.orion/observability.json:
{"enabled": true, "preset": "grafana",
"endpoint": "https://otlp-gateway-<zone>.grafana.net/otlp/v1/logs"}
| Preset | Format | Credentials (environment variables) |
|---|---|---|
grafana | OTLP | GRAFANA_CLOUD_INSTANCE_ID, GRAFANA_CLOUD_TOKEN |
datadog | OTLP | DD_API_KEY |
newrelic | OTLP | NEW_RELIC_LICENSE_KEY |
dynatrace | OTLP | DT_API_TOKEN |
honeycomb | OTLP | HONEYCOMB_API_KEY |
loki | Loki push API | LOKI_USER, LOKI_TOKEN (both optional for a self-hosted Loki) |
otlp | OTLP | none; add headers with headers_env |
- Credentials come only from the environment, never from the file.
user_envandtoken_envrename a preset's variables;headers_envmaps a header name to the variable holding its value, for the genericotlppreset. include_tool_events: truealso ships every agent tool call, which is most of the volume a backend bills for.orion logs KEYalready has them locally.- Events carry project, ticket, actor, model and verdict, as
service.name = orion. - Shipping never blocks the watch. A backend that refuses, or a full buffer, drops events with one warning. Anything credential-shaped is scrubbed before it leaves the machine.
The design is in ADR 0035.
Not documented yetβ
- The full field list of an exported event.
- How cost per run is computed.