Skip to main content
Version: Next

Budget and monitoring

Orion spends money every time an agent runs. You can cap that spend, read what happened without watching, and send Orion's events to your own observability backend.

Weekly budget checkpoints​

Orion accounts for what it spends over a rolling seven days and stops for confirmation at 50%, 75%, 90% and 95% of your limit.

orion budget status # spend, tokens, next checkpoint, recent runs
orion budget ack # confirm the current checkpoint and continue

Set the limit in orion.json:

"budget": { "weekly_usd": 40, "weekly_tokens": 0, "pause_at_percent": [50, 75, 90, 95] }
  • Zero means unlimited here. That is the opposite of the circuit-breaker convention, where zero restores a default.
  • Acknowledging one checkpoint covers only that one; the next threshold stops again.
  • A crossed checkpoint stops orion plan before it dispatches anything, and the web dashboard refuses to start work while the budget is spent.

Your budget and your plan's limit​

This budget is yours, separate from your Anthropic plan's weekly limit. Orion cannot read the plan's limit: claude has no usage command, and a run's result reports what that run consumed, never what remains on the plan. So Orion shows no percentage against the provider's quota.

What a run costs​

Every claude -p invocation carries a floor of roughly 30k input tokens. Measured, claude -p "say ok" reports about 34k input tokens and about $0.19, almost all of it system prompt and cached context. That is paid per invocation, so nine small stages cost nine floors before any work happens. Prefer fewer, larger stages over many trivial ones.

Orion cannot trigger compaction. It rarely needs to, since every stage is a separate claude -p run that reads committed artifacts, so context resets at each stage boundary. Within one long stage context can still climb: when a run's input passes 70% of the model's context window, Orion says so and suggests splitting the stage.

The digest​

orion report # failures, workspaces, budget, usage
orion report --since 24h # a narrower window (7d, 24h, 90m)
orion report KEY # one project
orion report --notify # also send it to your webhook

The digest leads with what is actionable: quota-parked workspaces, an unacknowledged budget checkpoint, failed runs with their log paths.

orion report exits 1 when something needs a person and 0 otherwise, so cron stays quiet unless there is a problem:

0 9 * * * /opt/homebrew/bin/orion report --notify

Orion reads its credentials from ~/.orion/config.env itself, so this works under cron and launchd, where your shell profile is not read.

Logs and post-run triage​

CommandShows
orion logs KEYwhat Orion is doing on a ticket; -f follows it live
orion logs KEY --actor implementeronly that agent's lines
orion logs KEY --transcriptthe raw agent output instead
orion aiops KEYreads a finished run's event log and reports what is worth filing, with draft tickets. It proposes only and never creates anything; --no-agent uses rules alone
orion dashboardwhether coding is outrunning integration: queue depth, batch cost, CI runs saved
orion status <id>a workspace's stage, breaker state and last run, including a pause on quota and when it resumes

The watch's own log is kept at ~/.orion/logs/watch-*.log, and each workspace keeps a full transcript per run and its events.jsonl under .orion/.

Quota walls​

On a provider limit, Orion parses the reset time, waits, retries, and tells you. For a wait longer than 90 minutes, Orion records a resume time and hands back instead of holding a process open. An estimated wait is always labelled as an estimate. In a watch, a ticket that hits a quota wall is held, not failed, and costs no retry.

Detection is a list of known error wordings. A wording it does not recognise is treated as an ordinary failure.

Exporting events to an observability backend​

Off by default. Create ~/.orion/observability.json:

{"enabled": true, "preset": "grafana",
"endpoint": "https://otlp-gateway-<zone>.grafana.net/otlp/v1/logs"}
PresetFormatCredentials (environment variables)
grafanaOTLPGRAFANA_CLOUD_INSTANCE_ID, GRAFANA_CLOUD_TOKEN
datadogOTLPDD_API_KEY
newrelicOTLPNEW_RELIC_LICENSE_KEY
dynatraceOTLPDT_API_TOKEN
honeycombOTLPHONEYCOMB_API_KEY
lokiLoki push APILOKI_USER, LOKI_TOKEN (both optional for a self-hosted Loki)
otlpOTLPnone; add headers with headers_env
  • Credentials come only from the environment, never from the file. user_env and token_env rename a preset's variables; headers_env maps a header name to the variable holding its value, for the generic otlp preset.
  • include_tool_events: true also ships every agent tool call, which is most of the volume a backend bills for. orion logs KEY already has them locally.
  • Events carry project, ticket, actor, model and verdict, as service.name = orion.
  • Shipping never blocks the watch. A backend that refuses, or a full buffer, drops events with one warning. Anything credential-shaped is scrubbed before it leaves the machine.

The design is in ADR 0035.

Not documented yet​

  • The full field list of an exported event.
  • How cost per run is computed.