nj-agents and other toolkits
Orion delegates review, security scanning, testing and PR authoring to nj-agents. Its secret scanner is required, with no heuristic fallback, and security findings get an adversarial verification pass.
Orion sequences the stages, and a toolkit supplies the method inside one.
nj-agents is the default; spec-kit is the default for four planning steps of
a new project; and the toolkit block lets a project point stages
elsewhere.
What is delegatedβ
| Concern | Delegates to |
|---|---|
| Test planning and authoring | /test-plan, /test-author, /test-suite-author |
| Running tests and builds | /review-tests-build |
| Test repair and triage | /test-repair, /test-triage, /flake-watch |
| End-to-end tests | /e2e-suite, /e2e-run |
| Coverage gaps | /test-gap-finder |
| Review | /pre-push-review (an umbrella over five review dimensions) |
| Security, standard | /review-secrets (a hard gate that blocks on a hit) |
| Security, high risk | /security-deep-review (see the cost note below) |
| Pull request description | /pr-describe (draft only) |
| Commit messages | /commit-assistant |
| CI failure triage | /test-triage, /flake-watch |
| Release notes | /changelog, /release-notes |
| Decomposition | /pm-plan (Epic to Stories to Tasks) |
These stages have no fallback, so orion doctor FAILs when the toolkit is
missing.
The contract with a delegated skillβ
A delegated skill advises. It never runs git push and never bypasses a
hook. Orion runs it non-interactively, with NJ_AGENTS_CI=1, and acts on its
exit code:
| Exit code | Verdict |
|---|---|
0 | PASS or WARN |
| non-zero | BLOCK |
Ambiguity resolves to BLOCK, so a suspected secret blocks instead of waiting
for confirmation. Reports go to
${NJ_AGENTS_REPORT_DIR:-<repo>/.nj-agents-reports}/. Whether to merge, and
what runs next, stays Orion's decision.
Cost interactionβ
/security-deep-review runs roughly 5 finder calls plus up to 15 verifier
calls, commonly 15 to 30 agent calls for a mid-size diff. Orion's breaker
defaults to 400 tool calls per session and counts every one, so a deep
review late in a session could trip the breaker mid-review.
Two settings in the delegation block of orion.json handle it:
| Setting | Default | Effect |
|---|---|---|
extra_tool_calls_for_review | 200 | widens the tool-call envelope while a delegated review runs. Deliberately generous, because too small a number means a review that cannot finish. |
deep_security_review_when | high-risk | runs the deep pass only when the change touches high-risk paths, and the standard /review-secrets gate otherwise. always if the cost is acceptable, never to disable. |
The high-risk paths are delegation.high_risk_paths, by default: auth,
security, crypto, payment and migrations directories, anything with
deserial in its name, Dockerfiles and Terraform (*.tf).
Finding the checkoutβ
Orion looks for the checkout in this order:
delegation.nj_agents_dirinorion.json;- the
ORION_NJ_AGENTS_DIRenvironment variable; - resolving an installed skill's symlink back to its clone;
- Orion's own managed clone under
$ORION_HOME/vendor.
Your copy always wins. Orion reads a global install, possibly the
repository you develop nj-agents in, and never writes to it. Orion's own clone, fetched by orion doctor --fix, is Orion's to
maintain.
orion njagents status # where it is, which commit, how stale
orion njagents update # fast-forward Orion's own clone; for a global install, prints the git pull to run yourself
orion njagents install --project # wire Orion's clone into a directory, only when there is no global install
orion doctor resolves each skill's symlink back to its clone, where the
shared contract the skills read lives.
How the skills reach the agentβ
On Linux, a supervised run sees only the toolkit's skills and agents from
the checkout orion doctor finds, and no MCP servers, through a
configuration directory at $ORION_HOME/agent-config. A skill installed only
in your own ~/.claude is invisible to a run. On macOS, runs inherit your own Claude
Code configuration; see
The sandbox.
delegation.inherit_operator_config lists stages or actors that get your
own configuration on purpose (empty by default).
Using a different toolkitβ
The optional toolkit block in orion.json names the skill repository Orion
delegates to, and what each stage invokes inside it. Without it, the
repository falls back to nj-agents, and a stage with no command declared runs
Orion's own built-in prompt.
{
"toolkit": {
"repo": "https://github.com/navjyotnishant/nj-agents.git", // the default
"ref": "v1.4.0", // pin a tag; empty clones the default branch
"dir": "/path/to/nj-agents", // an existing clone; overrides the vendor path
"stages": {
"review": "/pre-push-review",
"pr": "/pr-describe"
}
}
}
A toolkit may fill only some stages:
{
"toolkit": {
"repo": "https://github.com/github/spec-kit.git",
"stages": {
"spec": "/speckit-specify",
"plan": "/speckit-plan"
}
}
}
The stage names are the stages Orion runs: intent, constitution, spec
(or design), plan, analyze, ticket, scaffold, decompose, build
(or implement), verify (or test), review, pr (or ship). Either
spelling of a pair means the same stage. Naming a stage twice with two
different commands is refused, with both keys quoted, and so is a stage name
Orion does not run.
stages is a map of what each stage runs. Sequencing is Orion's, so a block
that expresses order (a list, an order key or a sequence key) is rejected
with an error citing
ADR 0001.
Use the command names the installer put on disk, which are not always what a
toolkit's documentation calls them. spec-kit's README writes
/speckit.specify, while its Claude integration installs the skill as
speckit-specify.
Orion clones a toolkit it manages into <ORION_HOME>/vendor/<repo-name>, so
two toolkits never share a directory; toolkit.dir overrides that entirely.
Background:
ADR 0019
and
ADR 0020.
spec-kitβ
A project orion plan creates delegates constitution, spec, plan and
analyze to spec-kit. spec-kit is a Python CLI whose commands become skills
only through its own installer, so cloning it is not enough.
-
orion plan's first step installs it per project: it runsspecify initin the workspace repository when any stage names aspeckit-*command, and commits what it wrote. You install only the CLI (Install). -
It is pinned to v1.0.4: Orion reads skill names, the
[NEEDS CLARIFICATION]marker, the analyze report's critical-issue count and the constitution template's slots from what that version installs.orion doctorprints the installed version beside the pin. An existing workspace keeps the templates it was initialised with untilspecify integration upgradeis run inside it. -
For an existing repository, run the same command the chain does:
cd <repo> && specify init --here --force --non-interactive --integration claudeThat writes
.claude/skills/speckit-*/SKILL.md, which Orion and Claude Code both read.orion doctorthen FAILs anytoolkit.stagescommand with no file behind it, naming the stage.
Orion uses only spec-kit's commands, inside its own stages, and not its workflow engine, extension hooks or bundles: ADR 0021, ADR 0022.
The tracker tree from a spec-kit task listβ
orion decompose turns the specs/<nnn-feature>/tasks.md that spec-kit's
/speckit-tasks leaves (phases, [P] parallel markers, [USn] story groups
and file paths) into the tracker tree, with no skill in the middle:
orion decompose KEY # finds specs/*/tasks.md
orion decompose KEY specs/003-feature/tasks.md
- One Epic, one Story per
[USn]group, and each task as a child of its story. A task in no story group (Setup, Foundational, Polish) hangs off the Epic directly. - The
[P]marker, the phase, the dependency section and the file paths survive into the descriptions. Orion sets the routing marker on each item. - The whole tree is printed first,
+for new and=for what a previous run made, and one answer covers all of it. With nobody present to answer, the run creates nothing and says so. - A re-run finds earlier items by their
orion-spec-<feature>label, links them and creates only the rest, so the same command resumes a run that failed halfway.
orion plan runs this as its decompose step. The native tree is Jira-only
for now: a project whose tracker.provider names another tracker is refused
and told to decompose through the stage, where the /pm-plan fallback works
on any tracker.
Not verified yetβ
- Whether
/security-deep-reviewhonours a diff scope when driven non-interactively. Theextra_tool_calls_for_reviewenvelope assumes a diff-scoped run; a full-repository review would cost far more. - The exact argument syntax each skill accepts when invoked through
claude -p.