Skip to main content
Version: Next

nj-agents and other toolkits

Orion delegates review, security scanning, testing and PR authoring to nj-agents. Its secret scanner is required, with no heuristic fallback, and security findings get an adversarial verification pass.

Orion sequences the stages, and a toolkit supplies the method inside one. nj-agents is the default; spec-kit is the default for four planning steps of a new project; and the toolkit block lets a project point stages elsewhere.

What is delegated​

ConcernDelegates to
Test planning and authoring/test-plan, /test-author, /test-suite-author
Running tests and builds/review-tests-build
Test repair and triage/test-repair, /test-triage, /flake-watch
End-to-end tests/e2e-suite, /e2e-run
Coverage gaps/test-gap-finder
Review/pre-push-review (an umbrella over five review dimensions)
Security, standard/review-secrets (a hard gate that blocks on a hit)
Security, high risk/security-deep-review (see the cost note below)
Pull request description/pr-describe (draft only)
Commit messages/commit-assistant
CI failure triage/test-triage, /flake-watch
Release notes/changelog, /release-notes
Decomposition/pm-plan (Epic to Stories to Tasks)

These stages have no fallback, so orion doctor FAILs when the toolkit is missing.

The contract with a delegated skill​

A delegated skill advises. It never runs git push and never bypasses a hook. Orion runs it non-interactively, with NJ_AGENTS_CI=1, and acts on its exit code:

Exit codeVerdict
0PASS or WARN
non-zeroBLOCK

Ambiguity resolves to BLOCK, so a suspected secret blocks instead of waiting for confirmation. Reports go to ${NJ_AGENTS_REPORT_DIR:-<repo>/.nj-agents-reports}/. Whether to merge, and what runs next, stays Orion's decision.

Cost interaction​

/security-deep-review runs roughly 5 finder calls plus up to 15 verifier calls, commonly 15 to 30 agent calls for a mid-size diff. Orion's breaker defaults to 400 tool calls per session and counts every one, so a deep review late in a session could trip the breaker mid-review.

Two settings in the delegation block of orion.json handle it:

SettingDefaultEffect
extra_tool_calls_for_review200widens the tool-call envelope while a delegated review runs. Deliberately generous, because too small a number means a review that cannot finish.
deep_security_review_whenhigh-riskruns the deep pass only when the change touches high-risk paths, and the standard /review-secrets gate otherwise. always if the cost is acceptable, never to disable.

The high-risk paths are delegation.high_risk_paths, by default: auth, security, crypto, payment and migrations directories, anything with deserial in its name, Dockerfiles and Terraform (*.tf).

Finding the checkout​

Orion looks for the checkout in this order:

  1. delegation.nj_agents_dir in orion.json;
  2. the ORION_NJ_AGENTS_DIR environment variable;
  3. resolving an installed skill's symlink back to its clone;
  4. Orion's own managed clone under $ORION_HOME/vendor.

Your copy always wins. Orion reads a global install, possibly the repository you develop nj-agents in, and never writes to it. Orion's own clone, fetched by orion doctor --fix, is Orion's to maintain.

orion njagents status # where it is, which commit, how stale
orion njagents update # fast-forward Orion's own clone; for a global install, prints the git pull to run yourself
orion njagents install --project # wire Orion's clone into a directory, only when there is no global install

orion doctor resolves each skill's symlink back to its clone, where the shared contract the skills read lives.

How the skills reach the agent​

On Linux, a supervised run sees only the toolkit's skills and agents from the checkout orion doctor finds, and no MCP servers, through a configuration directory at $ORION_HOME/agent-config. A skill installed only in your own ~/.claude is invisible to a run. On macOS, runs inherit your own Claude Code configuration; see The sandbox.

delegation.inherit_operator_config lists stages or actors that get your own configuration on purpose (empty by default).

Using a different toolkit​

The optional toolkit block in orion.json names the skill repository Orion delegates to, and what each stage invokes inside it. Without it, the repository falls back to nj-agents, and a stage with no command declared runs Orion's own built-in prompt.

{
"toolkit": {
"repo": "https://github.com/navjyotnishant/nj-agents.git", // the default
"ref": "v1.4.0", // pin a tag; empty clones the default branch
"dir": "/path/to/nj-agents", // an existing clone; overrides the vendor path
"stages": {
"review": "/pre-push-review",
"pr": "/pr-describe"
}
}
}

A toolkit may fill only some stages:

{
"toolkit": {
"repo": "https://github.com/github/spec-kit.git",
"stages": {
"spec": "/speckit-specify",
"plan": "/speckit-plan"
}
}
}

The stage names are the stages Orion runs: intent, constitution, spec (or design), plan, analyze, ticket, scaffold, decompose, build (or implement), verify (or test), review, pr (or ship). Either spelling of a pair means the same stage. Naming a stage twice with two different commands is refused, with both keys quoted, and so is a stage name Orion does not run.

stages is a map of what each stage runs. Sequencing is Orion's, so a block that expresses order (a list, an order key or a sequence key) is rejected with an error citing ADR 0001.

Use the command names the installer put on disk, which are not always what a toolkit's documentation calls them. spec-kit's README writes /speckit.specify, while its Claude integration installs the skill as speckit-specify.

Orion clones a toolkit it manages into <ORION_HOME>/vendor/<repo-name>, so two toolkits never share a directory; toolkit.dir overrides that entirely. Background: ADR 0019 and ADR 0020.

spec-kit​

A project orion plan creates delegates constitution, spec, plan and analyze to spec-kit. spec-kit is a Python CLI whose commands become skills only through its own installer, so cloning it is not enough.

  • orion plan's first step installs it per project: it runs specify init in the workspace repository when any stage names a speckit-* command, and commits what it wrote. You install only the CLI (Install).

  • It is pinned to v1.0.4: Orion reads skill names, the [NEEDS CLARIFICATION] marker, the analyze report's critical-issue count and the constitution template's slots from what that version installs. orion doctor prints the installed version beside the pin. An existing workspace keeps the templates it was initialised with until specify integration upgrade is run inside it.

  • For an existing repository, run the same command the chain does:

    cd <repo> && specify init --here --force --non-interactive --integration claude

    That writes .claude/skills/speckit-*/SKILL.md, which Orion and Claude Code both read. orion doctor then FAILs any toolkit.stages command with no file behind it, naming the stage.

Orion uses only spec-kit's commands, inside its own stages, and not its workflow engine, extension hooks or bundles: ADR 0021, ADR 0022.

The tracker tree from a spec-kit task list​

orion decompose turns the specs/<nnn-feature>/tasks.md that spec-kit's /speckit-tasks leaves (phases, [P] parallel markers, [USn] story groups and file paths) into the tracker tree, with no skill in the middle:

orion decompose KEY # finds specs/*/tasks.md
orion decompose KEY specs/003-feature/tasks.md
  • One Epic, one Story per [USn] group, and each task as a child of its story. A task in no story group (Setup, Foundational, Polish) hangs off the Epic directly.
  • The [P] marker, the phase, the dependency section and the file paths survive into the descriptions. Orion sets the routing marker on each item.
  • The whole tree is printed first, + for new and = for what a previous run made, and one answer covers all of it. With nobody present to answer, the run creates nothing and says so.
  • A re-run finds earlier items by their orion-spec-<feature> label, links them and creates only the rest, so the same command resumes a run that failed halfway.

orion plan runs this as its decompose step. The native tree is Jira-only for now: a project whose tracker.provider names another tracker is refused and told to decompose through the stage, where the /pm-plan fallback works on any tracker.

Not verified yet​

  • Whether /security-deep-review honours a diff scope when driven non-interactively. The extra_tool_calls_for_review envelope assumes a diff-scoped run; a full-repository review would cost far more.
  • The exact argument syntax each skill accepts when invoked through claude -p.