A Blueprint’s rules see only what a Package passes to check(). A Package
can check one customer and still return another customer’s data, and
neither its tests nor the compiler need notice. The tag and the check
agree, and a test that asks for a customer’s charges passes when it gets
too many. An agent that reads the Package with that question in mind can.
This guide shows you how to review a Package’s authorization with a coding agent. Verify a Package in CI walks through one review end to end with Codex.
Run a review
Section titled “Run a review”From the Package project, with the agent installed and signed in:
submilli build security-review -a codex -m gpt-6.1-sol -e high --fail-on high --output review.json-a picks the agent, -m its model, and -e how hard it reasons.
--fail-on is the lowest severity that fails the command, and
--output saves the report to a new file. The command refuses to
overwrite one. -p @acme/billing limits the review to one Package and
the local Packages it depends on.
| Agent | -a |
Install | Sign in locally | Model, for example |
|---|---|---|---|---|
| Codex | codex |
npm install -g @openai/codex@0.160.0 |
codex login |
gpt-6.1-sol |
| Claude Code | claude |
npm install -g @anthropic-ai/claude-code@2.1.288 |
claude auth login, or ANTHROPIC_API_KEY set |
claude-opus-5-5 |
The agent runs with its tools turned off and sees only the Package’s
source, which goes to that agent’s model provider under your account’s
terms. Install the agents from a pinned version, as here, because the review runs
the executable it finds on PATH.
Read the result
Section titled “Read the result”The exit code is the verdict:
| Exit | Meaning |
|---|---|
0 |
The review finished, with no finding at or above --fail-on |
1 |
The review finished, with at least one such finding |
2 |
The review didn’t finish because the agent couldn’t sign in, timed out, or didn’t cover every file |
The findings print with their file, line, evidence, and fix. The report adds what was reviewed, down to a hash of every file:
jq '{status, agent, agent_version, model, files, findings}' review.json{ "status": "complete", "agent": "codex", "agent_version": "codex-cli 0.160.0", "model": "gpt-6.1-sol", "files": { "capabilities.yaml": "88ac5694770e8ac47c8de59f0f7413ab729c243c856a075a13592cb4338cdb83", "docs/readme.md": "cd2a343d7d5b078ebcb806ce7cc2fdd7c740421a00e5d078fda576965e808ef2", "src/lib.ts": "1774da90940199f002442c372b0b1ba592da9d434c3ba5edd3d20b9c90a19248", "submilli.toml": "489b52f52df72d837b686d3199344d2cf0af1224ef228c4203d53b4cea605b0f", "tests/lib.test.ts": "f122f65346cdf5325160697bb9873f5bd3f3f5e195dacae6b96b3592de617dcd" }, "findings": []}Keep the report with the revision it reviewed, because the hashes say
which source the verdict covers. A Package that depends on a Package
from another repository reviews incomplete, with a coverage gap for the
dependency, and exits 2. Review that dependency in its own project.
Require it in CI
Section titled “Require it in CI”The review job in Verify a Package in
CI
runs on every pull request and push to main, installs the agent and a
pinned Submilli release, reviews the Package, and uploads the report
even when the review fails. For Claude Code,
the install, the credential, and the agent and its model change:
| Agent | Install step | The review step’s environment |
|---|---|---|
| Codex | npm install -g @openai/codex@0.160.0 |
CODEX_API_KEY: ${{ secrets.CODEX_API_KEY }} |
| Claude Code | npm install -g @anthropic-ai/claude-code@2.1.288 |
CLAUDE_CODE_OAUTH_TOKEN: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }} |
The review step then names the agent and its model:
- name: Install Claude Code run: npm install -g @anthropic-ai/claude-code@2.1.288 # … install Submilli, as in the tutorial - name: Review authorization env: CLAUDE_CODE_OAUTH_TOKEN: ${{ secrets.CLAUDE_CODE_OAUTH_TOKEN }} run: >- submilli build security-review -a claude -m claude-opus-5-5 -e high --fail-on high --output "$RUNNER_TEMP/security-review-${{ github.run_id }}-${{ github.run_attempt }}.json"Give the credential to the review step only, and require the review job
in the branch’s protection rules. Don’t add continue-on-error. A
review that can’t sign in exits 2 and should fail the check, not skip
it.
Codex in CI uses an OpenAI API key, billed to the API project, not to
a ChatGPT subscription. Create one under API
keys and save it as the
CODEX_API_KEY repository secret. A ChatGPT subscription works only on
a private repository’s self-hosted runner that keeps Codex’s sign-in
between jobs, which run one at a time, and never restore an older
sign-in over the one Codex refreshed. OpenAI excludes public
repositories from it. See Codex
automation.
Claude Code in CI uses your Claude subscription (Pro, Max, Team, or
Enterprise) through a long-lived token. Create it, after claude auth login, and save it as a repository secret. gh prompts for the value,
so it never lands in your shell history:
claude setup-tokengh secret set CLAUDE_CODE_OAUTH_TOKENDon’t also set ANTHROPIC_API_KEY for the step, because Claude Code
prefers it and bills the API instead. Renew the token when it expires. See Claude
authentication.
Make deployment wait for it
Section titled “Make deployment wait for it”Review the Package’s own repository, at the revision you deploy. Reviewing
the repository that holds your Blueprints doesn’t inspect the Packages it
installs. When the Packages live elsewhere, as in Manage Blueprints in
Git, pin a commit in
packages.txt only after its review passed, and keep that review’s report
with the deployment.
When the review and the deployment are jobs of one workflow, make the deployment need both checks, and keep its condition:
deploy: if: github.event_name == 'push' needs: [check, review]A merge commit is a different revision from the one the pull request reviewed. The workflow’s push to main reviews it again before the deployment runs.