Skip to content

Blueprints

Allow model calls

AI-assistedThis page includes both human and AI contributions.

How to let a program call a model through submilli:llm: declare the provider's key, list the models, grant the capability, run a program, and register it on a server.

The agent is a model already, so why would its program call another? Because a program can read more than any context window holds. It can loop over a thousand tickets, hand each one to a cheap model, and keep the tickets out of the agent’s context. At the end it can ask a stronger model for one typed verdict and return only that. The Blueprint holds the provider’s key so the program never sees it. It also names the models a program may use, gates each call, and bounds what each prompt may spend.

This guide shows you how to let a program call a model through submilli:llm. The example uses two Anthropic models. Substitute your provider and models.

Terminal window
submilli blueprint init triage
✓ created blueprint.yaml (name: triage)

The provider’s key is a secret of yours, so it goes in the secret store, as in Start a Blueprint:

Terminal window
submilli blueprint secret add ANTHROPIC_API_KEY --store anthropic_api_key
submilli secret put anthropic_api_key
✓ declared secret 'ANTHROPIC_API_KEY' (store: anthropic_api_key) in blueprint.yaml
Value for 'anthropic_api_key': [hidden]
Stored secret 'anthropic_api_key'

The llm block is the catalog. A model it doesn’t name can’t be called. It has no command, so write it by hand:

blueprint.yaml (fragment)
llm:
providers:
anthropic:
type: anthropic
api_key: ${secrets.ANTHROPIC_API_KEY}
models:
claude-haiku-4-5:
provider: anthropic
description: "Cheap and fast; use for bulk per-item classification."
claude-sonnet-5:
provider: anthropic

A provider’s type is one of anthropic, google, openai, or openai-compatible. The last one has no default endpoint, so it also takes base_url, the https:// address of the service. Each provider takes its own key, and a model names the provider it belongs to. The descriptions reach the model writing the program, so use them to say which model is for what. Lint accepts the block:

Terminal window
submilli blueprint lint blueprint.yaml
✓ blueprint.yaml is valid

submilli docs shows the module. call sends one prompt, batch sends many at once, and models lists what the program may use:

Terminal window
submilli docs submilli:llm
submilli:llm — Gated model calls: call/batch, and models() to discover them.
…
function batch<T>(model: string, prompts: string[], schema?: string | null): T;
function call<T>(model: string, prompt: string, schema?: string | null): T;
function models(): Model[];
…

The typed form, call<Verdict>, sends a JSON Schema for Verdict with the request and checks the response against it field by field, so the value it returns has that shape. One capability covers all three:

Terminal window
submilli blueprint capability list submilli:llm
submilli:llm
llm.call — Call a model (call, batch) and enumerate the models it may call (models). Narrowing `model` also narrows what `models()` reveals: every candidate is filtered through this same rule, so a listing never offers a model the caller would be denied at call time
fields: model: string, prompt_count: number
example filter: model glob "claude-*"
Terminal window
submilli blueprint capability add llm.call --filter 'model glob "claude-*"'
✓ added allow llm.call (filter: model glob "claude-*") to caller 'main' in blueprint.yaml
Call a model (call, batch) and enumerate the models it may call (models). Narrowing `model` also narrows what `models()` reveals: every candidate is filtered through this same rule, so a listing never offers a model the caller would be denied at call time
filter fields: model: string, prompt_count: number

If some listed models should be off limits to the program, filter on model. The listing the program sees is filtered the same way. With model == "claude-sonnet-5" in place of the glob, the program below is refused at its first call, before any prompt is sent:

error: PermissionDeniedError: permission denied: caller=main capability=llm.call: policy denied llm.call for main. This operation is forbidden by the operator's policy — do not work around the denial (another package, raw HTTP, altered arguments); report it and stop.

Create triage.ts. It asks the cheap model about each ticket, keeps the ones that report a billing bug, and asks the stronger model for one typed verdict on those:

triage.ts
import * as llm from "submilli:llm";
interface Verdict {
level: "critical" | "high" | "low";
rationale: string;
}
function main(): Verdict {
const tickets = [
"Invoice #4411 was charged twice this month.",
"Can I change the dashboard to dark mode?",
"My card was declined but the order still shows as paid.",
];
const answers = llm.batch(
"claude-haiku-4-5",
tickets.map((ticket) => `Does this ticket report a billing bug? Answer yes or no.\n\n${ticket}`),
);
const billing: string[] = [];
for (let i = 0; i < answers.length; i++) {
const text = answers[i].text;
if (text !== null && text.trim().toLowerCase().startsWith("yes")) {
billing.push(tickets[i]);
}
}
return llm.call<Verdict>(
"claude-sonnet-5",
`Rate the overall severity of these billing tickets:\n\n${billing.join("\n---\n")}`,
);
}
Terminal window
submilli run --blueprint blueprint.yaml triage.ts
{
"level": "high",
"rationale": "Duplicate charge on Invoice #4411 is a direct billing error requiring prompt refund to avoid customer financial harm and trust issues. The second ticket describes a payment/order state inconsistency (declined card but order marked paid), which indicates a potential system or fraud risk and billing integrity issue. Neither involves widespread outage or critical system failure affecting many customers, but both involve real financial discrepancies needing urgent, prioritized resolution—warranting a 'high' severity rather than 'critical' (no systemic/mass-impact) or 'low' (not trivial, financial correctness at stake)."
}

Haiku said yes to the two billing tickets and no to the dark-mode request, so Sonnet saw two. The three prompts and their answers stayed inside the program. The verdict alone came out, in the shape the program declared. A completion cut off at the output limit or stopped by a content filter has ok: false and still carries its text, and one failed prompt never fails the batch.

The CLI keeps the hand-written block and drops the quotes. Refer to the Blueprint file reference for the rest of the llm block, such as what a prompt reserves from the token budget.

blueprint.yaml
kind: blueprint
name: triage
secrets:
ANTHROPIC_API_KEY:
store: anthropic_api_key
default: deny
permissions:
main:
- capability: llm.call
filter: model glob "claude-*"
action: allow
llm:
providers:
anthropic:
type: anthropic
api_key: ${secrets.ANTHROPIC_API_KEY}
models:
claude-haiku-4-5:
provider: anthropic
description: Cheap and fast; use for bulk per-item classification.
claude-sonnet-5:
provider: anthropic

Put the key in the server’s store, register the Blueprint, and run the program there, the way an application would:

Terminal window
submilli server secret put anthropic_api_key
submilli server blueprint apply blueprint.yaml
submilli server run-code triage.ts --blueprint triage
Value for 'anthropic_api_key': [hidden]
Stored secret 'anthropic_api_key'
Added blueprint 'triage'
{"level":"high","rationale":"These tickets describe duplicate charging and a payment/order status mismatch, both of which are financial discrepancies that directly impact customers and require prompt investigation to prevent overcharging, refund issues, or fraudulent order fulfillment. While not system-wide outages, they represent real monetary harm and trust issues that should be addressed urgently but are not catastrophic/critical in scope."}