Your application, or the agent framework it uses, is the harness, the code that runs the agent’s loop. Submilli replaces none of it. The harness keeps the model and the loop, and gains one tool that takes a program the model wrote and runs it on the server under your Blueprint.
The tutorials in this part build the same research agent on five harnesses, one per page. This page sets up what they share, the server and the Blueprint, then says what every harness decides when it connects. Take the page for yours when you are done:
| Harness | Language | Connects over |
|---|---|---|
| Mastra | TypeScript | MCP |
| LangChain deepagents | Python | MCP |
| OpenAI Agents SDK | Python | MCP |
| Claude Agent SDK | TypeScript | MCP |
| Vercel AI SDK | TypeScript | The HTTP API |
The agent answers questions by searching the web and reading pages, and
keeps notes so that the next conversation can start from what the last
one learned. It runs on behalf of whoever is signed in to your
application. The examples stand in for that person with one user id,
u_ada.
1. Save the Blueprint
Section titled “1. Save the Blueprint”Make a directory for the tutorials, harnesses, and save this in it:
kind: blueprintname: research
# Bound once per connection by the application, never by the model.variables: userId: required: true
packages:- '@submilli/jina'
secrets: JINA_API_KEY: store: jina_api_key ANTHROPIC_API_KEY: store: anthropic_api_key # GOOGLE_API_KEY: # store: google_api_key # OPENAI_API_KEY: # store: openai_api_key
# A model the programs may call themselves, to summarize a page or sort# results without bringing the text back into the conversation. One# provider is enough; uncomment yours.llm: providers: anthropic: type: anthropic api_key: ${secrets.ANTHROPIC_API_KEY} # google: # type: google # api_key: ${secrets.GOOGLE_API_KEY} # openai: # type: openai # api_key: ${secrets.OPENAI_API_KEY} models: claude-haiku-4-5: provider: anthropic description: Fast and cheap; use for summarizing pages and ranking results. # gemini-3.8-flash: # provider: google # description: Fast and cheap; use for summarizing pages and ranking results. # gpt-4o-mini: # provider: openai # description: Fast and cheap; use for summarizing pages and ranking results.
# Notes outlive the conversation. The `notes` volume keeps one directory# per user; a session sees only its own user's, at /notes, and starts# there. Everything else is scratch space that ends with the session.vfs: mode: per_session cwd: /notes mounts: /notes: mode: named volume: notes subPath: ${vars.userId}
default: deny
permissions: # What generated code may do. main: - capability: jina.ai/search action: allow - capability: jina.ai/read action: allow - capability: llm.call action: allow - capability: fs.read action: allow - capability: fs.write action: allow - capability: fs.list action: allow - capability: fs.stat action: allow - capability: fs.mkdir action: allow
# What the package itself may do, as `add-package` wrote it: reach # Jina, read its key, and save a download where the program asks. '@submilli/jina': - capability: fs.write action: allow - capability: http.download filter: host == "r.jina.ai" action: allow - capability: http.download filter: host == "s.jina.ai" action: allow - capability: http.post filter: host == "r.jina.ai" and path == "/" action: allow - capability: http.post filter: host == "s.jina.ai" and path == "/" action: allow - capability: secrets.get filter: name == "JINA_API_KEY" action: allowIn short, the agent may search and read through the curated
@submilli/jina Package, call one model, and keep notes at /notes.
That path is the same for each user, but what is behind it isn’t. The
notes volume holds a directory per user, and subPath mounts the one
the session’s userId names, so neither the program nor the Package
can reach another user’s notes, and no rule has to name a user. The file names Anthropic as the
provider and keeps Google and OpenAI entries commented out. Uncomment
yours. For what each block does, refer to Keep files and
state, Allow model
calls, and Mount a shared
volume.
2. Save a test program
Section titled “2. Save a test program”Beside it, save a program to prove the setup with before any model is
involved. It lists the notes already there, then writes one. A relative
path is under /notes, where every program starts:
import fs from "submilli:fs";
function main(): string { const before: string[] = []; for (const entry of fs.list(".", false)) before.push(entry.name); fs.writeText("check.md", "# Written by a check\n"); return `notes before: ${before.join(", ") || "none"}`;}3. Save the agent’s brief
Section titled “3. Save the agent’s brief”The harness supplies no system prompt for Submilli. The instructions that teach a model the language arrive as the execute tool’s description, with this Blueprint’s Packages and rules filled in. The harness does supply the agent’s brief, and each tutorial’s agent reads it from this file:
You are a research assistant. You answer questions by searching the weband reading pages, and you keep a notebook so the next conversation canstart from what this one learned.
## Work in programs
Do the work by writing and running TypeScript on Submilli. Read `docs`before using an unfamiliar package or API: the runtime is not Node.js,has no shell, and has no npm packages. Prefer one coherent program forrelated reads, filtering, and summaries, and return the evidence theanswer needs, not whole pages. Keep predictable follow-up steps insidethe same program: a URL or an id one call returns is used by the nextcall in code, not in another turn. When a program fails, read thediagnostic and repair it. A permission denial is final; do not look foranother route to the same effect.
## Resources
- `@submilli/jina`: web search, and reading a page as clean text.- `submilli:llm`: a model you may call from a program, to summarize a long page or rank results without bringing the text back here.- `submilli:fs`: your notebook, the directory /notes, where every program starts. Your first program lists it, reads what is there, and gets today's date from `Temporal.Now.plainDateISO()`, before any search. When you are done, write what you learned, with its sources, updating an existing note rather than replacing what it got right. Use no other file tool for it.
## Answer
Prefer the newest source and check its date against today's before youcall something the latest. Lead with the answer and cite the pages youused. Distinguish what theevidence establishes from what you infer and what remains unknown. Neverclaim you read or saved something unless a program's result shows it.Its shape works. It says what the agent is for, how to work in programs rather than one call at a time, what it may reach, and how to answer. Edit it for your agent.
4. Install the Package
Section titled “4. Install the Package”submilli install submilli/submilli-runtime @submilli/jinafetched github.com/submilli/submilli-runtime at 528fd656c38ainstalled @submilli/jina v0.1.0 -> ~/.submilli/packages/@submilli/jinaA server on the same machine reads this store.
5. Write the server’s config file
Section titled “5. Write the server’s config file”In harnesses:
secret_store: key_file: store.keyvolumes: notes: kind: managed-local size_limit: unlimited6. Start the server, store the keys, register the Blueprint
Section titled “6. Start the server, store the keys, register the Blueprint”head -c 32 /dev/urandom | base64 > store.keyexport SUBMILLI_SERVER_TOKEN=$(openssl rand -hex 32)submilli-server --config server.yaml &
submilli server secret put jina_api_keysubmilli server secret put anthropic_api_key# submilli server secret put google_api_key# submilli server secret put openai_api_keysubmilli server blueprint apply blueprint.yamlValue for 'jina_api_key': [hidden]Stored secret 'jina_api_key'Value for 'anthropic_api_key': [hidden]Stored secret 'anthropic_api_key'Added blueprint 'research'Jina issues a key at jina.ai. The second key is your
model provider’s. The checks make no request to either, so any value
will do for them, but the conversations need real ones. Keep
SUBMILLI_SERVER_TOKEN exported and the server up. The agents send the
token with every request, and the submilli server commands read it
too.
7. Prove it
Section titled “7. Prove it”Run the test program the way an application would, as u_ada, then as
another user, then as u_ada again:
submilli server run-code note.ts --blueprint research --var userId=u_adasubmilli server run-code note.ts --blueprint research --var userId=u_gracesubmilli server run-code note.ts --blueprint research --var userId=u_adanotes before: nonenotes before: nonenotes before: check.mdNotice the second run. u_ada had already written check.md, and
u_grace, running the same program at the same path, didn’t see it.
The binding, not the program, decides whose notes /notes holds,
whichever harness opens the session.
Now the model. This program reads a page through the Package and asks
the model to sum it up. llm.models() lists the models the Blueprint
allows, so it works whichever provider you uncommented:
import jina from "@submilli/jina";import llm from "submilli:llm";
function main(): string { const model = llm.models()[0].name; const page = jina.read("https://blog.rust-lang.org/2026/10/01/Rust-1.99.0/"); return llm.call<string>(model, `In two sentences, what is new in this release?\n\n${page.slice(0, 20000)}`);}submilli server run-code summarize.ts --blueprint research --var userId=u_adaRust 1.99.0 stabilizes defining C-ABI variadic functions with "C" and "C-unwind" ABIs, allowing variadic functions to be written in Rust itself, as well as stabilizing functions for retrieving size and alignment information from raw pointers to both sized and unsized types. The release also includes updated documentation recommending against unsafe round-trip unleaking patterns after `Box::leak`, along with numerous other stabilized APIs and standard library improvements.The page never reached the conversation. The program fetched it, the model read it, and two sentences came back.
What every harness does
Section titled “What every harness does”Whatever the harness, three things are its decision, and the model has no part in them:
- The address names the Blueprint. The MCP endpoint is
http://127.0.0.1:8128/mcp/research. Every program the harness sends there runs under that Blueprint, and no tool takes a Blueprint as an argument. - A header binds the variables.
submilli-variables: userId=u_ada. The server checks the values against the Blueprint before it accepts the connection and refuses one that leaves out a required variable. Take the value from what your application knows (the signed-in user), never from the conversation. - One connection is one session. The variables, the session’s state,
and its files outside
/noteslast as long as the connection. Open one per user and close it when the conversation ends.
A secret that belongs to the user, such as their own token for a
service, is declared with a harness source and sent the same way, in
a submilli-secrets header. Start a
Blueprint
shows the declaration, and the HTTP API
page shows the request. Connected, the model gets the tools Your
application describes, with
the execute tool’s description already carrying this Blueprint’s
Packages and rules. The harness adds the brief from step 3 and nothing
else.
Now take the tutorial for your harness: Mastra, deepagents, OpenAI Agents, Claude Agent SDK, or the HTTP API.