Armin Ronacher: What is Codemode

Wait 5 sec.

More than a year ago I wrote a few posts here that recommended people not toload custom tools into their context (orMCP servers) but to justuse more scripts. Most importantly I wrote that Code Is All YouNeed and I wrote about that MCP needscode. With Pi 1.0 we now added MCP support via Codemodewhich in some ways is a long time coming, but then also maybe somewhatsurprising to some. So I want to share some updated thoughts on this blog onwhat this all means.What Are ToolsWhen a harness like Pi provides tools for an LLM to call, it does so bysupplying some tool definitions which then translate into some token structureon the server side. Whether a model is encouraged to call a tool is the resultof the reinforcement learning process. Something I wrote aboutbefore if you want to learn more.One of the reasons we strongly lean towards CLI and bash is because it allowseasy composition of calls, and because the model also learns how the file systemworks when it’s trained. So when it invokes a tool like echo foo > /tmp/test.txt the model also learns that after that tool call, there is now afile called test.txt in /tmp.However bash has one fundamental limitation which is that it can only composeprograms that run. And there are some things, which are not programs, butnative tools to the LLM and they sort of have to be.The most obvious example here is read or view_image. If a multimodal modelneeds to read an image, it cannot use cat for that because the harness needsto inject the actual image payload into the protocol of the LLM.Another quite vivid example are sub agents. In order to spawn and orchestratesub agents, it’s tricky to avoid tools that are provided by the harness. Whilein theory the agent could provide a CLI tool that talks to the outer harnessvia environment variables and Unix sockets, it’s a rather crude process. Ithowever has another issue, and that is where the code runs.Brains vs HandsTo better understand that, it’s important to think a bit more about where all thebits and pieces run. There really usually are two different systems involved.The first is the brain, the harness: it runs on one machine. It’s trusted. Thesecond is often the same machine, but it’s really where the tools areexecuting: the hands. In Pi we now call this the execution environment, but youcan think of it as the target of all the operations.Crucially what is important for us, is that there is a dividing line between theharness brain and the target environment that runs bash and executes the tools.And splitting this in half has some really important consequences. For a startit means that they are running on different file systems and they have differentlevels of trust. If you for instance use a sandboxing solution likeGondolin your bash stuff will besandboxed just fine, but the harness itself will not be.Orchestrating The HarnessWhich brings us to what Codemode really does: it’s a way for the LLM to expressand orchestrate complex operations on the harness side, but not the executionenvironment side. Codemode runs in the harness, in its own sandbox. In case ofPi it’s running in QuickJS within a WASM runtime with intentional limitations:no network, no file system, no timers, limited RAM. The only way is to callmore tools. You could also imagine that Codemode could run Scheme or some otherlanguage as well.If you are not familiar with Codemode, it’s basically just a way to issuetool calls from within some language, in our case JavaScript. That allows youto compose those calls without necessarily going through the LLM’s context.Credit for naming goes to our friends at Cloudflare who coinedit.For instance if you issue a bash call as a regular tool call in the LLM, thenwe only throw the trailing 2000 lines into the context and if the agent wantsmore, it needs to look at the overflow file itself. If however the agent issuesthat invocation via Codemode, then the Codemode side gets larger outputssent structurally.Most importantly, because Codemode is JavaScript the agent can expressconcurrent operations and basic workflows. A common way in which you see agentsnow use this, is to first probe at 5-10 items from some tool response to seewhat it looks like, and to then write a Codemode script that processes the nextn items.Codemode also allows you to throw state into the transcript! That means thatone Codemode invocation can stash away data, that the next call in the sessioncan load again. And remember: this is on the harness host, not the sandbox.In case of Pi, Codemode also allows you to issue calls that naturally do notmake any sense in Pi’s traditional interface. For instance if you want togenerate images with an image model or you want to classify some text with aone shot classifier model, those Pi APIs are exposed via Codemode, but not viaregular tools where they would just waste context.What It Looks LikeSo now that we talked a bunch about it, it’s probably worth being a bit moreexplicit about it. Let’s walk ourselves through some invocations of Codemodeof recent Pi sessions of mine. Note that none of this code is human written.It’s from real sessions of Pi, just re-indented for your viewing pleasure. Theagent starts using Codemode automatically either because it’s a task where themodel already naturally picks up that tool, or because a user asked it to.Note that Codemode is by default only enabled in Pi when MCP is enabled, but youcan turn it on with "defaultTools": ["+codemode"] in the settings. Just askPi to enable it for you.Generating ImagesLet’s start simple with image generation. Image generation is a feature that Pisupports in the AI SDK core, but it’s not a tool that the agent can use. In thepast the only way to use image models has been to write a bespoke extension orto have the agent run node itself and use the internal image APIs. Howeverbecause we expose quite a few of the internal model APIs within Codemode, itmeans that the agent can use it:const [painter] = await models.getAvailableOfType("image");const result = await models.generateImages(painter, { input: [{ type: "text", text: "A cute little puppy sitting on a grassy " + "lawn, soft natural light, photorealistic" }],});if (result.stopReason !== "stop") return result.errorMessage;for (const block of result.output) { if (block.type === "image") image(block); else text(block.text);}Note that the call to image() sends the image back as image content to theLLM. On the harness side it feeds it directly into both the agent, as well asonto disk as a temporary artifact in case the agent wants to be able to passthat image back to bash.Classifying ThingsSimilar things apply to classifier models such as Jev.They also do not fit well into the workflows of an agent through the typicaltools. But rather than making a bespoke tool available, Codemode just allowsthe agent to reach into the AI SDK and invoke those directly. Here you can seehow Jev is used to mass process GitHub issues for a quick sentiment analysis:const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");const r = await tools.bash({ command: "gh issue list --state open --limit 100 " + "--json number,title,body,comments",});const issues = JSON.parse(r.output);const results = await Promise.all(issues.map(async (issue) => { const res = await models.classify(jev, { state: { title: issue.title, body: (issue.body || "").slice(0, 4000), comments: issue.comments.slice(-5).map(c => c.body.slice(0, 800)), }, questions: { sentiment: { type: "choice", instructions: "What is the overall sentiment of the author towards pi?", criteria: { positive: "Appreciative, happy, constructive praise", neutral: "Matter-of-fact report or request without emotion", negative: "Frustrated, annoyed, upset, or angry", }, }, frustration: { type: "score", instructions: "How frustrated is the reporter?", criteria: ["not at all", "mildly", "clearly frustrated", "very angry"], }, kind: { type: "choice", instructions: "What kind of issue is this?", criteria: { bug: "Bug report or regression", feature: "Feature request or enhancement", question: "Question or support request", other: "Docs, discussion, meta, spam", }, }, }, }); if (res.stopReason !== "stop") { return { n: issue.number, title: issue.title, error: res.errorMessage }; } return { n: issue.number, title: issue.title, ...res.answers };}));store("sentiment_results", results);return results .filter(r => !r.error) .sort((a, b) => b.frustration.score - a.frustration.score) .slice(0, 12) .map(r => `#${r.n} ${r.frustration.score.toFixed(2)} [${r.kind.choice}] ${r.title}`);Note how in that above example we also call store() which dumps the result ofthat execution into the session transcript. A future invocation of Codemode canthus read back that result if it wants to.The Promise.all here is fine, because Pi limits the total number of concurrenttool executions itself to four and maintains a queue for the rest.A more adventurous example is to use Jev to drive a game engine for debuggingpurposes:Codemode with Jev for Game DebuggingHere it knows about my tankctl command and it built itself quickly a minimalharness around it to drive a game loop to assist a user with debugging aproblem. Note how it built a 30 step loop in which each step goes back to boththe game engine to get a text dump of what’s going on, and then to Jev todetermine what to do next:const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");const tank = async (cmd) => (await tools.bash({ command: `tools/tankctl "${cmd}"` })).output;await tank("start --map assets/maps/night_arena.map");const questions = { action: { type: "choice", instructions: "You control the tank '@' in a top-down tank game. " + "Choose the best next action.", criteria: { attack: "an enemy has line of sight to you and you can fire at it", approach: "no enemy has line of sight; drive toward the nearest enemy", dodge: "an enemy shot is heading at you and will hit soon", powerup: "a powerup is close and no enemy threatens you", }, },};function commandFor(choice, st) { const p = st.player; const enemy = st.enemies.filter(e => !e.dead) .sort((a, b) => (b.los - a.los) || (a.dist - b.dist))[0]; if (choice === "attack" && enemy) { return `fire_at tank ${enemy.id}; frames 30 until clear,damage,kill`; } if (choice === "dodge") { // move perpendicular to the closest incoming shot const s = st.projectiles.filter(s => !s.yours) .sort((a, b) => a.eta - b.eta)[0]; const dir = s && Math.abs(s.vel[0]) > Math.abs(s.vel[1]) ? (p.pos[1] > s.pos[1] ? "+down" : "+up") : (p.pos[0] > (s ? s.pos[0] : 0) ? "+right" : "+left"); return `input ${dir}; frames 20 until damage; input stop`; } const powerup = st.powerups.filter(u => u.available) .sort((a, b) => a.dist - b.dist)[0]; if (choice === "powerup" && powerup) { return `goto ${powerup.pos[0]} ${powerup.pos[1]} 180`; } return enemy ? `goto ${enemy.pos[0]} ${enemy.pos[1]} 90` : null;}const log = [];for (let step = 0; step !s.yours && s.miss_dist ${await tank(cmd)}`);}return log.join("\n");Calling MCP ServersLastly, Codemode obviously is great for calling MCP servers. And because wedo not actually expose any of the MCP tools to the LLM, the agent first usesprovided APIs to issue a tool search within Codemode to discover what it mightbe able to do with the connected servers. This form of progressive discoverymakes the whole MCP business work well enough for a lot of use cases today.Here for instance you can see the agent reach for the Sentry MCP straight away,even without discovering the tools, presumably because it has learned during theRL process already about what the Sentry MCP looks like. But it learns fromwhat we inject into the system prompt, that the Sentry server is available tobegin with. It’s not completely guessing here.const orgs = await tools.mcp__sentry__find_organizations({});const { organizations } = orgs.structuredContent;const results = await Promise.allSettled(organizations.map(org => tools.mcp__sentry__find_projects({ organizationSlug: org.slug, regionUrl: org.regionUrl, })));return organizations.map((org, i) => { const r = results[i]; if (r.status !== "fulfilled") return { org: org.slug, error: String(r.reason) }; if (r.value.isError) return { org: org.slug, error: r.value.content }; return { org: org.slug, projects: r.value.structuredContent.projects.map(p => p.slug), };});Modern MCP Is A FightI really don’t want to talk too much about MCP here, but MCP is in fact aprotocol that greatly benefits from Codemode. The problem in parts is that MCPin practice often targets harnesses that do not (yet?) use Codemode. But thetide is shifting. In the meantime, a temporary crutch has been to do whatCloudflare did, and do Codemode within the MCP server. But now we have Codemodein Codemode which is pretty bad. It means double JSON escaping, easy forsmaller models to get confused by and the inner code cannot call the outertools. So if you for instance use the Cloudflare MCP servers in Pi, the agentneeds to write JavaScript and funnel it through more JavaScript. This is reallynot optimal, but it’s also understandable that this is happening:const accRes = await tools.mcp__cloudflare__execute({ code: `async () => { const r = await cloudflare.request({ method: "GET", path: "/accounts" }); return r.result.map(a => ({ id: a.id, name: a.name })); }`,});const accounts = JSON.parse(accRes.content.map(c => c.text).join(""));const out = [];for (const account of accounts) { const r = await tools.mcp__cloudflare__execute({ account_id: account.id, code: `async () => { const r = await cloudflare.request({ method: "GET", path: \`/accounts/\${accountId}/workers/scripts\`, }); return r.result.map(s => ({ id: s.id, modified: s.modified_on })); }`, }); out.push({ account: account.name, workers: r.content.map(c => c.text).join("") });}return out;MCP DesiresSo to end things off: how well does Codemode work with MCP today? Well … notamazingly well. That’s because MCP servers are not really targeting harnessesthat use Codemode yet (though at this point I think most harnesses support it).For this to work well some recommendations:Structured content: Codemode wants calls to return some nicelyformatted JSON. So that needs to come back from the server, and many don’tdo that yet. The outputSchema system in MCP is great for that.Consistent results: an interesting failure case is when an MCP serverdoes not return consistent data. For instance because it tries to tokenoptimize things depending on how many items are in the result set. This cancause an initial probe with 5 items to succeed, but then fail when the serverreturns the maximum batch size.Large binary data: today MCP does not yet support large binary dataso quite a few use cases that are really interesting do not work well at allyet. You end up with all kinds of weird workarounds such as pre-signed URLsto allow file uploads then to happen through non MCP channels.Composable tool search: the MCP server might know better than the MCPclient which tool is appropriate for a task. But there is no good mechanismtoday that allows a harness to fan out tool searches across multiple MCPservers. It’s all emergent behavior and it does not scale well to multipleactive servers.Future of CodemodeSo where does this leave us? Is this a reversal of what I wrote a year agowhere I encouraged CLIs? I don’t think so. In fact, the MCP ecosystem from myperspective picked up on exactly what we pointed out a year ago works: code.But Codemode goes beyond MCP in that it can act as a capable mechanism withinthe harness to express more freedom for the agent.There are however also some things that we still need to figure out. For one,durability with Codemode is trickier. We might have to adopt some ideas fromdurable workflow engines here to snapshot invocations. Or maybe, something likeStarlark is a better composition language than JavaScript given itsdeterministic nature.Images, binary data and just the inability of this pattern to work with smallermodels is also something that needs to be fleshed out. So it’s for sure not aperfect solution yet, but it’s quite a useful pattern that I expect us toleverage more.