Blog · 23 September 2026
Captain Code is open source
One terminal for every coding agent you already pay for. Captain Code picks a model for each task, carries the work on when one hits its limit, and puts open models to work when they are good enough.
On 22 September we released Captain Code under the MIT license. It is the controller we use every day at Lemma to run Claude Code, Codex CLI, Cursor and a bench of hosted and open-weight models from a single conversation. It runs on your machine, with your own logins and keys, and sends no telemetry.
You type the task. Captain Code does the rest. It works out what kind of task it is, sends it to the agent that is good enough at the lowest cost, and when that agent hits a usage limit, hands the work, context included, to the next one.
At a glance
- Open source, MIT. Code, documentation and release notes are on GitHub.
- Local. The controller, run history and settings stay on your machine. There is no hosted service and no telemetry.
- Your subscriptions. It drives the Claude Code, Codex CLI and Cursor agents you already pay for, with your own logins.
- Open models within reach. Grok, GLM, DeepSeek, Kimi, Qwen, Gemini, MiniMax and free tiers run through opencode, including open-weight models on OpenRouter, NVIDIA NIM and Hugging Face.
- Installed by your agent. Paste one prompt into the coding agent you use today. It sets Captain Code up and stops when
captain doctorpasses.
Why we built it
Three things kept getting in our way:
- We no longer knew which model to use for what. New models ship every week, and every coding agent ties a session to one of them.
- We were under-using what we already pay for. One subscription would hit its limit mid-task, with the plan and the half-finished work stuck inside it, while the others sat idle in other windows.
- We wanted to hand more work to cheaper, capable open models. A rename should not cost as much as an architecture migration just because both went to the same model.
Captain Code sits in front of your models instead of inside one of them.
How it works
- Triage in a fraction of a second. Each prompt is classified by kind of task and domain. With a TypeSafe key, Captain Code asks Jev, TypeSafe's System One decision model, which answers typed questions with a calibrated probability in about 0.1 to 0.8 seconds, for a fraction of a cent. Its answer is used when its confidence clears 0.6; otherwise built-in rules and a free model decide. Without a key, nothing breaks; triage is just slower.
- Route on value. Every leg, meaning one model running through one tool, has a quality score, a price and a measured latency. Captain Code picks the cheapest leg that clears the quality bar for the task: routine work leans cheap, hard work leans toward quality. A director model plans harder work and reviews combined results, and it never assigns work to itself.
- Hand off at the wall. When a leg hits a usage limit, Captain Code reads when the limit resets instead of guessing, keeps useful partial work, and hands the task with its context to the next best leg. Long conversations are trimmed predictably before they are summarised, and your original instruction stays at the top, so the next model does not start cold.
- Record everything, locally. Every run records the leg, the kind of task, duration, tokens, estimated cost, outcome and the full worker log, in plain files you own. Routing draws on that record, and so can you:
captain runslists every run andcaptain whyexplains the last routing decision.
Type a task, or take the helm
No model name or special syntax is required. A bare prompt is enough:
Fix the broken pagination and add a regression test.
When you want more control, a few words go in front of the prompt. Type /captain in the terminal for the full guide.
| What you want | What you type |
|---|---|
| Let Captain Code choose | Your prompt, without a prefix |
| Quality over cost | /quality Review this migration for data-loss risks |
| A fast answer, or the lowest plausible cost | /speed Explain this test failure · /save Tighten this README |
| Claude at maximum effort | /frontier Challenge this architecture |
| A particular agent | /cursor Fix the pagination |
| Any leg at its strongest settings | /frontier /grok Draft the migration plan |
| Open-weight models only | /oss Refactor this module |
| Only serving stacks measured as reproducible | /deterministic Generate the migration |
| Several workers, organized by the director | /team /quality Review this change from independent angles |
| A workflow described in plain English | /wf Have grok implement this, then cursor review it |
| A note for the worker already running | /btw keep the public API stable |
| Stop, but keep the partial work | /interrupt |
/quality leaves the choice of worker or team to the director, with cost set aside. /frontier on its own sends the task to Claude at maximum effort; in front of a named leg or a workflow it becomes a modifier that runs each leg at its most capable settings. Both express how to spend effort. Tests and review still decide whether the result is good.
/oss narrows the pool to open-weight models. /deterministic narrows it to serving stacks that the Agentic Determinism Index scored byte-exact on its latest run, and pins the run to that provider at temperature 0. Green is a snapshot of a deployment, not a certification, and only stacks ADI measures can be green: the local CLI agents are not probed, so /deterministic never selects them. Both words compose with the others, in either order.
Prompts you type while a turn is running are queued. Each then runs as its own turn, in the order you typed it, with its own routing.
Workflows: describe the sequence, then run it
Suppose a change needs an implementation followed by two independent reviews:
/wf Have grok implement pagination, then have cursor review
correctness and codex-cli check edge cases in parallel, without editing.
Captain Code turns that request into a plan showing the stages, workers and estimated duration. You inspect it, then send the /run wf_… command it gives you. The same workflow also fits on one line, and runs immediately:
/grok implement pagination > /cursor review correctness without editing + /codex-cli check edge cases without editing
The > operator means next; + means together. Each stage receives the previous stage's output and the original request, and the director returns one combined answer with findings attributed to their workers. A stage can carry a gate: test command that must pass before execution proceeds.
Use /team when you want the director to organize the work, and a workflow when you want to set the order yourself.
New in the launch releases
Four releases, v0.2.0 to v0.2.3, shipped between 20 and 22 September. The last one packages the first two pull requests merged in the public repository, both from @rafa-js.
- Setup your agent can follow. Paste one prompt into Claude Code, Codex, Cursor or any other coding agent. It follows
AGENT_SETUP.md, asks before installing anything, hands every login and key back to you, and finishes whencaptain doctorpasses. - Legs that cannot run say why. One readiness verdict is now shared by
captain doctor, the director, the router and the fallback ladder, so a leg with a missing login or key is skipped up front and names its cause. On one task, about 92 seconds of opaque failures became about 19. - **
/frontierreaches every leg's ceiling.** In front of a leg or a workflow, it runs each stage at that leg's strongest settings, including Grok 4.7 where it is available. - A queue is a sequence. Prompts typed during a turn run in order, each routed on its own.
- Bring your own decision model. An open, local decision backend can run beside Jev. It records its answers and decides nothing until you promote it against its own measured accuracy.
- Shield, the action gate and supervision. Recognisable secrets in tool output and provider requests are masked before a model or remote API sees them. With Jev configured, Captain Code also screens what a worker is about to do and whether it is stuck; both only record by default, and
CAPTAIN_ACTION_GATE=enforceblocks risky actions. - A vetted skills shelf.
captain skills syncpins a skills source at a named commit, vets it and locks its hashes. Workers are then stocked with the skills that fit the task, andcaptain skills reportshows which ones proved useful.
The full notes are on the releases page.
Privacy and security, plainly
- Local by design. The controller, run history and settings live on your machine. Nothing is sent anywhere except the prompts you route to a provider.
- Hosted models still see what you send them. Selected context goes to whichever providers you route to, including the director. Local orchestration does not make remote inference local.
- Workers run with approvals off by default. A background worker that stops to ask a question would wait forever, so vendor approval prompts are turned off. Read SECURITY.md before pointing Captain Code at anything you do not own; it explains how to keep approvals on or sandbox a worker.
- Your models, your rules. Choose your director with
CAPTAIN_DIRECTORor/captain <leg>, and the legs routing may use withCAPTAIN_LEGS. That list narrows automatic routing, but explicit commands can override it, so it is not a data boundary. - Your credentials, your work. Captain Code runs vendor CLIs with your logins, on your machine, for your work. It does not pool, share or resell subscriptions.
What it changed for us
Our starting point was an almost entirely Claude-based workflow. Here is where we stand today, after 54 days of running our work through Captain Code.
Open models doing real work
- 1,823 worker runs since 1 August, spread across 16 legs.
- 308 of them ran on open-weight models: about one in six (17%). Each is a task that spent neither a subscription's usage window nor a frontier model's price.
- Seven open-weight lanes did the work: GLM, DeepSeek, Kimi, Qwen, MiniMax, StepFun and the free tier.
One in six is a floor, not a ceiling, and it is the number we most want to raise. The next section shows how.
Many projects at once
The clearest change is in our commit history, not in any one repository. Deduplicated across clones and counting only our own commits:
| May to July | Since 1 August | |
|---|---|---|
| Repositories with commits | 7 | 22, 20 of them new |
| Commits per week | 118 | 167 |
| Repositories touched per working day, on average | 1 | 4 |
| Working days touching three or more repositories | 3 of 56 | 28 of 41 |
The volume went up, but the shape changed more. Before, a working day meant one project. Now most days move several forward in parallel, up to nine in a single day. The work spans product code, two public websites, a research index, decks and strategy documents, and shared kits for sites, decks and demos. About a quarter of the commits are documents, decks and site content rather than code.
Three things made that possible:
- Limits stop stalling the day. When one subscription hits its limit, the task moves on with its context instead of waiting for the reset.
- The right model for each piece. Routine edits go to cheap and open models, and the frontier models are kept for the hard parts, so running several projects at once stays affordable.
- One way of working everywhere. The same routing, checks and secret masking apply in every repository, so switching projects takes almost no setup.
How we counted: worker runs are counted by the leg that ran them, including team and workflow workers, failures and explicit choices. Director planning and review calls are excluded. The free tier's roster rotates, and we count it as open-weight. Captain Code does not record which repository a run worked in, so the commit table describes the period, not a controlled measurement. Percentages are rounded.
Scaling onto open models
We want far more of our work on open-weight models, and Captain Code is built for it. New models appear on OpenRouter every week, and one command makes a new release routable:
captain legs add <id> openrouter/<slug> --aa <slug>
That pulls the current price and context length from OpenRouter's public catalog into the registry, with no rebuild. The --aa slug links the leg to its Artificial Analysis benchmarks, which captain priors sync fetches with a free AA key to keep the routing priors fresh. Keys for NVIDIA NIM and Hugging Face open more open-weight legs.
From there, value routing weighs the new leg against the rest on every prompt, /oss pins a task to open models, and the run records tell you whether the cheaper legs deliver on the work you actually do.
Try it
Paste this into Claude Code, Codex, Cursor or any other coding agent:
Set up Captain Code on this machine by following
https://github.com/lemma-ventures/captaincode/blob/main/AGENT_SETUP.md
step by step. Ask me before installing anything, and hand me any login or
API key instead of doing it yourself.
Or install it by hand:
git clone https://github.com/lemma-ventures/captaincode && cd captaincode
go build -o ~/.local/bin/captain ./cmd/captaincode
(cd plugin/captain-ui && bun install)
captain init
captain doctor
./captaincode.sh
captain doctor is the real quickstart. It lists every agent as ready, blocked or skipped, and each blocked line includes the command that fixes it. You do not need every agent: use whatever you have.
We use Captain Code every day for our own work. We publish it because it is useful, not as a product, so there is no support contract. Issues and pull requests are welcome.
Try it, then tell us what you routed first
Star the repository, open an issue or send a pull request. If you run several coding agents, we would like to hear which workflow you want to hand to Captain Code first.