Eight days of thunc
thunc went from an empty repository on 3 October 2026 to its twelfth release on 10 October. This is the whole story so far: what was built each day, what the benchmarks showed, what went wrong, and what changed because of it.
The releases
- 0.1.0 First release: typed functions on Claude, Claude Code and Codex
- 0.1.1 The OpenAI backend and local models
- 0.1.2
cache=True,system=and a parser hardened over twelve rounds - 0.1.3 The Jev backend
- 0.2.0 Agents
- 0.2.1 Durable agents with Temporal
- 0.2.2 Native tool calls on Claude Code, after a benchmark
- 0.2.3 Native tool calls everywhere agents run; the last of 0.2
- thunc-watch 0.1.0 A live dashboard in the terminal
- 0.3.0
thunc watch, and thunc write as an experiment - 0.3.1 The changes that missed 0.3.0's package
- 0.3.2 A stability pass
Each release has its full notes in the changelog. This post is the story around them.
A function that is a prompt
The idea behind thunc is that calling a model from code shouldn't need a prompt template, a JSON schema and a page of parsing code. You write the function you want. The docstring is the prompt, the parameters are the inputs, and the return type is the contract: thunc checks the reply against it, sends a wrong answer back to the model to try again, and raises a ThuncError if it still doesn't fit.
import thunc
@thunc.function
def urgency(ticket: str) -> int:
"""Rate how urgent this ticket is,
from 1 (can wait) to 5 (customer is blocked)."""
...
urgency("I was charged twice!") # 4, a checked intThe name is think + function. Two constraints were set on the first day and haven't moved since: thunc has no dependencies beyond the standard library, and it runs on whatever you already have. 0.1.0 shipped with three backends, the Claude API and the Claude Code and Codex command-line tools, so anyone logged in to one of those could try it without an API key. It also had thunc.call for prompts built in code, thunc.map to run many calls at once, ensure= for your own checks, and JSONL tracing.
The first commit with code landed at 13:33. 0.1.0 was on PyPI at 13:49, through PyPI's Trusted Publishing: creating a GitHub release builds and uploads the package, with no tokens to manage. By the end of the afternoon 0.1.1 added an OpenAI backend on the Responses API, and with it local models through OPENAI_BASE_URL, tested with LM Studio and gpt-oss-20b. The same afternoon added a CONTRIBUTING guide, GitHub Discussions and this website.
Twelve rounds of trying to break the parser
Models don't always follow the contract. They wrap JSON in a code fence, open with a <think> block, answer {"rating": 4} when asked for an int, or reply NaN. The parser has to read the harmless near-misses, send the ambiguous ones back, and never return a wrong value as if it were right.
To find out where it failed, independent agents attacked it in rounds. Each round wrote 25 to 37 new tests against the contract, offline with scripted backends, on Python 3.10 and 3.11. Every finding was fixed with a regression test or written down as a known issue. Rounds 11 and 12 were regression hunts that compared the branch with main over more than 150,000 (reply, type) pairs. Against a fixed set of 102 realistic bad replies:
| Outcome | Before | After |
|---|---|---|
| Wrong value returned, no error | 18 | 5 |
| Crash that skipped the retry loop | 2 | 0 |
| Near-miss read correctly | 7 | 24 |
The last rounds taught the most useful lesson. They mostly found bugs in the branch's own cleverest rules: an exactness check for whole-number floats, unwrapping under a field name, an elaborate rule for a code fence's opening line. Instead of patching them again, the final commit removed them. The test suite went from 61 tests to 205, and only ThuncError can now escape a call because of a bad reply. The same release added cache=True, which asks the model once per input and keeps the answer on disk, and system= for your own system prompt.
The first release mistake. The 0.1.2 GitHub release was published about a minute before the version-bump pull request merged. The tag pointed at the old commit, the build produced 0.1.1 again, and PyPI rejected it: a version number can never be uploaded twice. The fix the next morning was to turn the release back into a draft, move the tag to the bump commit and publish again. The rule since then: merge the bump, check the version on main, then publish.
Agents that hand back a typed answer
The morning's 0.1.3 added a backend for Jev, a small judgment model that answers yes/no, labels and ratings in about 0.3 seconds. It also made the Codex backend ignore your own Codex config and tools, so a call behaves the same on every machine.
Then the big one. A function answers from what you pass it; an agent can look around first. 0.2.0 introduced thunc.Agent: a typed task, written like a function, that lists, searches and reads files in a working directory and, where its permissions allow, edits them and runs commands, before it finishes with a checked value of its return type.
fixer = thunc.Agent(
"fixer",
workdir=".",
permissions=["write:src/**", "run:pytest"],
)
@fixer.task
def fix_failing_tests() -> bool:
"""Run the tests and fix any problems that they surface."""
...The pieces landed as seven stacked pull requests and merged within two minutes of each other:
- Permissions as rules with globs:
write:,read:,run:and!to deny. Commands run without a shell and with a minimal environment. - A record of every run:
agent.runreturns the value with the files changed, the commands run and each step, and anAgentErrorcarries the partial record when a run fails. follow=gives the agent yourAGENTS.mdorCLAUDE.mdas instructions, and agents keep notes between runs in a memory file.- Native tool calls on the Claude and OpenAI APIs, with the fixed part of the prompt cached. On the command-line backends, agents used a text protocol instead: one JSON action per reply.
- Your own functions as tools with
tools=, and system prompt presets for coding, code review and analysis.
The agent's default system prompt was a deliberate choice. Copying the prompts of Claude Code or Codex was ruled out early: their tools differ from thunc's, they assume a person at the keyboard, and they change between versions. thunc has a short prompt of its own, and live_tests/eval_prompts.py compares no prompt, the default and each preset on fixture repositories; 90 of 90 runs passed on the Claude API and Claude Code. CI started running the offline tests on Windows the same day.
Durable agents, the same afternoon
An agent that edits files and runs commands for minutes at a time shouldn't lose its work when a process dies. 0.2.1 added an optional runtime on Temporal, installed with thunc[temporal]. Each model turn and tool call is recorded in a workflow, so a run survives a worker restart and can be picked up from another process. File writes go through an intent and a receipt, so a retried step doesn't write twice, and a command that may or may not have run before a crash waits for you to say what happened with resolve() rather than guessing.
To make that possible, the agent loop's decisions moved into one engine shared by local and durable runs. Local thunc stays dependency-free; the Temporal SDK is only needed by those who ask for it. The evening went to the README (quickstart first, with demo GIFs) and a new look for this site, with the docs section you can read today.
Measuring thunc's own cost
Before tuning anything, thunc got two ways to measure itself. thunc run --profile app.py runs a program and reports where its time went: per function, calls, cache hits, retries and model time against thunc's own; per agent task, steps and time in each tool. And benchmarks/ is a suite that times thunc against stand-in models that answer at once, so the numbers are thunc's work, not a model's. It found three speedups:
| Change | Measured | Before → after |
|---|---|---|
| The API backends reuse their connections | 20 calls, with 60 ms of simulated connection setup | 1.47 s → 69 ms |
| Text-protocol agents can send several actions per reply | Mean run time over 3 tasks on Codex, 30/30 passing both ways | 37 / 27 / 27 s → 29 / 19 / 21 s |
| Codex answers return when the turn ends, not when the CLI exits | 8 pairs of real calls | faster in 8 of 8, ~0.67 s each |
The benchmark that changed the plan
The next question was harder: when a thunc agent fails or runs slowly, how much of that is thunc rather than the model? The tool-use benchmark (live_tests/bench_tooluse.py) answers it with eight small repositories, each pressing on one part of a harness: a value hidden two hops from its call site behind a stale build copy, a bug near line 1,900 of a 2,400-line file, a rename across 14 files, a module that is mostly backslashes and quotes, tab-indented near-duplicates, one real error in 4,000 lines of build output, tests that only pass from a subfolder, and a structured answer gathered from four files with two decoys. The same model ran every task through thunc and through Claude Code itself.
The answer was uncomfortable. On Claude Sonnet 5.5, thunc agents on the Claude Code backend passed 12 of 24 runs. Claude Code passed 24 of 24. The model was the same in every column; only the harness changed.
The report traced most of the gap to the text protocol. Current models are trained to call tools natively, and asked instead to write one JSON action as plain text, they slip back. They wrote a correct action and kept going, making up the tool's result and the next steps. One failed model call ended a whole run. search and list waded through what git ignores, and run had no working directory and kept only the end of long output, where the first error wasn't.
The fix was to stop fighting the model. On Claude Code, thunc now serves the agent's tools as an MCP server that one claude -p process per run calls natively; a small relay hands each call back to thunc, which carries it out with its own tools, permissions and records. Failed steps are retried. The tool gaps closed. And the Claude Code backend now loads none of your Claude Code settings, so files in an agent's working directory can't give it instructions.
| Claude Sonnet 5.5, 24 runs | Passed | Seconds per task | $ per task |
|---|---|---|---|
| thunc, before | 12/24 | 99 | 0.084 |
| thunc, after (native calls) | 24/24 | 10 | 0.022 |
| Claude Code | 24/24 | 11 | 0.069 |
Same model, same tasks: thunc agents now pass every run, slightly faster than Claude Code and at about a third of its cost. That shipped as 0.2.2 on the morning of 5 October, together with thunc run --profile.
Closing out 0.2
The benchmark left a list of open items, and 0.2.3 was scoped to finish all of them before anything new: eight workstreams, one pull request each, stacked on one another. A few worth telling:
- Agents renamed symbols by writing throwaway scripts. Faced with 26 separate edits, models wrote a Python script to do the rename instead.
editcan now replace every occurrence or make several changes in one call. editno longer needs a priorread. With a shell, agents read files withcat, andeditrefused them 7 times in 24 runs. It only changes text the agent quotes exactly, so it can't overwrite what the agent hasn't seen.- The text protocol reads messy replies. A quarter of Sonnet's text-protocol replies had been sent back as invalid with a correct action inside them. Now the first complete action is used, and the model is told to send the JSON alone next time. One run had batched
[search, read, finish 0.0], a guess written before its own read came back, sofinishalongside other calls is now refused. - A reply that stalled for an hour. The first live run on the Claude API hung on one Opus reply: the stream sent nothing, while the API's keep-alive pings kept the SDK's read timeout from firing. A reply that sends nothing for
timeoutseconds is now stopped and asked again. Replies are streamed with room for large writes,max_tokensno longer ends a run, and agents think at efforthighby default. - Native calls on Codex, through the same relay as Claude Code: 30 seconds a task against 45 for the text protocol. The third round of runs hit the Codex subscription's usage limit, so that table has two runs per task instead of three.
- Durable runs with native calls on Claude Code. This one was gated on a spike: kill Claude Code in the middle of a tool call, then resume its session. It worked, with a twist. Claude Code marks the interrupted call as having an unknown outcome and the model asks for it again, so the durable runtime maps that repeated call to the same journal entry: replayed if it finished, held for
resolve()if nobody knows.
In the rerun, every harness passed 24 of 24, the text protocol included (20 of 24 in 0.2.2) in about half its previous time, and through the Claude API both Sonnet 5.5 and Opus 5.5 passed 8 of 8. 0.2.3 shipped on 6 October as the last of the 0.2 line.
Watching a program
A program that makes dozens of calls and runs agents is hard to follow from its logs. thunc watch app.py runs it with a dashboard in the terminal: the calls waiting on a model, retries and why each reply was rejected, timings for each function, each agent's steps as they happen, and the profile report when the program ends. --agents follows agent runs from any process, and --plain prints one line per event for CI.
The dashboard is written in Rust with ratatui, which raised a packaging question. Bundling a compiled binary would have turned thunc into a set of platform wheels and broken installs where no wheel exists. So the dashboard ships as its own package, thunc-watch, installed with pip install "thunc[watch]", and thunc stays pure Python. thunc only writes one JSON line per event to the file named by THUNC_EVENTS, and nothing at all when it isn't set. See Watching a program.
Functions that write themselves
The other 0.3 feature started from a simple question: if the model can answer a function's calls, why not have it write the function? With @thunc.function(write=True), the first call asks for three things side by side: this call's answer, a draft of the body from the docstring and signature, and five test calls from a separate request, each answered by the model. The draft is linted, run on all six inputs and must match every answer. A passing draft goes into your file in place of ..., the decorator is removed, the checked calls become doctest examples, and from then on the function is plain Python with no model calls.
thunc: writing minutes() in durations.py (first call)
thunc: checked against 6 model answers: all agree
thunc: wrote durations.py lines 5-27 in 14s (answer 4s, draft 11s, test calls 9s; side by side). Removed @thunc.function. Review: git diff durations.pyIt took two tries to build the right thing. The first attempt, a day earlier, built something else: hand-written function bodies that fall back to the model when they're unsure. It was a reasonable feature but not the one that had been asked for, so it was reverted the same day, unreleased. The second attempt started by writing down which idea it delivered. Its first version used a whole agent as the writer and took about 39 seconds on Codex; replacing it with one typed call, and asking for the test calls separately so they don't share the draft's blind spots, brought that to 20.
thunc write edits your source files, so it ships as an experiment. It's refused in CI, in installed code, in files outside the project and in files changed since they were imported, and every change is a diff to review. See thunc write.
The second release mistake. 0.3.0 was published on 7 October from a draft made by the release-notes workflow. Editing the draft to point at a newer commit didn't move it: the tag was made on the commit the draft was created on, and two changes merged just before the release missed the package. 0.3.1 shipped them 13 minutes later. Tags are now pushed by hand at the merge commit and checked before a release is published.
A stability pass
After a few days of new features, a review of 0.3.1 looked for ways a run could fail that it shouldn't. It found two.
When the Claude API sends an error in the middle of a streamed reply, say an overloaded server, the SDK raises it with the stream's HTTP status, 200. thunc read that as permanent and the agent run failed, though the docs promise such errors are retried. Now the error's type decides. In a simulation of 10-step runs with real thunc agents and a scripted client:
| Requests failing mid-reply | Runs finished in 0.3.1 | In 0.3.2 |
|---|---|---|
| 1% | 90.8% | 100.0% |
| 5% | 58.9% | 99.9% |
| 10% | 33.4% | 98.8% |
| 20% | 10.0% | 92.1% |
The second was a trace file that couldn't be written, in a missing folder for example. The error replaced the call's outcome: a successful call lost its answer, already paid for, and a failed one hid its own error. Now thunc warns once per path and the call's result stands.
What we learned
- Measure the harness, not only the model. The biggest improvement in these eight days, 12 of 24 to 24 of 24, didn't come from a better model. It came from a benchmark that held the model fixed and asked what thunc was doing wrong.
- Every change comes with a number. Each release note says what was measured and how: runs passed, seconds and dollars per task, runs that survive an error rate. It keeps claims honest and shows which ideas didn't pay off.
- Attack your own code, then simplify it. Adversarial rounds found real bugs, and then found bugs in the fixes. Removing the clever rules did more than patching them again.
- Releasing is code too. Two releases went out from the wrong commit. Both times the fix was a written-down order of steps, not more care.
- Keep the core small. Temporal and the dashboard are real dependencies, so they're an extra and a separate package.
pip install thuncstill installs nothing else. - Say which idea you're building. One day's work went to the wrong feature because the plan didn't say plainly what it would and wouldn't deliver.
What's next
thunc is in beta, and the 0.3 line is about using it on real programs. The open candidates are on the issue tracker: Enum and Pydantic return types (#6, #7), record and replay for tests (#8), a precise type for thunc.call(returns=...) (#10) and more examples (#11). Whether thunc write leaves its experimental label depends on how it holds up in your code. If you try it, tell us how it went.