Changelog
Every thunc release, newest first. Each one links to its full notes on GitHub and its page on PyPI. Upgrade with pip install --upgrade thunc.
| Version | Released | Headline |
|---|---|---|
0.4.0 | 11 Oct 2026 | Record and replay for tests and CI; Enum return types; thunc.acall and thunc.amap; precise types for returns=; deprecations |
0.3.2 | 10 Oct 2026 | Agent runs on the Claude API retry errors sent mid-reply; a trace that can't be written warns instead of failing the call |
0.3.1 | 7 Oct 2026 | The 0.3.0 changes that missed its package: a fixed order in thunc write's timing message, and the README |
0.3.0 | 7 Oct 2026 | thunc watch: a live dashboard in the terminal for calls and agent runs; thunc write, functions that write themselves (experimental) |
0.2.3 | 6 Oct 2026 | Native tool calls on Codex and in durable runs on Claude Code; durable tools=; a sturdier Claude API path; multi-edit |
0.2.2 | 5 Oct 2026 | Agents use their tools reliably; native tool calls on Claude Code; faster calls; thunc run --profile |
0.2.1 | 4 Oct 2026 | Durable agents with Temporal |
0.2.0 | 4 Oct 2026 | Agents: typed tasks that read, edit and run commands under permissions |
0.1.3 | 4 Oct 2026 | The Jev backend |
0.1.2 | 4 Oct 2026 | cache=True, system= and sturdier parsing |
0.1.1 | 3 Oct 2026 | The OpenAI backend and local models |
0.1.0 | 3 Oct 2026 | First release |
0.4.0 beta
Functions complete, and testable without a model. Return an Enum; get the precise type of thunc.call(..., returns=Literal[...]) from your type checker; await calls with thunc.acall and thunc.amap; and record a test suite's answers once, then replay them in CI with no model, backend or API key. A program that uses none of these behaves as in 0.3.2.
Added
- Record and replay (#8). With
configure(recordings="tests/recordings")orTHUNC_RECORDINGS, every call and agent run is answered from that folder and no model is asked; one that wasn't recorded raisesThuncError.THUNC_RECORD=1asks the model and saves every answer;THUNC_RECORD=missingonly the ones missing. Calls are matched on their function, system prompt and request, not the backend or model, so CI needs no key. Agent runs replay the model's replies while their tools run for real. While replaying, a durable run fails instead of reaching a model. - Enum return types (#6). The member comes back, matched on its value; the prompt shows names where a value doesn't say what it means (
3 (HIGH)). Onjev, an enum of strings is a choice and one of integers a score. - The async API.
await thunc.acall(...)isthunc.callto await;await thunc.amap(func, items, workers=8)runs up toworkerscalls at once in input order, forasync defand plain functions. thunc.ThuncDeprecationWarningand the deprecation policy: at least two minor releases of warnings before anything is removed. Nothing is deprecated yet.- Benchmarks:
record.*.
Changed
returns=is typed as a type form (#10), sothunc.call(..., returns=Literal["a", "b"])is aLiteral["a", "b"]to current mypy and Pyright, notAny. Nothing is added at run time.
Fixed
thunc.mapover anasync deffunction returns its values, not coroutines that were never awaited.
0.3.2 beta
Steadier agent runs on the Claude API, and a trace that can't be written no longer costs the answer. Two fixes from a stability review of 0.3.1. Nothing else is new.
Fixed
- An error the Claude API sends mid-reply is retried when asking again may fix it. An overloaded, rate-limit, timeout or server error that arrives while a reply streams comes with the stream's HTTP status, 200, so thunc took it for a permanent error and the agent run failed. It's now retried like the same error arriving before the reply: twice, with a pause, as the docs say. A bad request, a bad key or a billing error still fails at once. In a simulation of 10-step runs with 5% of requests failing this way, 99.9% of runs finish, up from 58.9% (#68).
- A trace file that can't be written warns instead of failing the call. With
trace=orTHUNC_TRACEnaming a folder that doesn't exist or a read-only file, every call and agent run raised the file's error: a successful call lost its answer, already paid for, and a failed one hid its own error. Now thunc warns once per path, and the call returns its value or raises its own error (#68).
0.3.1 beta
The 0.3.0 changes that missed its package. The v0.3.0 tag was made on an earlier commit than intended, so two changes merged before the release didn't reach PyPI. Nothing else is new.
Fixed
- thunc write's timing message lists its parts in a fixed order: answer, draft, test calls. The three run side by side and were listed in the order they finished, so the message changed from run to run (#65).
Docs
- The README on PyPI says beta v0.3 and its intro mentions
thunc watchand thunc write. The website's landing page has a section for each, the Functions guide coverswrite=True, and Get started links both guides (#66).
0.3.0 beta
Watch a program's calls and agent runs live. thunc watch app.py runs a program with a dashboard in the terminal: the calls waiting on a model, retries and why each reply was rejected, timings for each function, each agent's steps as they happen, and the --profile report when the program ends. --agents follows agent runs from any process, --plain prints one line per event for CI, and a saved run can be replayed. The dashboard is a compiled binary in its own package, thunc-watch 0.1, installed with pip install "thunc[watch]", so thunc itself stays pure Python with no dependencies. And, as an experiment, thunc write: a function declared with @thunc.function(write=True) writes its own body into your file on its first call, checked against the model's answers, and runs as plain Python from then on. A program that uses neither behaves as in 0.2.3. See Watching a program and thunc write.
Added
thunc watch, a live dashboard in the terminal for a program's thunc calls and agent runs: calls in flight, retries and why each reply was rejected, per-function timings, each agent's steps as they happen, and the--profilereport when the program ends. Click around with the mouse or use the keys.thunc watch app.pyruns a program and watches it;thunc watch --agentsfollows agent runs from any process;--plainprints one line per event for CI. It's a compiled binary in its own package, thunc-watch, installed withpip install "thunc[watch]", so thunc itself stays pure Python with no dependencies.THUNC_EVENTS=FILEwrites one JSON line per call, attempt and agent step to FILE, which thunc-watch reads. Inputs, replies and values are cut to short previews unlessTHUNC_EVENTS_CAPTURE=1. Nothing is written when it isn't set.- thunc write (experimental): functions that write themselves, with
@thunc.function(write=True). Its behavior, options and the code it writes may change, or it may be removed, in a later release without a deprecation period. On its first call the function writes its own body. Three requests start side by side: the call's answer, a draft of the body from the docstring and signature (with the whole file and any project types in the signature), and five test calls, each answered by the model on its own. The draft is linted and run on the call and the test calls within a time limit, and must match every answer; one that doesn't goes back with the failing calls, up to three drafts. A passing draft goes into the file in place of..., the decorator is removed (import thuncstays), the checked calls become doctest examples, and the call runs the new code: from then on it's plain Python. If the model says the task needs judgment, or no draft passes, the call returns the model's answer, the file is left as it was, and the reason is kept in.thunc_write/until the docstring or signature changes. Writing is refused, with a warning, in CI or withTHUNC_WRITE=0, outside the project, and for read-only, installed or changed files and nested functions. - The
thunc write FILE::FUNCTION [--dry-run]command (experimental) writes awrite=Truefunction ahead of its first call, or shows the change as a diff.
0.2.3 beta
Native tool calls everywhere agents run, durable runs that keep them, and the last of 0.2. Agents on Codex make native tool calls through the same MCP relay as Claude Code, durable runs on Claude Code do too and pick up after a crash mid-call, and durable agents can take tools=. On the Claude API, agents stream their replies, think at effort high, and no longer lose a run to max_tokens or a stalled reply. edit can replace every occurrence or make several changes at once, and no longer needs a prior read. In the tool-use benchmark (live_tests/bench_tooluse.py, 8 tasks, Claude Sonnet 5.5, 3 runs each), every harness passed 24 of 24, the text protocol included (20 of 24 in 0.2.2) in about half the time; through the Claude API, Sonnet 5.5 and Opus 5.5 passed 8 of 8; on Codex, native calls took 30 seconds a task against 45 for the text protocol.
Behavior changes
- Agents on
codexmake native tool calls, as on Claude Code since 0.2.2 (see Added). When Codex can't start them, a run falls back to the text protocol with a warning;protocol="text"keeps the old way, and durable runs on Codex still use it. finishcalled in the same reply as other calls is refused (except besideremember): its value can't account for results the model hasn't seen yet. The other calls run, and the model is told to callfinishon its own. In the tool-use benchmark, a run on the text protocol batched[search, read, finish 0.0]and returned the guess. Every protocol and durable runs get it.- Agents on the Claude API think at effort
highby default on Claude 4.6 and later. Claude Opus 5.5's own default ismedium, which is low for agentic coding.effort=changes it (below). editno longer needs the file to have been read first. It only changes text the agent quotes exactly, so it can't overwrite what the agent hasn't seen. With theshellpermission, agents often read files withcat, andeditrefused them until they read the file again withread: 7 times in 24 runs of the tool-use benchmark. A file the agent did read must still not have changed on disk since, andwritestill replaces only a file read withread(a file changed by an edit alone still counts as unread). Theruntool's description, withshell, now says to read files withread.
Added
- Durable runs on Claude Code make native tool calls. A run goes in segments: one activity keeps one
claude -pprocess for many model replies, and before each reply's calls are carried out it saves a checkpoint of the run and of Claude Code's session. If the worker stops or the CLI dies, the retried activity restores the session and continues it with--resume; Claude Code marks the call that was in flight as interrupted, the model asks for it again, and it gets the journal entry it had, so it's replayed, waits forresolve(), or runs (one the model doesn't ask for again keeps its entry). The session is removed when the run ends. When Claude Code can't start native calls, the run goes on with the text protocol. Runs already in progress keep the text protocol. - Durable agents can have
tools=. Each call of one of the agent's own functions is journaled like a command: the intent is recorded before it runs and its result after, so a retried activity replays the result instead of calling it again, and a call interrupted by a worker stopping waits forresolve().registry.agent_task(..., retry_safe_tools=["find_issue"])names the tools that may run again instead. A tool's description, arguments and retry marking are part of the task's fingerprint, so changing one needs a new version; tasks without tools keep their fingerprint. - Native calls on
codex: the agent's tools are an MCP server (thunc's relay, given with-c mcp_servers.thunc.*) that Codex calls; thunc carries out each call with its own tools, permissions and run record. Eachcodex execis a turn: one that ends withoutfinishis continued withcodex exec resume, as is one that fails (twice at most). Codex's own tools and your~/.codexconfig stay out and its sandbox stays read-only; only thunc's server is approved to run without asking, with a tool timeout abovecommand_timeout. Resuming needs the session saved, so thunc deletes the run's Codex session (codex delete --force) when the run ends. Agent(effort=...):"low","medium","high","xhigh"or"max", on every backend (output_config.efforton the Claude API,reasoning.efforton OpenAI,--efforton Claude Code,model_reasoning_efforton Codex; OpenAI and Codex go up to"xhigh"). Recorded inagent.jsononly when set, so durable tasks registered without it keep their fingerprint.- The model is told when few steps are left. In its last three replies before
max_steps, the last tool result says how many replies remain, so the model can finish with what it has instead of being cut off. Every protocol and durable runs get it; the run record keeps each tool's output. editcan replace every occurrence, and make several changes in one call. With"replace_all": true, every occurrence ofoldis replaced and the result gives the count. With"edits": [{"old": ..., "new": ..., "replace_all"?: ...}, ...](at most 50) instead ofoldandnew, the changes apply in order, each to the text the ones before it left; if one fails, none is made, and the error names it. In the tool-use benchmark, models renamed a symbol by writing a throwaway script instead of making 26 separate edits, and took twice Claude Code's turns on a multi-spot fix. In a durable run, a multi-edit is one effect, recovered as a whole.
Changed
- The text protocol reads replies that aren't only the action (finding 1 of the tool-use benchmark report). Models trained for native tool calls often wrap the action in prose, a code fence or
<invoke>markup, or carry on past it with results they make up: on Sonnet 5.5, 25% of text-protocol replies were sent back as "not valid JSON" with a correct action inside. Now the first complete action (or array of them) in the reply is used, with literal newlines in its strings accepted, and a reply written only as<invoke name="...">markup is read as its calls, each argument in its tool's type (or as one JSONargsparameter). A reply with no action in it is still sent back, as before. A reply read this way runs, but its results carry a note to reply with the JSON action alone: without it, a model that slipped into markup was never corrected, and on Sonnet 5.5 fell into repeating empty markup until the step timed out. This is the text protocol on Codex,protocol="text", durable runs on Codex, and Claude Code's fallback from native calls. - A text-protocol step on Claude Code stops once its action is complete. The reply is streamed (
--output-format stream-json --include-partial-messages) and the CLI is stopped as soon as a complete action has arrived, rather than left to make up the tool's result until the step times out (7 of 8 replayed first steps on Sonnet ran past 120 seconds that way). A batch that has begun is waited for, and two blocks of<invoke>markup end the step too (a model repeating itself). Plain@thunc.functioncalls on Claude Code aren't streamed. - The Claude API agent path keeps runs going (finding 7 of the tool-use benchmark report):
- Replies are streamed with
max_tokens=64000(was 16,000 without streaming), so a large write fits. - A reply cut off at
max_tokensno longer ends the run: its tool calls aren't run and get an error result saying so, and the model is asked again. Two in a row end the run. pause_turnis asked to carry on (up to 6 times in a row) instead of ending the run.model_context_window_exceededends it with a message that says so.- On Claude 4.6 and later, the API clears old tool results on long runs (context editing, beta
context-management-2025-06-27). - Tools are
strict(arguments guaranteed to match their schema) on the models that support it, when the schema allows it: the built-in tools exceptedit,remember, andfinishand custom tools whose schema is closed. - A reply that sends nothing for
timeoutseconds (300 by default) is stopped and asked again. The SDK's read timeout doesn't catch it, as the API's keep-alive pings count as reading: in the tool-use benchmark, one reply on Claude Opus 5.5 sent nothing for an hour. - A lost connection, a rate limit or a server error on the Claude and OpenAI APIs is retried as a step, like the CLI backends' errors, instead of ending the run.
- Replies are streamed with
Fixed
- A large system prompt no longer stops the
claude-codebackend from starting. It went on the command line, so largefollow=files and memory could pass the operating system's limit on its length (128 KB for one argument on Linux, 32,767 characters for the whole line on Windows), and startingclaudefailed with a rawOSError: Argument list too long. The prompt now goes in a temporary file (--system-prompt-file), as it already did for agents' native calls and on Codex. This covers@thunc.functionandthunc.call, agents on the text protocol, and durable runs on Claude Code. A command line that is still too long raises aThuncErrorthat gives its size.
0.2.2 beta
Agents that use their tools reliably, and faster calls. Agents on Claude Code make native tool calls instead of writing each action as JSON text, a failed step is retried instead of ending the run, and the agent's tools fill gaps a benchmark found. In that tool-use benchmark (live_tests/bench_tooluse.py, 8 tasks, Claude Sonnet 5.5, 3 runs each), agents on Claude Code went from 12 of 24 runs passing to 24 of 24, from 99 to 10 seconds a task, and from $0.084 to $0.022 a task; Claude Code itself took 11 seconds and $0.069. The API backends also reuse their connections, Codex answers return sooner, and thunc run --profile shows where a program's time goes.
Behavior changes
Nothing is removed, but these defaults change:
- Agents on
claude-codemake native tool calls (see Added). With aclaudeCLI too old for them, or with MCP servers turned off by a policy, a run falls back to the text protocol with a warning.protocol="text"keeps the old way. listandsearchleave out what git ignores in a git repository (build output, caches, vendored code). A folder named explicitly is still listed and searched.- Long command output keeps its start and its end (the first error and the summary), not only the end.
- The
claude-codebackend loads none of your Claude Code settings (--setting-sources ""): noCLAUDE.md, settings or hooks reach thunc's calls, plain function calls included, so an agent'sworkdircan't give it instructions unlessfollow=asks for them.
Added
- Native calls on
claude-code: the agent's tools are an MCP server that oneclaude -pprocess per run calls; thunc carries out each call with its own tools, permissions and run record. The CLI runs in the agent'sworkdir.protocol="text"keeps the old way, and durable runs on Claude Code still use it. When Claude Code can't start native calls (an older CLI, or MCP servers turned off by a policy), a run falls back to the text protocol with a warning and afallbackentry in its record;protocol="native"raises instead. runtakescwd, a folder insideworkdirto run the command in.- The
shellpermission runs command lines through the system shell, so pipes,&&,cdand redirects work. Off by default; it can't be combined with!run:rules. searchtakesglob(*.pyby file name,src/**/*.tsby path) to limit the files searched.thunc run --profile: runs a script (or-m module) and prints a performance report to stderr when it ends: per function, calls, cache hits, retries, failures, total/mean/p95/max time and the split between model time and thunc's own; for agents, steps and time in each tool; and the share of wall time spent in thunc, with the overlap fromthunc.map.
Changed
- A failed step is retried. A timeout, lost connection, rate limit, server error or CLI call that ended in an error (
thunc.errors.TransientError) is retried twice in an agent run, with a note in the run record, before the run fails. A text-protocol step on Claude Code or Codex may take 120 seconds before it's retried, instead of the wholetimeout. - The
anthropicandopenaibackends reuse their connections. One SDK client is shared by every call in the process (thunc.map's threads and agent runs included), instead of a new client, and so a new TCP and TLS handshake, for each call. A new client is made when the API key, the SDK's environment variables (ANTHROPIC_*,OPENAI_*) or the process change. In a local benchmark with 60 ms of connection setup, 20 calls in a row went from 1.47 s to 68 ms. - Agents on the text protocol can act several times per reply. On Codex and
protocol="text"(and on Claude Code when it falls back to the text protocol), a reply can be a JSON array of independent actions (reading three files) instead of one. They run in order, at most 16 per reply, and every result comes back together, as with native tool calls. Each turn resends the whole transcript and, on the CLI backends, starts the CLI, so fewer turns save both. A single JSON action works as before. On Codex, withlive_tests/eval_prompts.py(default prompt, 5 runs of each task), every run batched its first reads: replies went from 6.0 / 4.8 / 5.0 to 5.0 / 3.0 / 3.6 (fix / review / analysis) and the mean time from 37 / 27 / 27 s to 29 / 19 / 21 s, with the same work done and 30/30 passing. - The
codexbackend returns as soon as the answer arrives. It reads Codex's JSON events as they come (codex exec --json) instead of waiting for the process to exit and reading the answer from a file. Codex takes about 0.4 s to shut down after answering; that now happens in the background. Over 8 alternating pairs of real calls the new way was faster every time, by a median of 0.67 s on a call of about 4 s. Every agent turn on Codex is one call, so the saving repeats.
0.2.1 beta
Durable agents with Temporal. An optional thunc[temporal] runtime records each model turn and tool call in a Temporal workflow, so a run survives worker restarts and can be reattached from another process. Local thunc stays dependency-free. See the Temporal guide.
Added
thunc.temporal:Registry,Worker,RuntimeandHandleto register versioned tasks and start, reattach to, inspect, cancel and resolve durable runs. One coordinator per workspace runs requests in order; a repeated request ID reattaches to the same run. Agents withtools=ortimeout=can't be registered for durable runs yet (#37).- Recoverable tool effects: file writes and memory notes go through an intent and receipt journal with atomic replacement and content hashes. A command whose outcome is uncertain is never rerun automatically: the run waits for an operator's
resolve()(#37). thunc.temporal.adapters.execute_taskcomposes registered tasks from native Temporal workflows, with a classify → agent analysis → typed summary example (#38).
Changed
- The agent loop's decisions moved into a shared engine (
thunc/execution.py) that local and durable runs both use. Aremembercall that fails no longer appears inRun.notes(#36).
0.2.0 beta
Agents. An agent is a typed function that can look around before it answers: it lists, reads and searches files in a working directory, and, when its permissions allow, writes, edits and runs commands, then returns a checked value of the task's return type. See the Agents section of the README.
Added
thunc.Agent(name, workdir=...)and@agent.task: declare tasks like@thunc.function. Read-only by default, withlist,read,searchandremembertools. Sync and async tasks (#23).- Memory and run files in
.thunc_agents/<name>/:memory.md(notes kept between runs),agent.json, and one JSONL record per run. Runs of one agent take turns through an OS file lock (#23). - Permission rules:
write:,read:,run:and!denies with globs, plus thewriteandedittools. Edits need a fresh read of the file in the same run (#26). - The
runtool: commands allowed byrun:rules run without a shell, with a minimal environment and a time limit that also stops their child processes (#28). agent.run(task, ...)returns athunc.Runwith the value, files changed, commands, denials, notes and steps. A failed run raisesthunc.AgentErrorwith the partial record (#29).follow=gives the agentAGENTS.md/CLAUDE.md, or files you name, as instructions. Off by default (#31).- Native tool calls on the Claude and OpenAI APIs, with the fixed part of the prompt cached. Claude Code and Codex use a JSON text protocol;
protocol="text"picks it on an API too (#32). - System prompt presets:
thunc.prompts.CODING,CODE_REVIEWandANALYSIS, forsystem=(#34). tools=: your own typed, documented Python functions as agent tools (#34).agent.call(...), the agent version ofthunc.call, and the@thunc.agent(...)shorthand for a one-task agent (#34).timeout=bounds a run's time; command limits are cut to the time left (#34).- Files changed by commands are included in
Run.files_changed(#34). live_tests/eval_prompts.py, an evaluation of the agent system prompt (bare / default / preset). 90/90 runs passed on the Claude API and Claude Code (#34).- CI now also runs the offline tests on Windows (#34).
Changed
thunc.agentis the decorator for one-task agents.from thunc.agent import Agentstill works; onlyimport thunc.agent as mnow gives the decorator rather than the module.- The
codexbackend leaves out Codex's own permission notes, which made it refuse allowed edits (#26). - Agents refuse the
jevbackend with aThuncErrorbefore a run starts; it only answers typed questions (#33).
Nothing changes for @thunc.function and thunc.call.
0.1.3 beta
jevbackend for TypeSafe's Jev judgment model:boolandLiteralanswers in about 0.3 s (#27).- The
codexbackend runs with Codex's own tools off and ignores~/.codex/config.toml(#25). @thunc.functionbodies likereturn 1now raiseTypeErrorat definition (#24).
0.1.2 beta
cache=Truesaves valid answers on disk;thunc.clear_cache(),thunc.cache_info()and thethunc cachecommand manage them (#18, #19).system=replaces the opening of the default system prompt; Codex gets it as its instructions file (#21).- Sturdier parsing: common near-misses are read, and wrong values are retried instead of returned. Only
ThuncErrorescapes (#20). - An empty reply is no longer a valid
str.
0.1.1 beta
openaibackend on the Responses API, and local models throughOPENAI_BASE_URL.- The Claude API backend tested live; CONTRIBUTING.md and Discussions added.
0.1.0 beta
- First release:
@thunc.function,thunc.call, typed and validated results with retries andensure=,thunc.map, JSONL tracing; theanthropic,claude-codeandcodexbackends.