Claude Code skill
thunc-offload is a skill that teaches Claude Code when and how to hand work to thunc. Long, self-contained jobs go to a thunc agent on a cheaper model, and Claude Code checks the result. Batch work, such as classifying hundreds of log lines, runs as typed functions, so none of it lands in the conversation.
Install
The skill is one file. Download it into ~/.claude/skills/thunc-offload/ for every project:
mkdir -p ~/.claude/skills/thunc-offload
curl -fsSL https://eltarras.github.io/thunc/skills/thunc-offload/SKILL.md -o ~/.claude/skills/thunc-offload/SKILL.mdTo share it with everyone on a project, put it in the repository's .claude/skills/thunc-offload/SKILL.md instead and commit it. Claude Code finds new skills when a session starts.
It needs Claude Code, logged in, and Python 3.10 or later. thunc itself comes from uv when a script runs (uv run --with thunc), or from pip install thunc. The skill runs thunc on the claude-code backend, so it uses your Claude Code login and no API key. It only works in Claude Code, not in Claude apps without a local claude CLI.
Using it
The skill is for two kinds of task:
- Long agentic work with a checkable outcome: a wide refactor or rename, a debugging hunt with no obvious lead, a feature with tests. The rough bar is a task that would take Claude Code more than about 10 turns.
- The same judgment applied to many items: classifying, extracting or summarizing across more than about 10 files, log lines, rows or tickets.
Ask for it. Claude Code sees the skill in every session, but in our tests it never handed work off on its own, not even on a feature task where handing off was 21% cheaper. Say "offload this to thunc", or type /thunc-offload.
For agentic work, Claude Code then:
- writes the job as an agent task with a typed result, the narrowest permissions that work, and everything the agent needs in the docstring, because the agent can't see the conversation;
- runs it on Sonnet through your Claude Code login, with the agent's run records kept in
~/.cache/thunc/agents, out of your repo; - waits for the run to finish, then reads the diff and reruns the tests itself instead of trusting the agent's report;
- tells you what it handed off, to which model, and what it checked.
For batch work, it writes a typed function on Haiku, runs it with thunc.map, caches the answers when one answer per input is right, and reads back only the summary.
What it saves
We gave Claude Code on Opus 5.5 a feature to build from a spec: mid-month plan changes with prorated billing, storage, a CLI command and reporting, across 4 or 5 modules of a small billing app. 24 hidden tests the agents never saw graded each result, and each setup ran 3 times. The costs include everything: the main Opus session, the thunc agent, and the main session checking the work.
| Setup | Cost per run | Range | Hidden tests passed |
|---|---|---|---|
| Claude Code alone | $0.52 | $0.45–0.62 | 72/72 |
| Claude Code, handing off to a thunc agent on Sonnet | $0.41 $0.30 main session + $0.11 agent | $0.38–0.45 | 72/72 |
21% cheaper overall, with the same pass rate. Every handoff run cost less than every run of Claude Code alone, and the wall time was about the same (108 against 102 seconds).
Most of the remaining cost is the main session. The thunc agent built the feature for $0.11, but Claude Code still read the code, wrote the task and checked the diff, which cost $0.30. That's why the saving on the handed-off work is much bigger than the saving overall. In thunc's tool-use benchmark, a thunc agent and Claude Code each did 8 coding tasks on their own, with the same model, and every run passed:
| Model | Claude Code, per task | thunc agent, per task | Saving |
|---|---|---|---|
| Claude Opus 5.5 | $0.174 | $0.039 | 78% |
| Claude Sonnet 5.5 | $0.078 | $0.021 | 73% |
Agent on its own, not counting a main session: thunc's prompt is about a fifth of the size of Claude Code's. Everything ran on a Claude Code subscription; costs are tokens at API prices. Details: the benchmark report.
Since the main session's part is roughly fixed, the overall saving depends on the size of the task. On small bug fixes across three files, handing off cost about the same as Claude Code fixing the bugs itself, about $0.27. These numbers come from one feature task and a few small ones, 3 runs or fewer each, so treat them as a guide.
Rule of thumb: let Claude Code fix small things itself. Ask it to hand off long jobs, and batch work that would otherwise fill the conversation with files and outputs.
What the agent may do
The skill tells Claude Code to grant the narrowest permissions that work, such as write:src/** and run:pytest, and never git commit or git push unless you asked. Permissions limit which tools the model uses, but they aren't a sandbox: a permitted command such as pytest runs your project's code. Work in a git repository with a clean tree, so every change is a diff you can review.