Caching, tracing and profiling
Save answers on disk so the model is asked once per input, clear them from Python or the command line, record every call to a file, see where a program's time goes, and replay recorded answers in tests.
Caching answers: cache=True
cache=True saves each answer on disk and reuses it when the same inputs come again, so the model is asked once.
@thunc.function(cache=True)
def category(ticket: str) -> Literal["bug", "billing", "other"]:
"""Classify this support ticket."""
...It's off by default, because it only suits some functions:
- Use it for functions that should give one answer per input: classify, extract, score.
- Don't use it for functions meant to vary (drafting a reply, brainstorming), or whose answer depends on something that isn't an input, like today's date. Make that an input instead,
def overdue(deadline: date, today: date) -> bool, and caching becomes safe.
When a saved answer is reused
Only for the exact same function, prompt, backend and model, so changing the docstring, the return type or the model asks again. A saved answer is checked against the return type and ensure= before it's reused, and failed calls are never saved. The function's name is part of the key, so renaming a function starts its cache fresh.
Where answers go
In .thunc_cache/ in the working directory; change it with configure(cache_dir=...) or THUNC_CACHE_DIR. Each call is one JSON file holding the full prompt in plain text, inputs included, so treat the folder like the data you send.
Clearing the cache
Clear everything, or one function's answers, from Python:
thunc.clear_cache() # everything
thunc.clear_cache(urgency) # one function
thunc.clear_cache("urgency") # the same, by name
thunc.clear_cache(older_than=timedelta(days=30)) # answers saved more than 30 days ago
thunc.cache_info() # what's saved, one group per functionOr from the command line:
thunc cache list # saved answers per function
thunc cache clear # everything
thunc cache clear --function urgency # one function (repeat for several)
thunc cache clear --older-than 30d --dry-run # what would go, without deletingA name is the function's name (urgency, or Triage.urgency for a method), optionally with its module (support_inbox.urgency). For thunc.call, pass name="..." to group its answers the same way; unnamed calls are cleared only with everything or by age. Ages count from when the answer was saved. clear_cache returns how many answers it deleted.
Clearing deletes only cache entries, never other files in the folder, and it's safe while another process is using the cache. The thunc command (also python -m thunc) reads THUNC_CACHE_DIR, or takes --cache-dir; it can't see a configure(cache_dir=...) in your code.
Tracing every call
thunc.configure(trace="calls.jsonl")Every call is appended to the file as one JSON line, or set THUNC_TRACE instead. An agent run is one line too, with every model reply in it. log_triage.py and repo_guide.py read their traces back.
Profiling: thunc run --profile
Run your program through the thunc command to see where the time went when it ends:
thunc run --profile support_inbox.py --limit 20 # a script and its arguments
thunc run --profile -m myapp.triage # a module, as with python -mThe report goes to stderr. Per function: the calls, cache hits, retries and failures, the total, mean, p95 and slowest time, and how much of it was the model and how much thunc's own work (building the prompt, parsing, the cache). Agent runs get their steps, model time and time in each tool. It also says what share of the program's wall time was spent in thunc, and how much calls overlapped under thunc.map.
CALLS
FUNCTION CALLS CACHED RETRIES FAILED TOTAL MEAN P95 MAX MODEL LOCAL
urgency 11 1 1 0 1.70s 155ms 309ms 309ms 1.69s 12ms
In thunc: 774ms of 980ms wall time (79%); the rest was the program's own code
Model time: 1.69s, 99% of the time in calls (anthropic/default model 1.69s)
Concurrency: calls overlapped 2.2x on average (thunc.map or threads)
Slowest: urgency took 309msWithout --profile, thunc run just runs the program and records nothing. The program's exit code is passed through.
To watch the same numbers while the program runs, with each call and agent step as it happens, use thunc watch.
Testing code that calls a model: record and replay
Record the answers once, commit them, and replay them in tests and CI with no model, backend or API key. Point thunc at a folder:
import thunc
thunc.configure(recordings="tests/recordings") # or set THUNC_RECORDINGS=tests/recordingsThen fill it once, and replay from then on:
THUNC_RECORD=1 pytest # ask the model, and save every answer in tests/recordings
pytest # replay: no model is asked, and a call that wasn't recorded fails
THUNC_RECORD=missing pytest # replay what's there; ask for, and save, only what's missingWith a recordings folder set, every thunc.call, @thunc.function and agent run is answered from it. A call that isn't there raises ThuncError saying how to record it, instead of reaching a model, so CI never spends money or waits on a login. THUNC_RECORD=1 asks the model and saves (or replaces) every answer; THUNC_RECORD=missing keeps the answers the folder has and asks only for the rest. THUNC_RECORD without a folder is an error, not a silent no-op.
What's matched
A call's function name, system prompt and request (instructions, inputs and return type), so changing any of them needs a new recording. The backend and model are saved with each answer but not matched: CI needs no backend, and a recording made on one model replays under another, so re-record after changing models. The same call twice gets the same recorded answer.
A recorded answer is parsed into the return type and run through ensure= before it's used; one that no longer fits fails with a message saying to record it again. Failed calls are never recorded, only the answer that passed. Replaying never reads or writes the cache, and recording saves a cache hit's answer too.
Agent runs
A run is matched on its agent, task, request, tools and system=, and the model's replies are replayed in order. The tools run for real, so files are written and commands run as when it was recorded: give each test a fresh copy of its workspace. A run whose tools go differently can run out of replies, which fails it with a message saying to record it again. Memory and followed files aren't matched. Durable runs aren't recorded; while replaying, they fail instead of reaching a model.
The recordings folder
Recordings use the cache's format, one JSON file per call or run, holding the full prompt in plain text, so they're easy to review in a diff. thunc cache list --cache-dir tests/recordings lists them per function, and thunc cache clear --cache-dir tests/recordings --function urgency deletes one function's. Answers no test asks for any more stay until you delete them: to prune, delete the folder and record again. The trace marks replayed calls with "replayed": true (and "cached": true: no model was asked), and thunc watch shows them as replayed (from the next thunc-watch release).