smolagents replays its whole memory every step: the O(n²) token bill nobody mentions
Cost-audit series, episode 5. This series began with an AI agent that burned 136M tokens overnight →.
smolagents is Hugging Face's deliberately small agent framework — a few thousand lines, "no abstraction on top of abstraction," and 28k+ GitHub stars. Its CodeAgent is genuinely elegant: the model writes Python, the runtime executes it, the result comes back, repeat until a final answer.
The elegance hides a cost curve. A smolagents run that takes n reasoning steps does not cost n times a single step. On input tokens it costs closer to n²/2, because every step re-sends the entire accumulated memory of every previous step. The default step budget is 20. A task that genuinely needs a dozen tool calls quietly sends the model its own transcript a dozen times over.
This audit shows the exact lines, gives you a formula you can evaluate on your own workload, and shows the one-hook fix.
Where the tokens go
Every step, the agent rebuilds the full message list it sends to the model. Here is the method that does it, verbatim (agents.py:758-770, v1.26.0):
def write_memory_to_messages(
self,
summary_mode: bool = False,
) -> list[ChatMessage]:
"""
Reads past llm_outputs, actions, and observations or errors from the memory into a series of messages
that can be used as input to the LLM. Adds a number of keywords (such as PLAN, error, etc) to help
the LLM.
"""
messages = self.memory.system_prompt.to_messages(summary_mode=summary_mode)
for memory_step in self.memory.steps:
messages.extend(memory_step.to_messages(summary_mode=summary_mode))
return messages
Read the loop: it walks every entry in self.memory.steps and appends its messages. memory.steps only ever grows — it is a plain list initialised empty and appended to, never trimmed, except by an explicit reset() between runs (memory.py:230, memory.py:232-234).
That method is called at the top of every action step, with no summary_mode, so the full history is replayed each time (agents.py:1284-1286):
memory_messages = self.write_memory_to_messages()
...
input_messages = memory_messages.copy()
The default ceiling on how many times this can happen is 20 (agents.py:300, max_steps: int = 20). Nothing in the default path caps or summarises the growing memory — summary_mode=True is used only for planning messages (agents.py:684, agents.py:886), not for the main action loop.
The math (evaluate it on your own numbers)
Let:
-
P= tokens in the system prompt (fixed, sent every step) -
s= tokens each completed step adds to memory — the model's code/thought plus the tool observation it produced
At step k (1-indexed) the input the model receives is P + (k-1)·s — the prompt plus everything the previous k-1 steps left behind. Summed over an n-step run, cumulative input tokens are:
Σ (k=1..n) [ P + (k-1)·s ] = n·P + s · n(n-1)/2
The s · n(n-1)/2 term is quadratic in n. Compare it to the intuition most people price with — "n steps ≈ n × one step" = n·(P + s). The history you re-pay for is n(n-1)/2 · s instead of n · s:
| Steps n | History replays (× s), naive |
History replays (× s), actual |
Overpay factor |
|---|---|---|---|
| 4 | 4 | 6 | 1.5× |
| 8 | 8 | 28 | 3.5× |
| 12 | 12 | 66 | 5.5× |
| 20 (max) | 20 | 190 | 9.5× |
(Table is illustrative of the formula above — it is the closed form n(n-1)/2 vs n, not a measured run. Plug in your own P and s to get dollars.)
The observation size s is where it bites hardest. If a tool returns a chunk of a web page, a file, or a dataframe, that payload is now re-sent on every subsequent step for the rest of the run. Long, tool-heavy tasks are exactly the ones that hit max_steps, so the worst tasks pay the worst multiplier.
Good news: smolagents already measures this for you. Every ActionStep carries a token_usage field with input/output token counts (memory.py:63). After a run, sum step.token_usage.input_tokens across agent.memory.steps and you will see the curve directly. The problem is that by the time you read it, you have already paid.
The fix: prune the replayed history with a step callback
smolagents gives you the exact hook you need. The agent accepts step_callbacks — callables invoked at the end of each step, and you can register them per step-type (agents.py:282, agents.py:304, wired in _setup_step_callbacks, agents.py:416-425). Because memory.steps is just a list you own, a callback can cap how much history survives into the next write_memory_to_messages call:
from smolagents import CodeAgent, ActionStep
KEEP_LAST = 6 # replay only the most recent N action steps
def trim_memory(step, agent):
action_steps = [s for s in agent.memory.steps if isinstance(s, ActionStep)]
for stale in action_steps[:-KEEP_LAST]:
# collapse the bulky observation; keep a short marker so the model
# still knows the step happened
stale.observations = "[trimmed to control context cost]"
agent = CodeAgent(tools=[...], model=..., step_callbacks=[trim_memory])
(Illustrative usage of the real step_callbacks API — tune KEEP_LAST and what you collapse to your task. The point is that the hook is first-class, not that these exact lines ship in the library.)
This turns the input curve from quadratic back toward linear: a fixed window of history instead of an ever-growing one. You trade some long-range recall for a bounded bill — for most tool-loop tasks that is the right trade, and you make it deliberately instead of discovering it on an invoice.
Other levers, in order of bluntness: lower max_steps from the default 20 so a wandering run can't rack up 190× history replays; truncate large tool return values before they enter memory; and use the planning/summary path smolagents already has for long-horizon tasks.
See the bill before you run it
The pattern in this series is always the same: the framework is fine, the default is expensive, and the cost is invisible until it shows up on the invoice. smolagents is the most honest case yet — it even hands you token_usage — but you still have to run the task, at full quadratic cost, to see it.
That is the gap tokenscope closes — it shows what a run actually cost, and estimates a source tree's token footprint before you spend it. See a real cost breakdown in five seconds, no setup or logs required:
npx @wartzar-bee/tokenscope --demo
Then run it on your own most-recent Claude Code session (just npx @wartzar-bee/tokenscope), or estimate a directory's token cost before a run — the static check that powers the guardrail:
npx @wartzar-bee/tokenscope scan --dir .
If you want that check enforced automatically — a bot that comments the predicted token-cost delta on the responsible files in every pull request and can block a regression — that is what we build the ci-guardrail GitHub Action for.
Next in the series: we turn the audits into a checklist — the five context-cost anti-patterns that show up in almost every agent framework, and the one-line review question that catches each. Follow @wartzarbee so you don't miss it.
Found an error in this audit? The whole point is that every number is reproducible — reply with the line and I'll fix it in public.
Top comments (3)
The
n²/2framing is the right one, and it's worth stressing that the pricing side makes it worse than the token count alone suggests: those early steps are identical prefixes every turn, so whether that quadratic actually hurts depends entirely on whether your provider's prompt caching is landing. We watched a long-running agent's bill drop hard just by making the system prompt + early tool schemas byte-stable so the cache prefix wouldn't invalidate — no change to the loop at all.The one thing I'd add to the summary-mode fix: summarizing memory is lossy in a way that bites on tasks needing an exact earlier observation (a file path, an ID). We ended up keeping the last k steps verbatim and only summarizing the older tail — a sliding window rather than a global summary. Curious whether you found smolagents'
summary_modeaccurate enough to trust unattended, or whether it silently dropped detail the model later needed?The write_memory_to_messages pattern is naive chat-history accumulation in a production framework — every agent harness that started as a simple append loop discovers this same cost cliff once it reaches multi-step tasks. The structural fix beyond the one-hook patch is to treat replay as a retrieval problem: run a recency-weighted sparse retrieval over past steps at each decision point so effective context grows as O(k log n) rather than O(n), which the summary_mode parameter already hints at but isn't wired into the hot path. The 136M token overnight burn from episode 1 is exactly what happens when O(n²) accumulation combines with a loop condition that doesn't converge. This series is the most practically useful cost-audit work I've seen for agent frameworks — most teams discover this the expensive way.
This is the "repeat tax" in the wild: replaying the whole memory each step doesn't just cost O(n2) tokens, it also evicts any prefix cache you might have had - so you pay full input on top of the quadratic blowup. The fix isn't just truncation; it's cache affinity. Keep a byte-stable system+history prefix at the front and append-only after it. Most frameworks optimize the model call and ignore cache behavior, then wonder why the bill is O(n2).