10 min
Your model might not be the bottleneck
Notes from Alex Zhang on recursive language models, context rot, and why organizing the work may matter as much as improving the model.
An agent reads a document correctly. Writes a useful function. Spots an inconsistency. Then, halfway through the larger job, it forgets a constraint, repeats completed work, or combines two findings that do not belong together.
The pieces work. The whole thing does not.
I pulled the full transcript of this conversation with MIT researcher Alex Zhang using ytmd, my local tool for saving and reading timestamped YouTube captions. His work on recursive language models, or RLMs, starts with a useful question: how much of what looks like a model limitation is actually a failure to organize the work?
These are the ideas I am keeping, and what I think they mean for people building AI products. The experimental results are reported in the video, not independently verified here. The implementation examples are mine.
The model can do the step. The system cannot do the job.
We tend to treat an agent failure as a request for more intelligence. A better model. A longer prompt. More examples. More context.
Sometimes that is exactly what is missing. But sometimes the model already knows how to do the individual steps. It can extract a fact, compare two claims, classify a passage, and write the code. What breaks is the connection between those abilities.
Zhang calls his broader view the mismanaged geniuses hypothesis: models have substantial capabilities that we are not managing well enough to turn into reliable work. It is a hypothesis, not proof that orchestration can fix everything. But it is a much better starting point than assuming every failure needs another model release. (39:56)
If your model succeeds on the parts and fails on the whole, investigate the workflow before replacing the model.
Context is working space, not the whole company
The familiar agent loop is simple. Give it a task, let it call tools, append the results, and keep going. Eventually the conversation fills up. Compact it. Continue.
Now the same context holds instructions, source documents, failed attempts, old plans, intermediate results, and decisions. The model has to find the current task inside the history of doing the task.
Compaction helps, but it creates another problem: you have to decide what matters before you know everything you will need later. A detail that looks disposable now may be the thing that explains a contradiction ten steps from now. Zhang argues that this pattern gives weak guarantees about preserving important information. (31:32)
A larger context window buys room. It does not decide what belongs in that room.
My takeaway: keep source data, durable decisions, and execution state outside the conversation where possible. Bring in what the current step needs. Keep a way back to the originals.
The context window should not have to be the database, the execution log, and the institutional memory at the same time.
What an RLM actually changes
An RLM has two core ideas.
Put the context in an environment. A large input lives in a programmable workspace, such as a Python REPL, rather than being pasted into the prompt in full. The model writes code to inspect it, filter it, split it, and retrieve relevant pieces.
Let the model call models. It can programmatically delegate a bounded problem to another model call, with the context that call needs. The same pattern can be applied recursively. (5:01)
Think about how you work with a large dataset. You do not memorize every row. You inspect its shape, write a query, check a sample, and narrow the question. The data remains available even though it is not all in your head.
That is the shift: context becomes something the model can operate on, not just something it must consume.
The interesting part is not giving five agents different personalities. It is giving each call a clear job and the information required to do it.
Keep the local problems familiar
Zhang describes the property he wants as locally in-distribution.
Roughly: the overall task can be unfamiliar, long, or complicated, while each individual model call handles a problem resembling something it already knows how to solve. The system gets somewhere new by composing familiar operations. (28:38)
That is more useful than simply saying the prompts should be shorter.
A short prompt can still be impossible. It might omit the relevant contract clause, hide a dependency, or ask for an ambiguous judgment. And we generally cannot inspect a frontier model’s training data to prove a task is in-distribution.
The practical test is whether the model reliably solves the bounded problem you give it, and whether the system combines those answers correctly.
Decomposition moves the hard parts. It does not make them disappear. Someone still has to choose the boundaries, preserve dependencies, and check the result.
Learn the procedure, then scale the work
The video discusses two forms of generalization: training on short tasks and evaluating on longer ones; training in one domain and evaluating in another with a similar structure.
The host reports transfer to tasks 8–32 times longer than the training tasks. The experimental walkthrough describes stronger generalization in the RLM setup than in a raw-transformer baseline. At the time of the conversation, Zhang was still developing the compositional-generalization write-up. (2:14, 10:11)
The interesting possibility is that the model learns a way to organize work that survives changes in scale and content.
Inspect the data. Select relevant pieces. Delegate analysis. Aggregate results. Check exceptions. That procedure may remain useful even when the input changes.
This is not software beating model weights. The model still has to learn how to use the environment. It is model plus runtime, trained and evaluated as a system.
For builders, the question becomes: are we teaching a reusable way of working, or tuning a workflow until it passes one particular test?
A concrete example: did the release break checkout?
Suppose an agent has to work out whether a product release caused an increase in checkout failures. It has support tickets, release notes, and logs.
The easy implementation is to feed it a large pile of evidence and ask for a conclusion. A more deliberate one separates the work:
- Inspect. Check dates, versions, missing fields, duplicates, and identifiers.
- Select. Build relevant before-and-after cohorts. Keep uncertain cases visible.
- Interpret. Use bounded model calls to classify ticket excerpts, returning source references with each finding.
- Calculate. Use code to count categories and compare periods.
- Challenge. Check whether several tickets describe one incident, reporting volume changed, or the failure existed before the release.
- Conclude. Give the final call verified aggregates, representative evidence, and unresolved questions.
This is my example, not a workflow prescribed in the interview. Before-and-after counts can point to a regression; they do not prove the release caused it. Compare unaffected versions or try to reproduce the failure where possible.
The point is not more calls. It is putting each operation where it belongs. Code does exact counting. Models interpret language. Source records remain available.
But there is a trap. A worker looking at one ticket cannot necessarily see that it belongs to a larger incident. A clean-looking aggregate can hide a collection of bad classifications.
Split the work around the problem’s dependencies, not every few thousand tokens. Otherwise you just distribute the confusion.
Code makes the work explicit
Zhang strongly favors code as the main interface for agents. Code can express loops, conditions, transformations, and multiple tool or model calls in one program. (87:02)
For a builder, that means less reliance on the model remembering to do the same thing the same way. Counting, output validation, budget enforcement, and state transitions can live in executable logic.
It also opens up optimization. If two operations are independent, the runtime can potentially run them together.
Zhang’s speculative programmatic tool calling explores starting eligible calls while the model is still generating the rest of a program. He also describes a shadow execution environment for identifying calls ahead of normal execution. The implementation discussed does not use a smaller draft model, and he explicitly cautions against speculating over operations with unsafe side effects. (83:34)
It is an early systems idea, not a production recipe. My order of operations would be correctness first, ordinary parallelism second, speculation only after measuring where the time goes.
Faster is not automatically cheaper. Work you discard still costs something.
Evaluate the harness, not just the model
A harness is the software around the model: tools, context management, the execution loop, retries, and stopping rules.
Zhang argues that harness design is still too unscientific. Every benchmark has its own setup, and small changes to that setup can substantially change what the model appears capable of doing. (25:01)
For a product team, this means holding the model fixed and testing the workflow changes separately.
Start with your current agent. Move source data outside the prompt. Then add bounded delegation. Then parallelize independent work. Test on cases you did not tune against, with the same model version, tool access, and acceptance checks. Compare under similar budgets, or make the extra compute explicit.
Measure missing evidence, lost constraints, repeated work, recovery after failures, total cost, and latency. Increase the task size. Try a different dataset with the same underlying structure. A workflow that only works on the examples you tuned it against has not demonstrated much transfer.
Keep the more complicated version only if the reliability gain is worth its cost and latency. More calls are not a win by themselves.
Keep acceptance checks outside the agent’s control. Zhang is skeptical that asking a model to invent its own scaffolding will reliably produce good designs rather than exploit the evaluation setup. (88:56)
An agent saying it finished is not evidence that it finished correctly.
Prototype the idea before changing the architecture
The conversation goes beyond application code. Zhang sees harnesses and model architecture as more closely related than we usually assume. Some behaviors implemented outside the model today might eventually move inside it.
His practical suggestion is to test the behavior in a harness first. If you think a new retrieval or reasoning mechanism will help, use an existing model to build the simplest version and find out. It is a cheaper way to test the assumption before committing to an architectural change. (67:34)
That does not prove a neural implementation will work or fail the same way. But it forces a useful question: have we shown that the behavior we want is valuable, or are we just attracted to a more complicated implementation?
What I am keeping
- If the parts work and the whole fails, inspect the organization of the work.
- Keep durable data outside the conversation. Preserve a route back to the evidence.
- Give each call a bounded problem, not just a smaller prompt.
- Decompose around dependencies. Local correctness is not global correctness.
- Use code for exact operations and enforceable rules.
- Evaluate workflow changes with the same discipline as model changes.
- Test one recurring failure before rebuilding the entire agent.
Zhang is careful about the limits. Better composition does not establish that existing models can produce every scientific breakthrough. Nor do recursive calls, on their own, make a persistent agent reliable. (41:18, 70:44)
His closing point is not that everyone should use one best harness. It is that we should understand which properties let a system carry useful capabilities into new tasks. (93:23)
That is what makes this worth paying attention to. Better models expand what is possible. Better systems determine how much of that possibility survives contact with an actual job.
Before asking the agent to remember more, give it a better way to work.
Transcript pulled with ytmd. I read the full available automatic-caption transcript; captions can be wrong, and the figures and underlying papers were not independently audited. Full episode: RLMs are Compositional Generalizers with Alex Zhang from MIT.