ReAct and the planning family
Once you have a loop and tools, the obvious next question is whether the model should plan before it acts — and for three years the literature answered with a family of algorithms: ReAct, Plan-and-Execute, Reflexion, Tree of Thoughts, LATS. Almost none of them survives in production by name. What survived are their ideas, folded into a single loop with better context engineering. This part takes each algorithm on its own terms, shows where its shape pays off and where it only adds latency, and then maps every name onto the production mechanism that inherited its contribution.
ReAct: reasons and acts in one trajectory
Interleave thought with action, and let the transcript carry the state
ReAct (Yao et al., 2022, arXiv:2210.03629) is the algorithm the previous part quietly implemented. Its claim was that reasoning and acting should not be separated: the model emits a thought, then an action, then reads the observation, then emits the next thought. The thought conditions the action; the observation conditions the next thought. The model never plans the whole path — it navigates, one grounded step at a time, and the transcript is the state.
Plan-and-Execute splits the two (prompt-only form: Plan-and-Solve, Wang et al., 2023, arXiv:2305.04091). A planner call produces a list of subtasks up front; an executor then works through them one at a time, each with a smaller context because it only needs to know its own subtask and the plan. The argument for it is straightforward: a small executor context is cheaper and less prone to drift, and the plan is an explicit artefact you can inspect.
The trade is where the interesting part lives. When a task is long and genuinely decomposable, planning pays: it stops the loop from re-deriving the same strategy every step, and it keeps each executor call small. When a task is short or not really decomposable, the planner is pure overhead — one extra call, one extra artefact to keep in context, and a plan that the executor ignores or contradicts. Drag the controls and watch the two traces on the same task.
Two traces on one task. Below them, the model calls and the context cost projected by AppSim.agentRun for each run length.
Reflexion: retry with a critique
Verbal self-feedback across attempts
Reflexion (Shinn et al., 2023, arXiv:2303.11366) adds memory of failure. After an attempt, the model is asked to reflect — in words — on what went wrong, and that reflection is carried into the next attempt as context. The mechanism is not a gradient and not a new policy; it is a second piece of text that changes the distribution of the retry. Across attempts, the success rate rises.
It rises with diminishing returns, and that shape is the whole lesson. A good critique removes a class of errors the model can name; once those are gone, further reflections mostly restate the same advice, and the curve flattens. Compare it to a plain retry with no critique: the per-attempt probability never moves, and only the arithmetic of repeated tries helps you.
Set the critique quality and the number of attempts below. The solid line is reflexion, the dashed line is retrying the same thing, and the marker is the attempt budget you chose.
Cumulative success after k attempts. Reflexion lifts the per-attempt probability; plain retries only compound the original one.
Tree of Thoughts and LATS: search with a budget
Branch, estimate, and pay for exploration
Tree of Thoughts (Yao et al., 2023, arXiv:2305.10601) treats reasoning as search. Instead of one chain, the model proposes several candidate “thoughts” at each step, a value estimate scores them, and the search expands the most promising branches. LATS (Zhou et al., 2023, arXiv:2310.04406) makes that search explicit: Monte-Carlo-tree-search style exploration with a value function, backed by real tool observations, under a node budget.
Both buy accuracy with search budget, and both are fighting the same exponential: the number of leaves grows as bdepth. A branching factor of four at depth three is sixty-four leaves; one more level of depth makes it two hundred and fifty-six. The search is only worth it when a value estimate is cheap and informative enough that you rarely expand a bad branch — which is exactly why the useful residue in production is not “run MCTS” but “generate a few candidates and pick with a cheap check, under a step budget”.
Left: one search tree, with expanded leaves filled. Right: expected accuracy against the number of leaves evaluated.
What actually survived
The ideas outlived the algorithms
None of these papers ships as-written in a production agent, and that is not a failure of the papers — it is what happens when a good idea is absorbed. The named algorithm is a research artefact with a fixed control structure; the production mechanism is a single loop whose context carries the same information. Explicit state, decomposition, retry-with-critique and bounded search are all now things you engineer into that loop rather than algorithms you install.
Click each algorithm to see the mechanism that inherited its idea.
Each research algorithm mapped to the production mechanism that carries its contribution today. The highlighted row is the one you selected.
The same story is told from the other side by Volume 1: the mechanism that makes a single loop carry all this is context engineering, and the reason a chain of thought helps a multi-step task at all is the subject of the reasoning part. This part is what to do with that reasoning once the task needs tools and retries.
Cheat sheet
| Question | The answer that survives production |
|---|---|
| What is ReAct? | Interleaved reasoning and acting: thought → action → observation → thought, in one transcript. |
| What does Plan-and-Execute add? | A planner call and an explicit subtask list, so each executor call has a smaller context. |
| When does planning pay? | Long, genuinely decomposable tasks. On short tasks the plan is an extra call that buys nothing. |
| What is Reflexion? | A verbal self-critique after a failed attempt, carried into the retry as context. |
| Why do reflexion gains flatten? | Once the nameable errors are gone, further critiques restate the same advice — diminishing returns. |
| What is Tree of Thoughts? | Reasoning as search: propose several candidate thoughts, score them, expand the best. |
| What does LATS add? | MCTS-style exploration with value estimates and real tool observations, under a node budget. |
| Why is search expensive? | Leaves grow as bdepth. The harness has to bound the budget; the tree will not. |
| What survived to production? | The ideas: explicit state, decomposition, retry-with-critique, bounded best-of-n — folded into one loop. |
Further reading
- Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, arXiv:2210.03629, 2022 — the interleaved reason/act loop.
- Shinn et al., “Reflexion: Language Agents with Verbal Reinforcement Learning”, arXiv:2303.11366, NeurIPS 2023 — verbal self-reflection across attempts.
- Yao et al., “Tree of Thoughts: Deliberate Problem Solving with Large Language Models”, arXiv:2305.10601, NeurIPS 2023 — reasoning as search over candidate thoughts.
- Zhou et al., “Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models”, arXiv:2310.04406, ICML 2024 — MCTS-style exploration with value estimates and tool observations.
- Wang et al., “Plan-and-Solve Prompting”, arXiv:2305.04091, ACL 2023 — the planner/executor split in its prompt-only form.
- Anthropic, “Building effective agents”, December 2024 — why the evaluator-optimiser and the simple loop absorb most of what these algorithms were reaching for.