Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

ReAct: reasons and acts in one trajectory

Interleave thought with action, and let the transcript carry the state

ReAct (Yao et al., 2022, arXiv:2210.03629) is the algorithm the previous part quietly implemented. Its claim was that reasoning and acting should not be separated: the model emits a thought, then an action, then reads the observation, then emits the next thought. The thought conditions the action; the observation conditions the next thought. The model never plans the whole path — it navigates, one grounded step at a time, and the transcript is the state.

Plan-and-Execute splits the two (prompt-only form: Plan-and-Solve, Wang et al., 2023, arXiv:2305.04091). A planner call produces a list of subtasks up front; an executor then works through them one at a time, each with a smaller context because it only needs to know its own subtask and the plan. The argument for it is straightforward: a small executor context is cheaper and less prone to drift, and the plan is an explicit artefact you can inspect.

The trade is where the interesting part lives. When a task is long and genuinely decomposable, planning pays: it stops the loop from re-deriving the same strategy every step, and it keeps each executor call small. When a task is short or not really decomposable, the planner is pure overhead — one extra call, one extra artefact to keep in context, and a plan that the executor ignores or contradicts. Drag the controls and watch the two traces on the same task.

Two traces on one task. Below them, the model calls and the context cost projected by AppSim.agentRun for each run length.

💡 The lesson the papers already agree on: planning is a bet that the task decomposes cleanly. Take the bet when the task is long and modular; skip it when it is short, and do not keep a prescriptive planner/executor split that costs a call and buys nothing.
2

Reflexion: retry with a critique

Verbal self-feedback across attempts

Reflexion (Shinn et al., 2023, arXiv:2303.11366) adds memory of failure. After an attempt, the model is asked to reflect — in words — on what went wrong, and that reflection is carried into the next attempt as context. The mechanism is not a gradient and not a new policy; it is a second piece of text that changes the distribution of the retry. Across attempts, the success rate rises.

It rises with diminishing returns, and that shape is the whole lesson. A good critique removes a class of errors the model can name; once those are gone, further reflections mostly restate the same advice, and the curve flattens. Compare it to a plain retry with no critique: the per-attempt probability never moves, and only the arithmetic of repeated tries helps you.

Set the critique quality and the number of attempts below. The solid line is reflexion, the dashed line is retrying the same thing, and the marker is the attempt budget you chose.

Cumulative success after k attempts. Reflexion lifts the per-attempt probability; plain retries only compound the original one.

⚠️ The trap: a critique that only restates the failure is not feedback, it is a longer prompt. Reflexion buys you a better retry only when the reflection names a specific fix the next attempt can act on — otherwise you are paying tokens for a flatter curve.
4

What actually survived

The ideas outlived the algorithms

None of these papers ships as-written in a production agent, and that is not a failure of the papers — it is what happens when a good idea is absorbed. The named algorithm is a research artefact with a fixed control structure; the production mechanism is a single loop whose context carries the same information. Explicit state, decomposition, retry-with-critique and bounded search are all now things you engineer into that loop rather than algorithms you install.

Click each algorithm to see the mechanism that inherited its idea.

Each research algorithm mapped to the production mechanism that carries its contribution today. The highlighted row is the one you selected.

The same story is told from the other side by Volume 1: the mechanism that makes a single loop carry all this is context engineering, and the reason a chain of thought helps a multi-step task at all is the subject of the reasoning part. This part is what to do with that reasoning once the task needs tools and retries.

Cheat sheet

QuestionThe answer that survives production
What is ReAct?Interleaved reasoning and acting: thought → action → observation → thought, in one transcript.
What does Plan-and-Execute add?A planner call and an explicit subtask list, so each executor call has a smaller context.
When does planning pay?Long, genuinely decomposable tasks. On short tasks the plan is an extra call that buys nothing.
What is Reflexion?A verbal self-critique after a failed attempt, carried into the retry as context.
Why do reflexion gains flatten?Once the nameable errors are gone, further critiques restate the same advice — diminishing returns.
What is Tree of Thoughts?Reasoning as search: propose several candidate thoughts, score them, expand the best.
What does LATS add?MCTS-style exploration with value estimates and real tool observations, under a node budget.
Why is search expensive?Leaves grow as bdepth. The harness has to bound the budget; the tree will not.
What survived to production?The ideas: explicit state, decomposition, retry-with-critique, bounded best-of-n — folded into one loop.

Further reading

5

Check your understanding

0/4 answered