PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

arXiv:2609.40285

Yinghui He1,2*† Yapei Chang2,3 Khushi Bhardwaj2 Daniele Molinari2 Tugrul Konuk2 Jan Kautz2 Ali Hatamizadeh2†

1Princeton University 2NVIDIA 3University of Maryland

*Work done during an internship at NVIDIA.   †Corresponding authors.

PivotOPD is an on-policy distillation framework for multi-turn language agents. It finds the single early mistake that most failed rollouts hinge on, steers the student away from it, and, when the mistake happens anyway, teaches the student to recover from the state it created. The result is the best average against 13 baselines on ALFWorld, WebShop and Search-based QA, with gains that carry over to SWE-Bench Verified.

A student rollout: correct turns, one pivotal mistake, then recovery turns until the task completes.

The idea in one line: prevent the pivotal mistake, and learn to recover when it happens anyway.

Abstract

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In our preliminary experiments across three Qwen3 models (8B–235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%.

Failed rollouts trace back to one early, recoverable mistake

Three panels: pivotal turns arrive early in rollouts of Qwen3-8B, 30B and 235B; correcting the pivotal turn or the two turns after it raises replayed success to 59% and 58%; on-policy distillation reduces overall failures but not failures after a pivotal turn.
Pivotal turns arrive early, decide the outcome, and survive standard OPD. Using ALFWorld's symbolic oracle, 59% of the failed rollouts of three Qwen3 models contain a pivotal mistake, an action that lengthens the remaining optimal trajectory or makes the task unsolvable. (a) The first one lands at a median of turn 8 to 12 out of 30, and the agent wastes the rest of the episode. (b) Correcting the pivotal turn lifts replayed success from 8% to 59%; guiding only the two turns after it still reaches 58%. (c) Standard OPD lowers the overall failure rate from 79% to 56%, but failures after a pivotal turn only fall from 51% to 49%: the recovery action stays below 1% probability, so a group of eight rollouts never samples it and there is no learning signal to fix it.

PivotOPD

PivotOPD adds three components to group-based RL. A privileged self-teacher, the frozen student conditioned on a hint that names an action, turns every action the teacher names into a token-level target written in the student's own reasoning style, so the teacher only ever names actions. Both distillation terms enter a single PPO update as per-token distillation advantages.

01

Pivot detection

A teacher model reads each rollout in hindsight, selects candidate turns, and names a gold action at each. A turn is pivotal when the student's committed action disagrees with it. After each pivotal turn, the teacher names a recovery action at each of the next K turns.

02

Preventive distillation (reverse KL)

The student's own recorded response is re-scored against the self-teacher hinted with the gold action, which moves the student away from the committed mistake.

03

Recovery distillation (forward KL)

The self-teacher hinted with the recovery action writes recovery responses, and the unhinted student is trained on them. The mass-covering forward KL places probability on recovery actions the student almost never produces on its own.

Overview of PivotOPD: a teacher performs pivot detection on student rollouts, generates a gold action at the pivotal turn and recovery actions afterward; preventive distillation uses reverse KL with a privileged self-teacher, recovery distillation uses forward KL; both combine with RL in the PivotOPD objective.
Overview of PivotOPD. Pivot detection names the gold and recovery actions, a privileged self-teacher turns them into token-level targets, and preventive (reverse KL) and recovery (forward KL) distillation combine with group-based RL in one PPO update. Later recovery turns start from the state reached by executing the recovery action in a copy of the environment that replays all preceding actions.

Results

8 / 8
per-benchmark averages won against 13 baselines, for both Qwen3-1.7B and Qwen3-8B students
+5.5%
over the strongest baseline on ALFWorld with the 1.7B student (+5.9% on Search-based QA)
72.7%
of replayed pivotal mistakes recovered, against 8.3% for the base model and 20.3% for standard OPD
+3.2%
SWE-Bench Verified resolve rate for a Nemotron-3.5 student, where standard OPD adds +0.2%

ALFWorld, WebShop and Search-based QA

Against 13 baselines spanning RL and self-distillation, turn-level distillation for multi-turn agents, and guidance from skills or pivotal turns, PivotOPD attains the best average on ALFWorld, Search-based QA, and WebShop for both Qwen3-1.7B and Qwen3-8B students, ranking first on all eight per-benchmark averages over three seeds. With the 1.7B student it improves over the strongest baseline by +5.5% on ALFWorld and +5.9% on Search-based QA, with the largest gains on the task types where successful rollouts are scarcest and outcome-only RL receives no signal. It also turns partial progress into completed tasks: on WebShop it beats RLSD by only +1.2% in score but by +14.1% in success rate.

Table: success rates on six ALFWorld task types, exact-match accuracy on seven QA datasets, and WebShop score and success rate for 14 methods with Qwen3-1.7B and Qwen3-8B students. PivotOPD holds the best average in every column group for both students.
Main results. Success rate per ALFWorld task type, exact-match accuracy per QA dataset, and WebShop score and success rate, averaged over three seeds. Green marks the best result in a column and lilac the second-best. Every method with a teacher uses Qwen3-30B-A3B (1.7B student) or Qwen3.5-122B-A10B (8B student).

No stronger teacher required

With Qwen3-8B serving as its own teacher, and every baseline that uses a teacher given the same treatment, PivotOPD remains best on all three benchmarks, ahead of the strongest baseline on each by at least +1.5% and by +3.9% on average. Much of the gain comes from where and how the teacher intervenes rather than from teacher capacity alone.

Bar charts of self-distillation results on Qwen3-8B for ALFWorld, WebShop and Search-based QA; PivotOPD is best on all three.
Self-distillation results on Qwen3-8B, where the student serves as its own teacher.

Transfer to software engineering

Trained on a curated bug-fix curriculum with Nemotron-3-Super as the teacher, a Nemotron-3.5-SFT student gains +3.2% resolve rate on SWE-Bench Verified (62.8% to 66.0%), closing roughly a third of the gap to the teacher, whereas standard on-policy distillation from the same teacher improves it by only +0.2%. Because SWE-Bench episodes are long and containerized, this experiment audits the final committed action and exercises preventive distillation alone.

Bar chart of SWE-Bench Verified resolve rate: Nemotron-3.5-SFT student 62.8, standard OPD 63.0, PivotOPD 66.0, Nemotron-3-Super teacher 73.0.
SWE-Bench Verified. Standard OPD moves the student by +0.2 points; PivotOPD moves it by +3.2.

PivotOPD learns to recover

Replayed from the same 72 oracle-labeled pivotal mistakes, PivotOPD recovers in 72.7% of replays, against 8.3% for the base model, 20.3% for standard OPD and 45.8% for the preventive-only variant, and it does so in the fewest turns. It improves the recovery rate on 60 of the 72 mistakes and worsens none.

Case study: after taking a tomato instead of the egg, the base model never finds the egg and fails, while the PivotOPD model sets the tomato aside, finds the egg and completes the task. Right: PivotOPD recovers from 72.7% of replayed mistakes and in the fewest turns.
Recovery after the same pivotal mistake. (a) The base model commits the tomato to the microwave and never finds the egg; PivotOPD sets the tomato aside, finds the egg, and completes the task. (b) Over all 72 replayed mistakes, PivotOPD recovers most often and in the fewest turns.

How many recovery turns to train

The selected recovery budget raises the best validation score over the preventive-only variant (K = 0) by 5.4 points on ALFWorld, 21.5 on WebShop and 2.4 on Search-based QA. One recovery turn is enough on WebShop and Search-based QA, while ALFWorld, whose tasks chain longer sequences of fine-grained actions after a mistake, needs two.

Validation curves for recovery budgets K of 0, 1, 2 and 3 on ALFWorld, WebShop and Search-based QA with the Qwen3-1.7B student.
Ablation on the recovery budget with the Qwen3-1.7B student. Validation performance during training for K in {0, 1, 2, 3}, where K = 0 is the preventive-only variant.

BibTeX

@article{he2026pivotopd,
  title   = {PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents},
  author  = {He, Yinghui and Chang, Yapei and Bhardwaj, Khushi and Molinari, Daniele and Konuk, Tugrul and Kautz, Jan and Hatamizadeh, Ali},
  journal = {arXiv preprint arXiv:2609.40285},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.40285}
}

Return to top