Papers
arxiv:2610.08077

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

Published on Oct 6
ยท Submitted by
Haoxiang Zhang
on Oct 8
Authors:
,
,
,
,
,
,
,
,
,
,

Abstract

Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37--98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.

Community

Paper author Paper submitter
โ€ข
edited about 2 hours ago

Hi everyone! ๐Ÿ‘‹ Author here, happy to answer any questions.

TL;DR: RLVR learns from what happened after the agent acted. We ask whether the same experience can teach the agent what it should have anticipated before acting.

The problem. Group-relative RL (GRPO and friends) gives zero gradient whenever every rollout in a group gets the same reward. In our runs, that is 37โ€“98% of groups depending on model scale. Weak models fail everything and strong models solve everything, so the trajectories get thrown away, even though they show exactly how the agent failed or what it needed to know.

image

The idea: prospective learning. Hindsight supervises foresight.

  • Before interacting, the policy predicts the pitfalls it is likely to hit.
  • After the rollout, a stop-gradient copy of the same policy sees the completed trajectory and scores that prediction.
  • SRD (Self-Retrospection Distillation) aligns the two token-level distributions.

SRD is a lightweight auxiliary loss that plugs into GRPO, OPSD, or RLSD. Foresight is only a training target and is never needed at inference.

image

Results (Qwen3.5-4B/9B, 10 benchmarks across math, code, search, ALFWorld and WebShop):

  • ๐Ÿ“ˆ Consistent gains across GRPO, OPSD, and RLSD, especially on recovering the scaffold tax where agents fail more in tool calls than pure reasoning.
  • ๐Ÿ”ฅ At 2B with 98% all-fail groups, GRPO stays at 0.0%, while GRPO+SRD reaches 60.6% training success on the same rollout budget.
  • ๐Ÿ› ๏ธ SRD stabilizes pure self-distillation. 9B OPSD falls below the untrained model on HotpotQA, 2Wiki and LCB-v6; adding SRD lifts all three back above it.
  • ๐ŸŒ Gains hold under held-out formats and much longer horizons (BrowseComp-Plus, 10ร—+ tool calls).

If you like concrete examples, Appendix F walks through raw training groups.

๐Ÿ“„ Paper: https://arxiv.org/pdf/2610.08077
๐Ÿ’ป Code: [TB Released]

We'd love to hear your thoughts, especially on other kinds of foresight worth distilling beyond pitfalls and knowledge!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08077 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.08077 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.