Sea-FillOpen-Source Research Community University of LuxembourgPCOGUniversity of Luxembourg · SnT

Reward-Free Policy Optimization

Unlocking the Critic

Recent approaches to RL post-training for large language models increasingly remove the critic to reduce training instability and memory overhead. Even where a critic is trained, it is discarded once training ends, although it has learned to predict outcomes. We revisit this trend and show that a pretrained critic's ability to predict future outcomes can make it a valuable asset for efficient long-horizon reasoning.

The frozen critic's value at the last token

Share of rollouts per value bin, correct above the axis, incorrect below.

correctincorrect
47.6%mean validation over steps 10–300, against 46.6% for supervised PPO at the same cap.
~50%of RFPO's training rollouts are unfinished at every step and still receive a reward.
264 hGPU-hours for 300 steps, against 327 for PPO (−19%). No critic update, no verifier calls.
01 · Unlock the critic

The critic is not the problem

Instability in critic-based RL for long chain-of-thought reasoning is largely an optimization artifact. Keeping policy updates small and low in variance restores stable convergence, and the critic becomes what it was trained to be: a predictor of eventual success.

Contributions (i) and (ii) ↓
02 · Reward-free

No verifier, no labels in the loop

A single calibrated, frozen critic serves as the reward, the GAE baseline and the forecaster. Training needs no completed rollouts, no step-level annotations and no external reward labels, and the binarized critic matches supervised PPO.

Contributions (ii) and (iii) ↓
03 · Don't wait for the end

Reward a rollout before it finishes

The critic scores unfinished rollouts with accuracy comparable to complete ones, so rollouts can be rewarded before they finish. This suits long-horizon reasoning, where outcomes arrive late and generation dominates cost.

Contribution (iv) ↓

Unlock the critic · contribution (i)

Critic instability is an optimization artifact

Recent approaches to RL post-training increasingly remove the critic to reduce training instability and memory overhead. We find that this instability is largely an optimization artifact: it comes from the update recipe, not from the value network, and it recedes once each policy update is kept small and low in variance. Under this recipe, critic-based training converges stably, without length collapse or entropy collapse.

Unlock the critic, reward-free · contribution (ii)

One frozen critic, three roles

A well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. Its predictions therefore provide outcome-derived, dense, per-prefix learning signals that need neither completed rollouts, step-level annotations, nor external reward labels. RFPO repurposes a single calibrated, frozen critic as the rollout-level reward, as the value baseline for generalized advantage estimation, and as a success forecaster for unfinished prefixes. The RL loop contains no verifier and trains no value network.

If a critic is good enough to be trusted as a baseline, it is also good enough to be the reward.

ROLLOUTS (tokens →) generation cap frozen critic V(x, y≤t) calibrated once reward1[v − b(ℓ) > τ], last token baselineV at every token, for GAE forecastscores rollouts cut off illustrative scores: 0.93 → reward 1 · 0.08 → reward 0 · truncated rollout scored at its last visible token
RFPO in one picture. No value network is trained and no verifier runs once the critic is frozen. The raw score has a length bias, so we subtract a length baseline b(ℓ) and binarize at a threshold τ; a continuous score lets the policy exploit the bias.

The value histogram at the top of this page is why this works: the frozen critic pushes incorrect rollouts toward 0 and correct ones toward 1 (AUC 0.96 on rollouts from the converged policy). It ranks attempts at the same problem just as well when the attempt never finished, so the reward is not simply detecting truncation. From a prefix alone, within-problem AUC reaches 0.84 at 4,096 tokens, and 0.84–0.92 on prefixes that have not produced an answer yet.

Ranking attempts at the same problem

Within-problem AUC of the frozen critic, averaged over ten policy checkpoints (16,384 rollouts each, 19–54% truncated).

Truncated attempts are ranked as well as the full set. Finished-only is lower because a problem that mixes finished and truncated attempts offers easy contrasts, and removing the truncated attempts removes them.

Ranking from a prefix

Within-problem AUC when the critic sees only the first n tokens.

all prefixesno answer yetshare finished
Hover to read a prefix length.

Reward-free, safely · contribution (iii)

Calibration guards against length bias

The critic's raw score carries a length bias, and a continuous reward lets the policy exploit it in either direction. Binarizing the debiased score closes that channel. Used as a continuous reward, the same critic first climbs above fully supervised PPO, which shows how much it knows, and then gives the gain back as response length drifts. Binarizing trades that ceiling for stable training.

Response length drifts under a continuous reward

Change in mean training response length, last step against first, by debiasing strength. The shaded band is the range the binarized runs stay within.

Binarizing trades the ceiling for stability. With no debiasing the policy shortens its answers by 35% in 87 steps; with full debiasing it lengthens them by 34% and ends up truncating almost every rollout.

Don't wait for the end · contribution (iv)

Parity at lower cost, built for long horizons

Outcome rewards wait for the last token, so the longer the reasoning, the more training pays for waiting. A critic answers earlier.

Binarized, RFPO matches supervised PPO without a single label in the training loop, while substantially cutting compute and memory overhead. Because rollouts can be rewarded before they finish, training no longer has to pay for waiting on every trajectory to complete. Even when a tight generation cap leaves a large share of every batch unfinished, RFPO validates above supervised PPO trained under the same cap, at a lower cost.

Validation under a 4,096-token training cap

Macro-average of AIME 2026 and AMC 2023 at 12,288 tokens. Dots: single evaluations; lines: centered means of five.

RFPO, 0% labelsPPO, 100% labels
Hover the chart to read a step.

Share of training rollouts truncated

PPO shortens its answers to fit the cap; RFPO keeps about half unfinished.

RFPOPPO
Hover to read a step.

Median seconds per RL step

Same node, same cap. The gap is PPO's critic update.

At the standard 5,120-token cap

Separate 16-sample evaluation after 300 RL steps, pass@1 (%)

ModelAIME25AIME26AMC23GPQAAvg
Initial policy22.721.071.444.940.0
PPO, 100% labels25.221.773.646.641.8
RFPO, 50% labels22.721.272.848.541.3
RFPO, 0% labels25.421.770.047.341.1

Level or ahead on AIME and GPQA, behind on AMC 2023 (40 problems, 2.5 points each). An RL step takes 421 s, against 582 s for PPO.

As reasoning traces lengthen and agentic episodes stretch, waiting for the outcome buys less and less for what it costs, and the value network is the one component of the standard recipe that speaks before the outcome arrives. So far our evidence comes from mathematical reasoning with one 4B model family and traces up to 12,288 tokens; larger models and other tasks are next.

The question is not whether a critic is affordable, but how much of what it already knows the current recipe throws away.

Our findings challenge the prevailing critic-free paradigm and establish critic-based, reward-free optimization as a scalable and computationally efficient path for LLM post-training.

Authors

Hongyang (Kevin) Li, Xiao Li, Caesar Wu, Said Mammar, Grégoire Danoy, Pascal Bouvry University of Luxembourg · Sea-Fill Open-Source Community · Université Paris-Saclay

Cite

@article{li2026unlocking,
  title   = {Unlocking the Critic: Reward-Free Policy
             Optimization for LLM Post-Training},
  author  = {Li, Hongyang and Li, Xiao and Wu, Caesar and
             Mammar, Said and Danoy, Gr{\'e}goire and Bouvry, Pascal},
  journal = {arXiv preprint arXiv:2609.37119},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.37119}
}