Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecessary exploration that wastes computation and can even overturn the correct answer. We frame this behavior as a latent productive-to-redundant transition and show it is directly reflected in hidden states: around first-correct-solution (FCS) boundaries, late-layer representations separate efficient from overthinking tokens, while boundary-permutation and position controls collapse. We propose ROM, a streaming intervention framework that monitors a frozen LRM with a lightweight hidden-state detector (~0.1% of backbone parameters) and intervenes at well-formed reasoning boundaries; Counterfactual Self-Correction (CSC) balances supervision with wrong→correct trajectories, preserving useful pre-FCS self-correction. Unlike prior adaptive early-exit methods, ROM extracts no intermediate answers, launches no probe decoding, and updates no backbone weights. Across five backbones from three model families and five reasoning benchmarks, against ten recent baselines under a shared protocol, ROMCSC attains the highest accuracy in 19 of 25 model–benchmark settings, cuts response length by 28–77% (mean 45%) versus vanilla decoding, and is the only method on the accuracy–length Pareto front in every setting. The same MATH500-trained supervision transfers zero-shot across scales, families, and task domains, and end-to-end wall-clock latency drops by 46.5% with ~5% per-token overhead.
Existing methods approximate the productive→redundant transition through length-RL objectives, decoding statistics, or answer arrival. ROM directly supervises the boundary itself from token-level FCS labels, while keeping the backbone frozen and operating at every token. Extraction free: no online intermediate-answer extraction or probe/trial decoding. Transition supervised: the stopping signal is trained on the productive-to-redundant boundary itself.
| Method family | Frozen backbone | Token level | Extraction free | Transition supervised |
|---|---|---|---|---|
| Length retraining: L1, O1-Pruner | ✗ | ✗ | ✓ | ✗ |
| Answer-defined triggers: Dynasor, DEER, PUMA-RD, PMA, RP | ✓ | ✗ | ✗ | ✗ |
| Surface heuristics (step level): EAT, REFRAIN | ✓ | ✗ | ✓ | ✗ |
| Surface heuristics (token level): SyncThink, RCPD, RPDM, NEAT | ✓ | ✓ | ✓ | ✗ |
| Answer-arrival head: TERMINATOR | ✓ | ✓ | ✓ | ✗ |
| Latent-boundary head: ROM (ours) | ✓ | ✓ | ✓ | ✓ |
We align 61 MATH500 traces at each overthinking boundary, take the last 20 efficient and first 20 overthinking tokens, and probe Qwen3-8B's late-layer hidden states with logistic regression under response-level group cross-validation. The two phases are linearly separable at 85.9% accuracy / AUROC 0.928 despite often discussing the same content; probe scores rise sharply at the true boundary but stay flat under permuted alignment (which collapses to 48.5%, below chance); and the signal is not a position shortcut — a position-only classifier reaches 50.3%, while position-residualized (86.4%) and position-matched (86.1%) hidden states stay as separable as the originals.
(a) t-SNE of late-layer hidden states separates efficient (pre-FCS) from overthinking (post-FCS) tokens.
(b) Boundary-aligned probe scores rise sharply at the true FCS; permuted alignment stays flat.
ROM has four components: (1) latent boundary supervision that converts attempt-level correctness into token-level labels at the First-Correct-Solution (FCS) boundary; (2) Counterfactual Self-Correction (CSC), which synthesizes balanced wrong→correct trajectories so the detector learns a semantic boundary around sufficiency rather than an attempt-index shortcut; (3) a streaming detector on frozen late-layer hidden states that emits a per-token overthinking score pt in lockstep with decoding; and (4) boundary-aware intervention that, once pt crosses the threshold, backtracks to the nearest clean sentence/solution boundary and prompts a final answer.
The detector reads one frozen layer in a fixed late-layer band at 83–90% of depth — not a per-model search: Qwen3-8B L32/36, Qwen3-14B L36/40, Gemma-4-12B L40/48, DS-R1-32B L56/64, DS-Llama-8B L28/32. The architecture is shared across backbones; only the input projection scales with hidden size. Every head is trained from the same QwQ-labeled MATH500 traces, so the other four benchmarks are zero-shot transfer for the detector.
Accuracy (%, ±95% CI) / mean output tokens across five backbones and five benchmarks, against ten recent baselines. All methods share identical vLLM decoding, a uniform 8,192-token budget, and n=3 samples per problem (n=10 on AIME25). Overall is the macro-average over the five benchmarks. Bold = highest accuracy / shortest output per model–benchmark. ROM (w/o CSC) is our own ablation, not a baseline. †RP: official probes exist only for DS-R1-32B; baselines are run where their released assets apply.
| Method | MATH500 | GSM8K | AIME25 | GPQA-D | MMLU-Pro | Overall |
|---|---|---|---|---|---|---|
| Qwen3-8B | ||||||
| Vanilla | 90.3±4.2 / 4,569 | 98.9±0.6 / 1,969 | 27.0±14.3 / 6,258 | 37.9±5.8 / 5,791 | 76.7±8.6 / 2,784 | 66.15 / 4,274 |
| EAT | 83.3±7.0 / 2,088 | 96.3±0.8 / 1,563 | 26.0±13.2 / 3,874 | 38.9±5.3 / 4,712 | 76.7±8.8 / 2,052 | 64.23 / 2,858 |
| RCPD | 90.3±4.2 / 4,586 | 95.0±0.9 / 918 | 27.0±14.5 / 6,265 | 36.7±5.7 / 5,900 | 75.7±8.8 / 2,647 | 64.95 / 4,063 |
| RPDM | 90.0±4.2 / 4,519 | 98.8±0.6 / 1,943 | 27.0±14.2 / 6,149 | 37.5±5.8 / 6,137 | 76.2±8.6 / 2,751 | 65.91 / 4,300 |
| NEAT | 85.7±6.2 / 4,143 | 99.1±0.5 / 2,059 | 27.0±15.4 / 7,369 | 36.2±5.7 / 6,225 | 75.7±9.3 / 2,761 | 64.73 / 4,511 |
| TERMINATOR | 88.3±5.5 / 3,086 | 96.7±0.8 / 710 | 26.3±13.1 / 7,203 | 36.7±5.7 / 5,210 | 73.8±9.3 / 1,991 | 64.37 / 3,640 |
| REFRAIN | 86.7±5.8 / 3,345 | 96.7±0.8 / 1,576 | 26.7±13.3 / 5,997 | 43.3±5.8 / 4,643 | 77.6±9.1 / 1,749 | 66.18 / 3,462 |
| Dynasor | 88.7±4.3 / 3,982 | 99.6±0.3 / 1,657 | 26.7±14.2 / 6,184 | 41.8±5.9 / 5,060 | 80.0±8.7 / 2,094 | 67.33 / 3,795 |
| PMA | 84.3±6.2 / 3,164 | 99.2±0.4 / 1,299 | 25.0±13.3 / 5,342 | 38.2±5.8 / 4,731 | 74.8±8.6 / 1,733 | 64.30 / 3,254 |
| PUMA-RD | 77.3±6.2 / 3,210 | 92.5±1.1 / 1,548 | 26.7±14.3 / 6,189 | 33.3±5.5 / 5,248 | 70.5±8.3 / 2,011 | 60.06 / 3,641 |
| ROM (w/o CSC) | 90.3±4.0 / 2,703 | 99.6±0.3 / 669 | 27.0±14.0 / 4,587 | 43.4±5.7 / 4,285 | 78.1±8.8 / 1,095 | 67.70 / 2,668 |
| ROMCSC | 90.7±5.0 / 2,412 | 99.8±0.3 / 627 | 27.3±14.2 / 4,126 | 44.6±6.0 / 3,908 | 79.0±8.8 / 1,009 | 68.29 / 2,416 |
| Qwen3-14B | ||||||
| Vanilla | 89.3±5.2 / 4,038 | 98.3±0.7 / 1,570 | 37.0±15.0 / 5,847 | 53.4±6.1 / 6,060 | 84.3±7.9 / 2,362 | 72.46 / 3,975 |
| EAT | 82.7±6.8 / 2,251 | 96.7±0.8 / 1,284 | 33.7±13.2 / 4,483 | 49.2±6.1 / 4,166 | 82.4±8.1 / 1,937 | 68.91 / 2,824 |
| RPDM | 88.3±5.5 / 4,001 | 98.3±0.7 / 1,585 | 36.3±14.8 / 5,792 | 52.5±6.1 / 6,013 | 83.8±7.9 / 2,349 | 71.86 / 3,948 |
| Dynasor | 86.7±5.8 / 3,129 | 98.8±0.6 / 1,058 | 36.3±14.2 / 5,104 | 56.7±5.7 / 4,903 | 86.2±7.6 / 1,728 | 72.94 / 3,184 |
| PMA | 83.7±6.3 / 3,288 | 97.9±0.7 / 1,163 | 32.3±13.3 / 4,917 | 47.5±6.1 / 5,382 | 81.0±8.3 / 1,856 | 68.47 / 3,321 |
| PUMA-RD | 79.3±6.5 / 3,374 | 94.2±1.0 / 1,247 | 35.0±15.4 / 5,936 | 45.0±5.8 / 5,694 | 77.6±8.6 / 1,994 | 66.21 / 3,649 |
| ROM (w/o CSC) | 88.7±4.8 / 2,694 | 99.2±0.4 / 892 | 36.3±15.2 / 4,126 | 55.0±5.9 / 4,718 | 85.2±7.6 / 1,641 | 72.89 / 2,814 |
| ROMCSC | 89.0±4.3 / 2,437 | 99.6±0.3 / 810 | 36.7±14.7 / 3,728 | 58.2±6.0 / 4,382 | 87.1±7.4 / 1,523 | 74.13 / 2,576 |
| Gemma-4-12B | ||||||
| Vanilla | 78.3±6.7 / 3,789 | 92.5±1.1 / 2,095 | 32.3±16.2 / 7,636 | 31.3±5.7 / 7,005 | 64.3±8.8 / 3,787 | 59.75 / 4,862 |
| EAT | 77.3±6.5 / 1,445 | 95.8±0.9 / 1,537 | 32.3±16.2 / 7,426 | 32.8±5.2 / 5,699 | 70.5±8.3 / 2,206 | 61.76 / 3,663 |
| RPDM | 82.7±5.8 / 3,763 | 92.4±1.1 / 2,081 | 32.3±16.3 / 7,319 | 31.0±5.5 / 6,954 | 65.2±8.6 / 3,735 | 60.73 / 4,770 |
| Dynasor | 83.3±6.2 / 1,733 | 96.7±0.8 / 812 | 32.7±13.1 / 6,196 | 35.9±6.0 / 4,408 | 73.3±8.9 / 1,790 | 64.37 / 2,988 |
| PMA | 84.0±6.2 / 3,392 | 97.5±0.6 / 1,199 | 32.3±16.2 / 7,341 | 30.1±5.3 / 5,060 | 69.0±8.7 / 1,949 | 62.60 / 3,788 |
| PUMA-RD | 76.7±6.8 / 3,505 | 88.3±1.3 / 1,513 | 32.0±15.4 / 7,613 | 28.8±5.2 / 5,523 | 66.2±8.9 / 2,112 | 58.39 / 4,053 |
| ROM (w/o CSC) | 78.7±7.5 / 1,632 | 92.5±1.2 / 504 | 32.7±15.2 / 5,153 | 35.4±5.8 / 4,543 | 72.4±8.8 / 1,708 | 62.32 / 2,708 |
| ROMCSC | 85.0±6.0 / 1,532 | 98.0±0.7 / 479 | 33.0±14.2 / 4,358 | 36.4±5.7 / 4,090 | 73.8±8.4 / 1,633 | 65.22 / 2,418 |
| DS-R1-Distill-Qwen-32B | ||||||
| Vanilla | 83.3±6.2 / 3,367 | 92.8±1.2 / 456 | 29.7±14.0 / 6,145 | 55.6±6.1 / 3,551 | 66.2±8.1 / 1,313 | 65.50 / 2,966 |
| EAT | 77.0±7.0 / 1,693 | 92.5±1.2 / 446 | 27.0±12.3 / 3,762 | 53.9±5.4 / 2,950 | 63.8±7.9 / 1,030 | 62.83 / 1,976 |
| RCPD | 83.3±6.2 / 3,377 | 93.2±1.2 / 459 | 28.3±13.7 / 6,149 | 54.7±5.6 / 3,596 | 64.8±7.9 / 1,324 | 64.88 / 2,981 |
| RPDM | 83.3±6.5 / 3,344 | 92.8±1.2 / 461 | 28.7±13.5 / 6,156 | 55.2±6.0 / 3,498 | 66.2±8.1 / 1,332 | 65.24 / 2,958 |
| Dynasor | 84.0±6.2 / 3,003 | 92.5±1.3 / 337 | 28.7±13.3 / 5,882 | 59.6±6.1 / 2,660 | 72.4±9.1 / 898 | 67.43 / 2,556 |
| PMA | 79.0±6.5 / 2,574 | 92.7±1.2 / 457 | 27.7±13.2 / 5,174 | 51.5±5.6 / 3,098 | 59.5±8.3 / 1,135 | 62.08 / 2,488 |
| PUMA-RD | 81.3±6.5 / 2,790 | 92.7±1.2 / 453 | 28.3±13.7 / 6,073 | 54.0±5.8 / 3,541 | 66.2±7.9 / 1,288 | 64.51 / 2,829 |
| RP† | 80.7±6.8 / 2,440 | 92.2±1.3 / 442 | 28.7±13.8 / 5,838 | 57.1±5.7 / 3,199 | 70.5±8.4 / 1,166 | 65.83 / 2,617 |
| ROM (w/o CSC) | 83.3±6.3 / 1,994 | 93.0±1.2 / 318 | 28.7±12.7 / 4,390 | 58.8±5.9 / 1,679 | 69.5±8.9 / 636 | 66.66 / 1,803 |
| ROMCSC | 83.7±6.8 / 1,820 | 93.5±1.2 / 300 | 29.0±11.8 / 4,015 | 60.9±5.7 / 1,590 | 71.9±8.5 / 592 | 67.80 / 1,663 |
| DS-R1-Distill-Llama-8B | ||||||
| Vanilla | 80.7±6.5 / 897 | 91.0±1.3 / 245 | 17.3±11.7 / 2,613 | 35.7±5.7 / 2,426 | 62.9±9.5 / 484 | 57.52 / 1,333 |
| EAT | 77.3±6.4 / 712 | 90.5±1.3 / 238 | 13.0±9.5 / 1,488 | 33.7±5.1 / 1,360 | 60.5±8.9 / 391 | 55.00 / 838 |
| RPDM | 80.0±6.7 / 881 | 91.1±1.3 / 248 | 17.3±11.2 / 2,627 | 35.4±5.6 / 2,398 | 63.3±9.3 / 492 | 57.42 / 1,329 |
| Dynasor | 80.3±6.8 / 754 | 90.9±1.4 / 198 | 17.7±12.0 / 2,288 | 37.2±5.9 / 1,859 | 63.8±9.8 / 356 | 57.99 / 1,091 |
| PMA | 78.3±6.7 / 771 | 90.8±1.3 / 252 | 15.3±10.8 / 2,107 | 32.0±5.3 / 2,154 | 58.1±9.1 / 419 | 54.91 / 1,141 |
| PUMA-RD | 79.7±6.5 / 796 | 91.0±1.3 / 241 | 17.3±12.2 / 2,461 | 35.9±6.0 / 2,337 | 63.3±9.7 / 448 | 57.43 / 1,257 |
| ROM (w/o CSC) | 78.7±6.4 / 638 | 91.5±1.2 / 181 | 14.3±10.9 / 1,836 | 36.0±5.5 / 1,694 | 61.9±9.2 / 329 | 56.48 / 936 |
| ROMCSC | 81.0±6.1 / 542 | 92.2±1.2 / 161 | 18.0±11.7 / 1,607 | 37.7±5.4 / 1,443 | 64.8±8.9 / 283 | 58.73 / 807 |
Latent-boundary control gives the strongest accuracy–efficiency tradeoff. ROMCSC attains the highest accuracy in 19 of 25 model×benchmark settings while producing the shortest output in 16 of them. Relative to vanilla decoding it improves accuracy in 22 of 25 settings (+0.33 to +9.52 points, mean +2.56) while cutting output length by 28–77% (mean 45%), and it ranks first on both axes in all five Overall columns.
Gains scale with how much redundancy a backbone actually produces. The most verbose backbone gains 4.5× more than the tersest: Gemma-4-12B (24.3k vanilla tokens summed over the five benchmarks, macro gain +5.47 pp) at one end, DS-Llama-8B (6.7k, +1.21) at the other.
Effect of CSC. CSC adds +0.15 to +6.33 accuracy points (mean +1.62) while shortening output by a further 4.4–15.4%, improving both axes in all 25 settings. On DS-Llama-8B, ROM alone falls below vanilla's macro-average (56.48 vs. 57.52); CSC recovers all of it and more (58.73).
Across all 180 pairwise comparisons with vanilla or a baseline, ROMCSC dominates in 165 (higher accuracy and shorter output) and is dominated in none. It therefore lies on the accuracy–length Pareto front in every setting, and is the unique front point in 13 of 25 — reached with a single trigger configuration across all five backbones. Dynasor is the most consistent baseline (non-dominated in 22 of 25), but it stays 7.8–164% longer (+43% on average), and its savings shrink from up to 61% on GSM8K to 1–19% on AIME25: a confidence-gated, answer-defined exit rarely fires on problems the model cannot solve, precisely where compute is most expensive.
Accuracy vs. mean output length in all 25 model×benchmark settings; axis ranges are per panel, upper-left is better. The dashed staircase is the per-panel Pareto front.
If ROM were merely tracking confidence, it should agree with entropy-based stopping case by case. It does not: among 49 EAT-wrong MMLU-Pro samples on Qwen3-8B, ROMCSC corrects 6, spanning both entropy failure modes. ROM tracks whether the trace has crossed from solution construction into redundant continuation, not whether the answer distribution is stable.
| Method | Answer | Reasoning | Response | Stop signal |
|---|---|---|---|---|
| Case 1: confidently wrong (gold D) | ||||
| EAT | I ✗ | 1,570 | 2,194 | exit, H = 7.5×10−5 |
| ROMCSC | D ✓ | 906 | 1,641 | cut at 886 |
| Case 2: correct but uncertain (gold F) | ||||
| EAT | J ✗ | 6,802 | 7,734 | no exit, H̄ = 1.61 |
| ROMCSC | F ✓ | 2,898 | 3,938 | cut at 2,878 |
Two MMLU-Pro cases (Qwen3-8B) where entropy-based stopping fails in opposite directions: EAT exits early once entropy collapses onto a wrong option (Case 1), and never exits when entropy stays unsettled (Case 2).
MATH500 (Qwen3-8B, held-out 100-problem test split, n=3, identical cut/backtrace/continue harness). Each row below the shaded default changes one component of it. SL = mean output tokens.
| Configuration | Acc (%) | SL | |
|---|---|---|---|
| Vanilla (no cut) | 90.3 | 4,569 | |
| ROMCSC (L32, t=0.5, +BT) | 90.7 | 2,412 | |
| Detector | linear head | 90.0 | 4,105 |
| conf. MA (0.98) | 70.0 | 1,105 | |
| conf. MA (0.995) | 67.7 | 944 | |
| Control | w/o backtracing | 89.7 | 2,573 |
| Layer | L22 | 89.3 | 2,108 |
| L34 | 90.7 | 2,315 | |
| Threshold | 0.4 | 89.1 | 1,909 |
| 0.6 | 90.7 | 2,716 | |
| 0.7 | 91.0 | 3,036 | |
The recurrent state is necessary. A linear classifier over the same attention features collapses token-level training accuracy from 96.1% to 62.5% and end-to-end almost never fires; a confidence moving average fails in the opposite direction, triggering on locally confident derivation steps regardless of threshold. Robustness. Across probed layers and thresholds 0.4–0.7, accuracy stays within 1.2 pp of vanilla while compression varies smoothly from 34% to 58%.
Qwen3-8B, n=3. Base is vanilla decoding for open-ended MMLU-Pro (64 non-numerical problems, options removed, GPT-4o judge) and L1-Qwen3-8B-Max for MATH500 (40 problems).
| Setting | Accuracy (%) | Output tokens | ||
|---|---|---|---|---|
| Base | ROMCSC | Base | ROMCSC | |
| MMLU-Pro, open-ended | 80.21 | 81.77 | 2,457 | 1,587 |
| L1-Max, MATH500 | 90.83 | 90.83 | 2,684 | 2,105 |
Open-ended reasoning. With options removed and the free-form answer judged by GPT-4o, ROMCSC shortens responses by 35.4% at no accuracy cost (+1.56 pp). Without options, boxed outputs, or exact-match strings, the gain cannot come from answer formatting: only redundant thinking is removed, not the user-visible explanation.
Composability with RL length control. Stacking the same ROMCSC head on L1-Qwen3-8B-Max — a Qwen3-8B already RL-finetuned for length control — removes another 21.6% of tokens at exactly zero accuracy change. L1 shifts the expected length distribution globally, while ROM detects per-instance saturation.
The streaming head reads the forward pass the backbone already computes and decodes nothing extra, so its fixed per-token cost is quickly dominated by the shorter decoded sequence. On GSM8K with Qwen3-8B, ROMCSC reduces wall-clock time by 46.5% (53.3 → 28.5 s) while adding only ~5% per-token compute (26.1 → 27.4 ms).
| Vanilla | ROMCSC | Δ | |
|---|---|---|---|
| Wall-clock latency | 53.3 s | 28.5 s | −46.5% |
| Per-token compute | 26.1 ms | 27.4 ms | +5.0% |
End-to-end latency. GSM8K with Qwen3-8B.
@misc{wang2026romrealtimeoverthinkingmitigation,
title={ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention},
author={Xinyan Wang and Xiaogeng Liu and Ming Pei and Chaowei Xiao},
year={2026},
eprint={2603.22016},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2603.22016},
}