news

Aug 2026 Our updated ROM paper is now on arXiv. We frame overthinking in large reasoning models as a latent productive-to-redundant transition that is directly decodable from hidden states around first-correct-solution (FCS) boundaries. ROM turns this signal into control: a lightweight streaming detector (~0.1% of backbone parameters) monitors a frozen LRM and intervenes at well-formed reasoning boundaries — no answer extraction, no probe decoding, no backbone updates. Our Counterfactual Self-Correction (CSC) augmentation preserves pre-FCS self-correction. Across five backbones, five benchmarks, and ten baselines under a shared protocol, ROM attains the highest accuracy in 19 of 25 settings, cuts response length 28–77% (mean 45%), and is the only method on the accuracy–length Pareto front in every setting; the same MATH500-trained head transfers zero-shot and cuts wall-clock latency by 46.5%. Check out our project page, code, and dataset.
May 2026 Our new paper PW-OPSD is now on arXiv! We find that teacher-token reliability in on-policy self-distillation for reasoning is position-structured, and propose Position-Weighted On-Policy Self-Distillation (PW-OPSD) to up-weight reliable later tokens at no extra teacher cost. Check out our paper and code.
Apr 2026 Our paper ReasoningBomb has been accepted to ACM CCS 2026! We propose an RL-based framework that crafts short, natural-language prompts to trap LRMs into pathologically long reasoning, with a constant-time surrogate reward enabling 4.39×10⁵× training speedup. Just 10% malicious traffic cuts benign throughput by 49.8% and monopolizes 64.3% of compute. Check out our paper, website, code, and dataset.
Feb 2026 ReasoningBomb is now on arXiv. Check out our website, code, and dataset.
Sep 2024 Passed my Qualifying Exam.