← arXiv

arXiv·3 min read

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

It investigates whether confidence supervision can improve chain‑of‑thought efficiency without explicit length penalties.

Chain‑of‑thought models for math, science, and code tend to emit lengthy reasoning traces, inflating inference cost. Existing remedies rely on early‑stopping heuristics or reinforcement‑learning penalties that directly target trace length. The authors set out to see if a different supervision signal, confidence, could yield efficiency gains without such explicit constraints.

They fine‑tune large language models on a self‑generated dataset of 600 problems, training the model to predict its own confidence in the answer at intermediate steps of its own reasoning. The loss only penalizes confidence prediction error; there is no term for trace length or early stopping, and inference proceeds with the standard generation pipeline, never querying confidence.

Across Gemma, Qwen, Nemotron, and GPT‑OSS models, the confidence‑fine‑tuned versions cut generated tokens by as much as 25% while matching baseline accuracy on mathematical, scientific, and coding benchmarks. The efficiency gain rivals methods that explicitly optimize for shorter reasoning, and analysis shows the models retain their original high‑level reasoning structure.

TakeawaySelf‑supervised confidence fine‑tuning trims token consumption up to 25% without sacrificing accuracy.

Prodigy briefing — continue on the original for source material, discussion, and updates.

Read the paper ↗