Semantic search
↵ Search
Semantic
By claim
By author
Sort: Relevance
6 results, ranked by semantic relevance
Compute-Optimal Reasoning Without Scale
94% match
Bakshi, P. · Elorza, J.
…we present evidence that reasoning capability plateaus independent of parameter count once routing specialization saturates, directly contradicting the strong-scaling view…
Diminishing Returns in Chain-of-Thought Scaling
91% match
Marchetti, S. · Obi, C. · Reyes, D.
…our ablations show that beyond a critical model size, additional parameters yield no measurable gain in multi-step reasoning accuracy…
Small Models, Structured Reasoning
87% match
Lindqvist, A. · Tanaka, H.
…a 1.3B model with explicit routing supervision matches the reasoning benchmark performance of models 20x its size…
Routing Entropy as an Early Signal for Reasoning Onset
83% match
Alaya, R. · Whitcombe, N.
…tracking per-layer routing entropy during pretraining predicts downstream reasoning benchmark accuracy well before fine-tuning begins…
Scaling Laws Revisited for Sparse Architectures
78% match
Okonkwo, T. · Ferro, D.
…standard dense-model scaling laws systematically overpredict reasoning gains once expert routing is introduced…
Critical Capacity Thresholds in Multi-Step Inference
72% match
Voss, E. · Castellano, G.
…we identify a sharp capacity threshold below which chain-of-thought supervision fails to induce reliable multi-step reasoning…