Search papers, authors, claims…

2 processing

Semantic search

↵ Search

Semantic

By claim

By author

Sort: Relevance

6 results, ranked by semantic relevance

Compute-Optimal Reasoning Without Scale

94% match

Bakshi, P. · Elorza, J.

…we present evidence that reasoning capability plateaus independent of parameter count once routing specialization saturates, directly contradicting the strong-scaling view…

Diminishing Returns in Chain-of-Thought Scaling

91% match

Marchetti, S. · Obi, C. · Reyes, D.

…our ablations show that beyond a critical model size, additional parameters yield no measurable gain in multi-step reasoning accuracy…

Small Models, Structured Reasoning

87% match

Lindqvist, A. · Tanaka, H.

…a 1.3B model with explicit routing supervision matches the reasoning benchmark performance of models 20x its size…

Routing Entropy as an Early Signal for Reasoning Onset

83% match

Alaya, R. · Whitcombe, N.

…tracking per-layer routing entropy during pretraining predicts downstream reasoning benchmark accuracy well before fine-tuning begins…

Scaling Laws Revisited for Sparse Architectures

78% match

Okonkwo, T. · Ferro, D.

…standard dense-model scaling laws systematically overpredict reasoning gains once expert routing is introduced…

Critical Capacity Thresholds in Multi-Step Inference

72% match

Voss, E. · Castellano, G.

…we identify a sharp capacity threshold below which chain-of-thought supervision fails to induce reliable multi-step reasoning…