Claim Verification
Emergent Reasoning Chains in Sparse Mixture Models · 9 claims extracted
All · 9
Verified · 5
Disputed · 1
Unverified · 3
Reasoning-relevant experts specialize early in training, before step 12k.
Verified
91% confidence
4 linked sources
Routing entropy correlates with downstream reasoning accuracy across all three model scales tested.
Disputed
44% confidence
2 linked sources
The probe predicts benchmark performance before fine-tuning completes.
Unverified
0% confidence
0 linked sources
Sparse MoE models outperform dense models of equal active-parameter count on reasoning benchmarks.
Verified
78% confidence
3 linked sources
Reasoning specialization is detectable via linear probing of routing statistics alone.
Verified
85% confidence
5 linked sources
Expert utilization plateaus by the midpoint of training regardless of model scale.
Unverified
12% confidence
1 linked source
Routing entropy is a stronger predictor of reasoning accuracy than raw parameter count.
Verified
82% confidence
4 linked sources
Reasoning capability transfers zero-shot across unrelated benchmark domains.
Unverified
8% confidence
0 linked sources
Probe-based specialization signals generalize to dense transformer baselines.
Verified
73% confidence
2 linked sources