AI Models
Frontier Learning LLM Reasoners: Why Fixed Problem Pools Go Stale
A new arXiv preprint on frontier learning LLM reasoners argues that fixed problem sets go stale under GRPO training. Here is the claim, and what remains unproven.