New empirical follow-up: Part 2 compares seven trainer configurations on Qwen3-14B and our DEX gym , with no-think and thinking holdouts, interactive reward curves, and downloadable data. It is a separate experiment from the 27B table below and does not establish a universal trainer ranking. If you only ever train language models on toy math puzzles, reinforcement learning feels simple: did the model output 42? If yes, reward is 1. If no, reward is 0. The moment you try to train an aut...