We RL train language models how to reason about future events like "Which tech company will the US government buy a > 7% stake in by September 2025?", releasing all code, data, and weights for our model.
Our training makes an 8B model competitive with much larger models like GPT-OSS-120B across judgemental forecasting benchmarks and metrics.