Seven frontier models across four reasoning benchmarks. Hover any cell to see the exact score and source. N/R = not reported in primary paper.
One benchmark. Four years. The steepest capability jump in benchmarking history. Hover any point for the score and date.
A model that reasons well on math should reason well on PhD-level science. The scatter confirms it. Hover any dot.
GPT-4 vs. o1. Same company. One architectural shift. The gap is not incremental. Hover any bar.
Seven inflection points that changed what AI can think through. Hover any milestone for details.