AI Research  ·  Visual Essay  ·  August 2, 2026

The Word Is
Reasoning

From 5% to 96.4% on graduate math. In 4 years. Here is the data.
5%
GPT-3 on MATH
benchmark, 2021
96.4%
o1 on MATH
benchmark, 2024
4 yrs
Time elapsed
between those two
77.3%
o1 on GPQA Diamond
PhD-level science
Chart 01 of 05

Where Every Model Stands Right Now

Seven frontier models across four reasoning benchmarks. Hover any cell to see the exact score and source. N/R = not reported in primary paper.

Benchmark Heatmap: Score (%) by Model and Task
Sources: Hendrycks et al. arXiv:2103.03874 (MATH); Hendrycks et al. arXiv:2009.03300 (MMLU); Chen et al. arXiv:2107.03374 (HumanEval); Rein et al. arXiv:2311.12022 (GPQA Diamond). Model cards: GPT-4 (arXiv:2303.08774), Claude 3.5 Sonnet (Anthropic 2024), Gemini 1.5 Pro (arXiv:2403.05530), o1 (OpenAI system card August 2, 2026), DeepSeek R1 (arXiv:2501.12948), Kimi k1.5 (arXiv:2501.12599), Kimi K3 (arXiv:2607.24653). N/R = benchmark not reported in primary paper.

Chart 02 of 05

The MATH Climb

One benchmark. Four years. The steepest capability jump in benchmarking history. Hover any point for the score and date.

MATH Benchmark Score (%) Over Time
GPT-3 (2021): Hendrycks et al. arXiv:2103.03874. GPT-4 (2023): arXiv:2303.08774. Claude 3.5 Sonnet (2024): Anthropic model card. o1 (late 2024): OpenAI o1 system card.

Chart 03 of 05

Math vs. Science: Who Generalizes?

A model that reasons well on math should reason well on PhD-level science. The scatter confirms it. Hover any dot.

MATH Score (x-axis) vs. GPQA Diamond Score (y-axis)
GPQA Diamond tests PhD-level biology, chemistry, and physics. Correlation between MATH and GPQA scores suggests reasoning transfers across domains. Sources as above.

Chart 04 of 05

Before and After Reasoning Models

GPT-4 vs. o1. Same company. One architectural shift. The gap is not incremental. Hover any bar.

GPT-4 vs. o1: Benchmark Comparison (%)
GPT-4 scores from arXiv:2303.08774. o1 scores from OpenAI o1 system card, August 2, 2026.

Chart 05 of 05

The Reasoning Timeline

Seven inflection points that changed what AI can think through. Hover any milestone for details.

Key Milestones in AI Reasoning Capability (2021-2026)
Chain-of-Thought prompting: Wei et al. arXiv:2201.11903 (NeurIPS 2022). GPT-4: arXiv:2303.08774. Kimi k1.5: arXiv:2501.12599. DeepSeek R1: arXiv:2501.12948. Kimi K3: Moonshot AI, arXiv:2607.24653 (August 2, 2026).
19x
MATH score gain
2021 to 2024
97.3%
DeepSeek R1
on MATH, August 2, 2026
CoT
Chain-of-Thought:
the technique that started it all

Sources

  1. Hendrycks et al. "Measuring Mathematical Problem Solving." arXiv:2103.03874
  2. Hendrycks et al. "Measuring Massive Multitask Language Understanding." arXiv:2009.03300
  3. Chen et al. "Evaluating Large Language Models Trained on Code." arXiv:2107.03374
  4. Rein et al. "GPQA: A Graduate-Level Google-Proof Q&A Benchmark." arXiv:2311.12022
  5. OpenAI GPT-4 Technical Report. arXiv:2303.08774 (2023)
  6. Anthropic Claude 3.5 Sonnet Model Card (2024)
  7. Reid et al. "Gemini 1.5: Unlocking multimodal understanding." arXiv:2403.05530
  8. OpenAI o1 System Card. OpenAI, August 2, 2026
  9. DeepSeek-AI. "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs." arXiv:2501.12948
  10. Wei et al. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv:2201.11903 (NeurIPS 2022)
  11. Team Kimi. "Kimi k1.5: Scaling Reinforcement Learning with LLMs." Moonshot AI, arXiv:2501.12599 (2025)
  12. Kimi Team. "Kimi K3: Open Frontier Intelligence." Moonshot AI, arXiv:2607.24653 (August 2, 2026)

Excited about AI, innovation, and growth?

Start a conversation