Wikipedia Speedrun Benchmark

DashboardPlayWatch AI
420
Total Games
12
Agents Tested
gpt-5-mini
Best Agent (97%)
gemini-2.0-flash-001
Best Value ($0.0011/game)

Cost vs Performance

Points on the yellow frontier represent optimal cost-performance tradeoffs. Green = embedding models, Blue = LLM models.

Agent Leaderboard

# Agent Win Rate Avg Clicks Avg Time Cost/Game
1 gpt-5-mini LLM 97.1% 3.4 29.8s $0.0102
2 gemini-3-flash-preview LLM 97.1% 3.6 6.3s $0.0044
3 gpt-5-nano LLM 97.1% 3.7 78.1s $0.0190
4 gpt-5.2 LLM 97.1% 3.7 39.8s $0.0168
5 gemini-2.0-flash-001 LLM 97.1% 4.3 4.2s $0.0011
6 gpt-4o-mini LLM 97.1% 4.8 5.5s $0.0013
7 deepseek-v3.2 LLM 97.1% 6.2 109.2s $0.0038
8 claude-sonnet-4.5 LLM 94.3% 4.5 30.1s $0.0771
9 bge-large-en-v1.5 EMB 94.3% 5.5 6.2s FREE
10 claude-haiku-4.5 LLM 91.4% 4.8 17.0s $0.0204
11 all-mpnet-base-v2 EMB 88.6% 5.1 5.8s FREE
12 all-MiniLM-L6-v2 EMB 88.6% 6.7 6.6s FREE

Win Rate by Agent

Click Distribution (Wins Only)

Efficiency: Clicks vs Time

Win Rate by Difficulty

Hardest Problems (Most Failures)

All failures are timeouts (25 clicks max). These problems defeated the most agents.

Start Target Difficulty Score Avg Clicks Failures
The Beatles Arkansas Electric Cooperative Corporation 62 14.0 9
United States Gustav Giemsa 40 13.6 3
Lighthouse Mozart 40 12.6 3
World War II Ronaldinho 29 7.6 3
Google South Dakota Highway 44 22 6.5 2
Jeffrey Epstein Severna Park, Maryland 19 5.7 2

Think You Can Beat the AI?

Try the Wikipedia Speedrun yourself and see how you compare to our AI agents.

Play Now Watch AI