Key Takeaways
- Claude Opus 5 (Adaptive Reasoning, Max Effort) currently leads the Artificial Analysis Intelligence Index with a score of 61.
- Artificial Analysis says it has evaluated 170 models overall, including 94 open weights models and 126 reasoning models.
- The leaderboard is part of a broader comparison that also tracks speed, latency, pricing, context window, and other model metrics.
What happened
Artificial Analysis’s model leaderboard currently places Claude Opus 5 (Adaptive Reasoning, Max Effort) in first place on its Intelligence Index with a score of 61. The site describes the Intelligence Index as one of several ways it compares model performance, and it says the rankings are based on standardized prompts.
The leaderboard page shows a broader field around that top result. According to the source text, the next entries are Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) and Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback), both at 60, followed by GPT-5.6 Sol (max) at 59 and Claude Opus 5 (Adaptive Reasoning, High Effort) also at 59.
Artificial Analysis says its evaluation set spans 170 models in total. Of those, 94 are open weights models and 126 are reasoning models. The site also calls out that Claude Opus 5 (Adaptive Reasoning, Max Effort) leads among the reasoning-model subset.
The page is not limited to a single score. It presents a set of comparison categories that include intelligence, pricing, output speed, latency, end-to-end response time, context window size, and model size for open weights models. The source also mentions related indexes and evaluation tools such as AA-Briefcase, AA-Briefcase Elo, AA-Omniscience, and the Artificial Analysis Openness Index.
In addition to the Intelligence Index, the page highlights other top performers in separate categories. Mercury 2 is listed as the fastest model at 901.6 tokens per second. Nova Micro is shown as the least expensive at $0.03 per 1M tokens blended. Gemini 2.5 Flash-Lite (Non-reasoning) is said to have the lowest time to first token at 0.33 seconds.
For open weights models, GLM-5.2 (max) is the highest-ranked entry with an Intelligence Index score of 51. Artificial Analysis lists MiniMax-M3 and DeepSeek V4 Pro (Reasoning, Max Effort) next in that subset, both at 44.
Why it matters
A single leaderboard position does not settle the question of which model is best for every use case, but it does show how quickly the top tier of model performance is moving. Artificial Analysis’s page illustrates that model evaluation is no longer just about raw intelligence scores. It is also about tradeoffs among cost, latency, throughput, and context handling.

That matters because the most capable model is not always the most practical one. A model can lead on one benchmark while another is better suited to production systems because it returns a first token faster, costs less to run, or delivers higher output speed. The source’s layout reflects that reality by putting intelligence alongside operational metrics rather than treating the leaderboard as a one-dimensional race.
The breakdown also shows a split between proprietary and open weights models. Artificial Analysis notes that there are 94 open weights models among the 170 evaluated, and that GLM-5.2 (max) leads that category. For developers and buyers comparing deployment options, that distinction can be as important as the absolute top score.
The emphasis on reasoning models is another signal. Artificial Analysis says Claude Opus 5 (Adaptive Reasoning, Max Effort) leads among 126 reasoning models, and the page defines reasoning models as those that use extended thinking before answering. That makes the leaderboard relevant not only for general model ranking, but also for understanding how extended reasoning changes the competitive landscape.
What to watch
The immediate question is whether Claude Opus 5 holds the top spot as more models are added to the leaderboard. Artificial Analysis says its scores are based on standardized prompts, but the page also makes clear that rankings are organized by effort settings and model variants, which means the ordering can change as vendors release new versions or tuning modes.
It will also be worth watching how the surrounding metrics evolve. The source highlights speed leaders, latency leaders, low-cost models, and open weights standouts separately from the overall intelligence ranking. That suggests future model comparisons may increasingly be judged by combinations of performance measures rather than by a single headline score.
Another thing to watch is how users interpret the difference between the top proprietary models and the best open weights systems. Artificial Analysis’s current listing shows a gap between the overall leader and the open weights leader, but the page also makes clear that open weights models remain a substantial part of the evaluated field.
For now, the main takeaway is straightforward: on Artificial Analysis’s current leaderboard, Claude Opus 5 is at the top of the Intelligence Index, while the rest of the page shows how much more complicated model comparison has become.



