Rank.ai’s blended leaderboard puts Claude Fable 5 narrowly ahead of GPT-5.6 Sol

An October 7 blended leaderboard gives Claude Fable 5 an 85–84 edge over GPT-5.6 Sol, but its component boards use different methods and update dates.

2 min read

Illustrative close-up of a circuit board representing AI agent software infrastructure

Rank.ai updated its blended LLM leaderboard on October 7, placing Anthropic’s Claude Fable 5 first with a consensus score of 85, one point ahead of OpenAI’s GPT-5.6 Sol at 84. The result is not a new benchmark run. It is a normalization of results from Artificial Analysis, LiveBench, LMArena and Berkeley’s Function Calling Leaderboard (BFCL).

In the head-to-head shown by Rank.ai, Claude Fable 5 beats a slightly larger share of models on Artificial Analysis, 49.6% versus 47.0%, and has the higher LMArena rating, 1507 versus 1485. GPT-5.6 Sol is marginally ahead on LiveBench, 79.7% versus 79.5%. Neither model is represented in the BFCL comparison used by the page.

Rank.ai converts each source ranking into the share of other models beaten, averages those values with an added neutral score of 50, and requires coverage from at least two sources. It also keeps each model’s best run and removes dates and reasoning-setting labels when matching names. That makes the table easy to scan, but it can combine variants and settings that developers would normally evaluate separately.

Freshness is the main limitation. Artificial Analysis data was read on October 7, while Rank.ai dates its imported LiveBench board to June 25, LMArena to July 21 and BFCL to April 13. Berkeley’s own BFCL page says its leaderboard was last updated April 12. The four sources also measure different things: composite reasoning and coding tasks, contamination-limited tests, blind human preference, and function-calling accuracy.

The narrow 85–84 gap therefore should not be read as proof that one model is universally better. It is more useful as a discovery index: teams should inspect the underlying benchmark, exact model variant, reasoning setting, price and their own workload before choosing a model.

Circuit board representing the compute infrastructure behind AI model evaluation. Illustrative image.

Sources: Rank.ai blended leaderboard · Rank.ai methodology · Artificial Analysis leaderboard · LiveBench · LMArena leaderboard · Berkeley BFCL V4 · Unsplash image license

AnthropicbenchmarksClaude Fable 5OpenAIleaderboardsGPT-5.6 Sol