07/20/2026 | Press release | Distributed by Public on 07/21/2026 07:23
This article was originally posted on LinkedIn.
If you plot every model on the Aiera leaderboard by its composite score, you don't get a smooth slope. Instead, you get: two models sitting alone at the top, a long flat plateau, and then a cliff. Looking closely, those stages tell a very interesting story about 2026 so far.
Let's start with the top half of the leaderboard: starting with Claude Fable 5 at #3 (64.9) and moving down to Claude Sonnet 4.6 at #16 (58.1), fourteen models are separated by just 6.85 points. Most of these models were released in just the last three months; on the axis of "which new model is best overall," they are, for practical purposes, a near-tie, with scores differences that could be argued are mostly noise.
Meanwhile, below Sonnet 4.6 the floor drops out: about 8 points are lost by #20, and a further 8 points by #24. The bottom of the board (single-digit research scores) isn't even in the same conversation as the plateau. The field isn't gradually spreading out; it's being compressed into a dense band at the top, with a sharp descent.
The slate of models released in 2026 have mostly started to converge, and selecting one over the other has become a pricing and throughput exercise (with some specific task performance to be measured and considered). This is at the heart of the Kimi K3 freak-out, as the frontier closed-source providers realize their moat might be shallow.
The exceptions to this phenomenon (for now) are gpt-5.6-sol (69.6) and gpt-5.5 (68.9). The single step from #3 up to #2 is 3.95 points, which is more than half the entire 14-model plateau internal spread. These two aren't leading the pack; they've left it.
What separates them is almost entirely one thing: research. Our composite weights four tasks: research (60%), Q&A (24%), summarization (10%), sentiment (6%); and research is the only task measured by connecting the model to live, proprietary data and asking it to answer questions the open web can't. The two leaders average 61.7 on research; the plateau averages 53.5. Run the decomposition and ~70% of the gap from #3 to #2 is in the research axis alone. For our clients, that is quite significant.
The plateau isn't a coincidence; it's what happens when scoring axes saturate.
Three of the four core tasks we measure are things that modern models are simply good at now. When a capability saturates, it stops discriminating: nearly everyone scores well, so that slice of the composite becomes roughly constant across the field. Forty percent of the Aiera Score is now close to flat from #3 to #16. That compression is the plateau.
The remaining 60% of the score (research) is the one axis that hasn't yet saturated, because it isn't a test of what the model knows but of whether it can drive tools, retrieve the right proprietary data, and reason against that data to create a correct, well-sourced answer. That's still hard, and it's where the real spread survives.
The cliff below #16 is also telling: models that can't reliably drive the tools or effectively manage the data have research scores that collapse toward zero. Notably, which side of this cliff a model lands on isn't predicted by size or lab prestige: open-weight models intermix with frontier labs on the plateau, and some very large models fall off it.
For practitioners, the takeaway is that "best overall model" is becoming a less useful question than "best at the axis you care about." If your workload is general finance NLP (extraction, summarization, etc.) the field has converged. A dozen models are effectively interchangeable; choose on cost, latency, context window, or open weight, not leaderboard rank. But if your workload is agentic, deep financial analysis (reasoning over live transcripts, filings, and institutional research) a genuine gap remains, it's fairly wide, and it's where you should spend your evaluation budget.
The tidy version of the 2026 story is "models are plateauing." The board tells a sharper one: the common axes have plateaued, and model differentiation has migrated almost entirely onto the hard, data-connected one, where a small ceiling still stands and most of the field is bunched a few points below it. This is precisely what a research-weighted benchmark is built to surface, and why we pay such close attention to it.