AI is rewriting the rules of testing: Astra overtakes Claude at the top of the Index
The revision of the Artificial Analysis criteria favours agent-based automation and the management of complex documents. The top three places remain firmly in American hands, but the French firm Mistral (18th) is focusing on reduced costs and open standards to secure data sovereignty within businesses.
Which AI model is the ‘smartest’? It depends on which day we ask. As late summer 2026 draws to a close, we have learnt that one model may top the rankings today, whilst another may do so tomorrow – and not just because more powerful models (or at least those that appear to be so) are constantly emerging. It is the benchmark itself that is fluid and ever-changing. Take what happened in September with the current benchmark ranking from the specialist firm Artificial Analysis Intelligence. In version 4.3, Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra are tied for the top two spots. They are followed by Claude Opus 5 and Claude Fable 5 (third and fourth place). Five days earlier, however, Astra, OpenAI’s latest gem, was in fifth place. What happened? Artificial Intelligence realised that the benchmark was no longer entirely suitable for assessing the new models and therefore adjusted the weightings of the various tests it subjects them to.
It is important to note that the Intelligence Index does not measure a generic form of ‘intelligence’ in AI. It aggregates ten benchmarks, each with different weightings.
30 per cent relates to agents, that is, the ability to complete tasks comprising multiple operations. 20 per cent relates to coding, 30 per cent to general abilities such as knowledge, handling of hallucinations and comprehension of lengthy documents, and the remaining 20 per cent to scientific reasoning.
The tests include AA-Briefcase and GDPval-AA v2 for professional work; AutomationBench-AA for business workflows; Terminal-Bench and SciCode for programming; Humanity’s Last Exam and CritPt for reasoning; AA-Omniscience for knowledge and hallucinations; GDP.pdf and AA-LCR for very long documents and contexts.
Now, the latest version of the Index places greater emphasis on the ability to carry out long and complex professional tasks and to analyse very lengthy documents. In short, precisely the things at which Astra excels and where, until now, Anthropic had held an advantage over OpenAI. This refers to autonomous, ‘agent-based’ AI work on tasks that, until recently, were considered exclusively ‘human’. Experts agree that humans still hold an advantage when it comes to long-term planning capabilities, but the gap with AI is narrowing rapidly. For better or worse, this is evident when one considers how much long-term planning (spanning months) was required for the well-known attack (discovered in July) carried out by OpenAI agents on Hugging Face during a test. It involved not only planning, but also a complex collaboration between agents (unbeknown to humans).
-U25561633361BWT-1440x752@IlSole24Ore-Web.jpg?r=650x341)
