Digital Economy

AI is rewriting the rules of testing: Astra overtakes Claude at the top of the Index

The revision of the Artificial Analysis criteria favours agent-based automation and the management of complex documents. The top three places remain firmly in American hands, but the French firm Mistral (18th) is focusing on reduced costs and open standards to secure data sovereignty within businesses.

3' min read

Translated by AI
Versione italiana

3' min read

Translated by AI
Versione italiana

Which AI model is the ‘smartest’? It depends on which day we ask. As late summer 2026 draws to a close, we have learnt that one model may top the rankings today, whilst another may do so tomorrow – and not just because more powerful models (or at least those that appear to be so) are constantly emerging. It is the benchmark itself that is fluid and ever-changing. Take what happened in September with the current benchmark ranking from the specialist firm Artificial Analysis Intelligence. In version 4.3, Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra are tied for the top two spots. They are followed by Claude Opus 5 and Claude Fable 5 (third and fourth place). Five days earlier, however, Astra, OpenAI’s latest gem, was in fifth place. What happened? Artificial Intelligence realised that the benchmark was no longer entirely suitable for assessing the new models and therefore adjusted the weightings of the various tests it subjects them to.

It is important to note that the Intelligence Index does not measure a generic form of ‘intelligence’ in AI. It aggregates ten benchmarks, each with different weightings.

Loading...

30 per cent relates to agents, that is, the ability to complete tasks comprising multiple operations. 20 per cent relates to coding, 30 per cent to general abilities such as knowledge, handling of hallucinations and comprehension of lengthy documents, and the remaining 20 per cent to scientific reasoning.

The tests include AA-Briefcase and GDPval-AA v2 for professional work; AutomationBench-AA for business workflows; Terminal-Bench and SciCode for programming; Humanity’s Last Exam and CritPt for reasoning; AA-Omniscience for knowledge and hallucinations; GDP.pdf and AA-LCR for very long documents and contexts.

Now, the latest version of the Index places greater emphasis on the ability to carry out long and complex professional tasks and to analyse very lengthy documents. In short, precisely the things at which Astra excels and where, until now, Anthropic had held an advantage over OpenAI. This refers to autonomous, ‘agent-based’ AI work on tasks that, until recently, were considered exclusively ‘human’. Experts agree that humans still hold an advantage when it comes to long-term planning capabilities, but the gap with AI is narrowing rapidly. For better or worse, this is evident when one considers how much long-term planning (spanning months) was required for the well-known attack (discovered in July) carried out by OpenAI agents on Hugging Face during a test. It involved not only planning, but also a complex collaboration between agents (unbeknown to humans).

Well, if this is the cutting edge, the podium is an all-American affair, with Meta’s Muse Spark 1.3 (released in September) in fifth place – also adept at using a range of tools autonomously to achieve a goal. In sixth place is OpenAI’s GPT 5.6 Max. China follows closely behind with GLM 5.3 Max.

There are also differences in ability when it comes to specific tasks, which need to be taken into account.

Claude Fable 5.1 is particularly strong when it comes to complex professional tasks, such as work projects involving thousands of files where it needs to produce analyses, documents, spreadsheets or presentations. In this respect, it outperforms Astra, which, however, is better at programming (though Fable 5.1 is better at scientific programming).

Astra comes into its own when AI needs to act autonomously within business workflows in the finance, marketing, sales, human resources and operations sectors. Astra also costs less than Fable 5.1, but more than Glm 5.3 (among the top models). European models make their first appearance in 18th place with the French Mistral, which has just raised 3 billion euros. It is very inexpensive (although there are cheaper models, such as Deepseek 4.1 Flash). It is an open-weights model, like many of the Chinese ones, and unlike the top models from OpenAI and Anthropic. As it is open-source, it can therefore be freely downloaded and modified.

That 18th-place finish does not, therefore, mean that Mistral is irrelevant to businesses. In applications where infrastructure control, the ability to run the model on one’s own systems, data sovereignty or customisation are key, openness is an asset. Moreover, unlike benchmarks, this at least is an objective fact

Copyright reserved ©

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti