Tuesday 21 July 2026London --:--Frankfurt --:--Zurich --:--
France · News

Claude, ChatGPT, Gemini? Why Comparing AI Systems is Not So Simple

Claude Opus 4.7, GPT 5.5, DeepSeek-V4… Each release of a new AI language model is an opportunity for its creators to showcase better performance. While some rankings allow comparisons, their evaluation is in reality complex.

Claude, ChatGPT, Gemini? Why Comparing AI Systems is Not So Simple

Whether you are an occasional or daily user, you want to know which generative AI tool is the most effective for carrying out your tasks. Nothing could be more normal, especially as the striking claims made by the leaders of major artificial intelligence start-ups suggest, with each new release, spectacular advances.

Regarding its Claude Opus 4.7 model, Anthropic declares that it has “notably improved.” The American company boasts of images analysed with “higher resolution.” Opus 4.7 is also “more tasteful and creative” in performing professional tasks, Anthropic claims. In short, it does everything, but better. Several online rankings confirm this apparent dominance, ahead of Google’s Gemini, Meta’s Muse Spark, or OpenAI’s ChatGPT. But in reality, the question of the performance of these models is more complex than these rankings suggest.

The capabilities of large language models (LLMs) are today evaluated through a set of tests with obscure names: MMLU for general reasoning, SWE-bench for…

Topics
FranceEurope

More from France

FranceVictoria Hitchcock, the Stylist Silicon Valley Can’t Get EnoughFranceLate Spring Heat Wave? French Schools Need Thermal RenovationFranceMichel Barnier on the Neverending Lessons of BrexitFranceOld-School Lessons from a Special French Rural Education Program