Last week, we had the release of Anti Gravity and Gemini Pro 3, right after that Codex Max 5.1, and this week (11/26/2025) we have Opus 4.5 with a Benchmark surpassing its “adversaries.” But what, after all, does each benchmark actually measure? And how do they help us understand whether a model is better—or, if you prefer, closer to that so-called AGI?
Benchmark in a loose translation comes out as “reference point.” In the current context, the word is widely used as a performance measurement tool, whether for a product, investment, or an LLM. In our case, I’m going to explain (or at least try to), in a simple and direct way, the main benchmarks that show up in those comparison charts we see so much on social media.

Agentic coding — SWE-bench
SWE comes from the term Software Engineer, basically this benchmark evaluates a model’s ability to solve real tasks in Python repositories on GitHub. The model receives the project code, the task, and needs to propose a fix that actually solves the problem.
In this case, in Opus’s evaluation above, there was more or less a subset with ~500 hand-reviewed tasks, with tests and fixes validated by humans to avoid false positives.
The tasks can be broad, but they are usually more or less like this:
- “This test is breaking, fix the bug.”
- “This function needs to support a new parameter.”
- “This API is returning the wrong value, adjust the logic.”
The agent must:
- Understand the task description.
- Read the code (sometimes several files).
- Edit the correct files.
- Run tests and make sure everything passes.
When you ask CoPilot, Codex, Claude to execute a task, it needs to execute it and have an acceptable answer. It’s the same here.
A score above 80%, like the one achieved by Opus 4.5, is impressive.
But there are limitations: the benchmark is focused on Python, depends on fixed repositories, and there is always the risk of “contamination”—the model may already have seen parts of the code during training.

Agentic terminal coding — Terminal-Bench 2.0
This one is very similar to SWE, but the difference is that it runs as a real Linux through the command line. The scenarios range from compiling projects to training ML models, configuring servers, and resolving broken dependencies. For example, the agent can clone a repo, install dependencies, run tests and fix errors, or configure a web server, adjust the firewall, bring up a service.
The agent needs to navigate the terminal (git, pip, sed, systemctl, etc.), edit files with CLI editors and, in the end, leave the environment in the correct state. Basically, it’s what you currently do with Claude Code/Codex Cli.
A high score in Terminal-Bench indicates a model that can act like a “full stack dev Ops” in the terminal, executing end-to-end pipelines. In the same way SWE has its limitations, Terminal Bench cannot capture scenarios using a GUI (Graphical User Interface), and as realistic as the scenarios are, they are still several scenarios in Docker, which does not necessarily reflect the real world.

Agentic tool use — τ²-bench
τ²-bench (read as “tau-squared”) is a benchmark for conversational agents that use tools in customer service scenarios, mainly in domains such as retail and telecom. In other words, this benchmark is specific to be used as a ChatBot.
This benchmark evaluates the model’s ability to solve billing problems, diagnose connection problems using network tools. In the end, it is important for “call center agents,” corporate chatbots.
But there are important limitations:
- it does not measure interaction with confused or hostile customers,
- it does not test voice noise,
- nor real human ambiguity.
It is a great benchmark for an “ideal chatbot,” not for the chaos of the real world.
Scaled tool use — MCP Atlas
MCP Atlas is a benchmark made by Scale/SEAL to evaluate tool use at scale through the Model Context Protocol (MCP). The idea is to measure how the model performs when it has access to many different tools (databases, APIs, search engines, etc.) and needs to orchestrate calls in complex chains.
It tests scenarios where the model has access to:
- APIs,
- databases,
- search engines,
- internal systems,
- complex chains of actions.
The question here is:
Does the model know how to orchestrate several different tools to solve a complex problem?
Since it is recent, there is still no comparison base as solid as in other benchmarks.

Computer use — OSWorld
OSWorld is a benchmark of computer use: the agent controls a real desktop (or a simulated environment very close to it) with apps such as a browser, text editor, spreadsheet, and system tools. The tasks are real work routines, with multiple steps and windows. Such as:
- Download a file from the web, organize it into folders, unzip it.
- Fill in spreadsheets, make charts, and send them by email.
- Adjust system settings, install programs, etc.
The benchmark measures whether the task was completed and, in newer versions, also efficiency (how many extra steps the agent took compared to a human). The problem is that interfaces can change over time, and there is a finite number of tasks.

Novel problem solving — ARC-AGI-2
This is the most “mystical,” the most hyped, the most feared, the great ARC-AGI! ARC-AGI, created by François Chollet, tries to measure abstract reasoning and strong generalization, that is: the model’s ability to infer rules that are not explicit in text, but hidden in visual patterns.
Each problem shows small grids of colored squares (input/output), and the model needs to infer the implicit rule and apply it to new examples. ARC-AGI-2 is the new version, even harder, used today as a “stress test” of systems on the path toward AGI. ARC-AGI-2 is even harder and is used as a stress test of “system intelligence”—the famous “how far are we from AGI?”.
A “trained” human could in theory solve 100% of the ARC-AGI 2 panel, however if a model can solve 45% as was done by the latest version of Gemini 3 Deep Think, that is something far outside the curve.
But it is important to make it clear, the tasks you see in ARC AGI do not usually “happen” in real life. And the use of external tools also impacts the result.

Graduate-level reasoning — GPQA Diamond
GPQA is a multiple-choice exam in physics, chemistry, and biology at the graduate level.
The Diamond version is the elite of the elite:
- difficult questions,
- manually curated,
- solved only by PhD specialists.
A good score means the model can deal with complex reasoning and advanced content—something like a “good graduate student.”
The problem is that the focus stays only on three areas (physics, chemistry, biology); nothing from the humanities, specific engineering fields, etc. And because it is multiple choice, it limits the model’s “creativity.”

Visual reasoning — MMMU
MMMU (Massive Multi-discipline Multimodal Understanding) is a multimodal benchmark that mixes image + text in multiple-choice questions. It has more than 11 thousand questions collected from exams, books, and quizzes at the university level in several areas (art, business, medicine, sciences, engineering, humanities) and several languages.
A good result indicates that the model can reason and retrieve knowledge in several languages, not just literally translate from English, but also understand cultural and contextual nuances.
The problem is that translations can introduce noise, ambiguities, or cultural differences, besides a possibility of data contamination.
Conclusion
Benchmarks are essential to measure the progress of LLMs and understand where each model is strongest.
But there is no single benchmark that indicates “proximity to AGI.”
Each test measures a different piece of intelligence:
- some focus on practical capability,
- others on abstract reasoning,
- others on technical knowledge,
- others on tool use.
The correct interpretation requires looking at the whole set, not an isolated number on the chart. And it is important to remember, even with all this, it still does not mean AGI is right there, but rather that every time a model gets a higher score than the previous ones, we are one step closer to it. And there are still many steps left, many before we get there.
References and useful links
Agentic coding — SWE-bench (and variants)
- Official site / overview:
https://www.swebench.com/ - Original dataset:
https://www.swebench.com/original.html - GitHub repository:
https://github.com/SWE-bench/SWE-bench
Agentic terminal coding — Terminal-Bench 2.0
- GitHub repository (real terminal, DevOps tasks, etc.):
https://github.com/laude-institute/terminal-bench - Leaderboard (models and scores):
https://artificialanalysis.ai/evaluations/terminalbench-hard
Agentic tool use — τ²-bench (tau-squared bench)
- Paper / technical description (arXiv):
https://arxiv.org/abs/2506.07982 - GitHub repository:
https://github.com/sierra-research/tau2-bench - Research page / overview:
https://sierra.ai/resources/research/tau-squared-bench
Scaled tool use — MCP Atlas
- Official Scale AI leaderboard:
https://scale.com/leaderboard/mcp_atlas
Computer use — OSWorld
- Official benchmark site:
https://os-world.github.io/ - GitHub repository:
https://github.com/xlang-ai/OSWorld
Novel problem solving — ARC-AGI / ARC Prize
- Official ARC-AGI page:
https://arcprize.org/arc-agi - ARC Prize Foundation site (competition linked to the benchmark):
https://arcprize.org/ - GitHub repository (original ARC dataset):
https://github.com/fchollet/ARC-AGI
Graduate-level reasoning — GPQA / GPQA Diamond
- Original paper (GPQA):
https://arxiv.org/abs/2311.12022 - Official repository:
https://github.com/idavidrein/gpqa - Benchmark page / leaderboard:
https://epoch.ai/benchmarks/gpqa-diamond
Visual reasoning — MMMU
- Official MMMU site:
https://mmmu-benchmark.github.io/ - GitHub repository:
https://github.com/MMMU-Benchmark/MMMU - Paper (arXiv):
https://arxiv.org/abs/2311.16502
Multilingual Q&A — MMMLU / Multilingual MMLU
- Paper that discusses MMMLU (MMLU-ProX, with a section on the multilingual variant):
https://arxiv.org/abs/2503.10497 - Leaderboard / overview of models in multilingual benchmarks:
https://huggingface.co/spaces/StarscreamDeceptions/Multilingual-MMLU-Benchmark-Leaderboard