Last week, we had the release of Anti Gravity and Gemini Pro 3, right after that Codex Max 5.1, and this week (11/26/2025) we have Opus 4.5 with a Benchmark surpassing its “adversaries.” But what, after all, does each benchmark actually measure? And how do they help us understand whether a model is better—or, if you prefer, closer to that so-called AGI?

Benchmark in a loose translation comes out as “reference point.” In the current context, the word is widely used as a performance measurement tool, whether for a product, investment, or an LLM. In our case, I’m going to explain (or at least try to), in a simple and direct way, the main benchmarks that show up in those comparison charts we see so much on social media.

Table comparing Opus 4.5, Sonnet 4.5, Opus 4.1, Gemini 3 Pro and GPT-5.1 across nine benchmarks; the Opus 4.5 column is highlighted and leads in agentic coding at 80.9%, tool use, and ARC-AGI-2 at 37.6%.

Agentic coding — SWE-bench

SWE comes from the term Software Engineer, basically this benchmark evaluates a model’s ability to solve real tasks in Python repositories on GitHub. The model receives the project code, the task, and needs to propose a fix that actually solves the problem.

In this case, in Opus’s evaluation above, there was more or less a subset with ~500 hand-reviewed tasks, with tests and fixes validated by humans to avoid false positives.

The tasks can be broad, but they are usually more or less like this:

  • “This test is breaking, fix the bug.”
  • “This function needs to support a new parameter.”
  • “This API is returning the wrong value, adjust the logic.”

The agent must:

  1. Understand the task description.
  2. Read the code (sometimes several files).
  3. Edit the correct files.
  4. Run tests and make sure everything passes.

When you ask CoPilot, Codex, Claude to execute a task, it needs to execute it and have an acceptable answer. It’s the same here.

A score above 80%, like the one achieved by Opus 4.5, is impressive.
 But there are limitations: the benchmark is focused on Python, depends on fixed repositories, and there is always the risk of “contamination”—the model may already have seen parts of the code during training.

Terminal-Bench 2.0 leaderboard as horizontal bars, with Codex CLI (GPT-5.1-Codex-Max) first at 60.4% and Terminus 2 (Claude Opus 4.5) at 57.8%.

Agentic terminal coding — Terminal-Bench 2.0

This one is very similar to SWE, but the difference is that it runs as a real Linux through the command line. The scenarios range from compiling projects to training ML models, configuring servers, and resolving broken dependencies. For example, the agent can clone a repo, install dependencies, run tests and fix errors, or configure a web server, adjust the firewall, bring up a service.

The agent needs to navigate the terminal (git, pip, sed, systemctl, etc.), edit files with CLI editors and, in the end, leave the environment in the correct state. Basically, it’s what you currently do with Claude Code/Codex Cli.

A high score in Terminal-Bench indicates a model that can act like a “full stack dev Ops” in the terminal, executing end-to-end pipelines. In the same way SWE has its limitations, Terminal Bench cannot capture scenarios using a GUI (Graphical User Interface), and as realistic as the scenarios are, they are still several scenarios in Docker, which does not necessarily reflect the real world.

Figure from the τ²-bench paper: on top, the agent and the user each interact through their own tools and databases; below, a tech-support dialogue where the agent finds that roaming is disabled and turns it on.

Agentic tool use — τ²-bench

τ²-bench (read as “tau-squared”) is a benchmark for conversational agents that use tools in customer service scenarios, mainly in domains such as retail and telecom. In other words, this benchmark is specific to be used as a ChatBot.

This benchmark evaluates the model’s ability to solve billing problems, diagnose connection problems using network tools. In the end, it is important for “call center agents,” corporate chatbots.

But there are important limitations:

  • it does not measure interaction with confused or hostile customers,
  • it does not test voice noise,
  • nor real human ambiguity.

It is a great benchmark for an “ideal chatbot,” not for the chaos of the real world.

Scaled tool use — MCP Atlas

MCP Atlas is a benchmark made by Scale/SEAL to evaluate tool use at scale through the Model Context Protocol (MCP). The idea is to measure how the model performs when it has access to many different tools (databases, APIs, search engines, etc.) and needs to orchestrate calls in complex chains.

It tests scenarios where the model has access to:

  • APIs,
  • databases,
  • search engines,
  • internal systems,
  • complex chains of actions.

The question here is:
Does the model know how to orchestrate several different tools to solve a complex problem?

Since it is recent, there is still no comparison base as solid as in other benchmarks.

OSWorld figure with two sample tasks — updating a bookkeeping sheet from receipts and fixing the code of a snake game — and, below, the diagram of the environment where the agent receives a screenshot and an accessibility tree and responds with mouse and keyboard.

Computer use — OSWorld

OSWorld is a benchmark of computer use: the agent controls a real desktop (or a simulated environment very close to it) with apps such as a browser, text editor, spreadsheet, and system tools. The tasks are real work routines, with multiple steps and windows. Such as:

  • Download a file from the web, organize it into folders, unzip it.
  • Fill in spreadsheets, make charts, and send them by email.
  • Adjust system settings, install programs, etc.

The benchmark measures whether the task was completed and, in newer versions, also efficiency (how many extra steps the agent took compared to a human). The problem is that interfaces can change over time, and there is a finite number of tasks.

ARC-AGI-2 scatter plot crossing score against cost per task; Gemini 3 Deep Think sits alone near 45% at over $50 per task, while Opus 4.5 with 64K thinking reaches 37% at about one dollar.

Novel problem solving — ARC-AGI-2

This is the most “mystical,” the most hyped, the most feared, the great ARC-AGI! ARC-AGI, created by François Chollet, tries to measure abstract reasoning and strong generalization, that is: the model’s ability to infer rules that are not explicit in text, but hidden in visual patterns.

Each problem shows small grids of colored squares (input/output), and the model needs to infer the implicit rule and apply it to new examples. ARC-AGI-2 is the new version, even harder, used today as a “stress test” of systems on the path toward AGI. ARC-AGI-2 is even harder and is used as a stress test of “system intelligence”—the famous “how far are we from AGI?”.

A “trained” human could in theory solve 100% of the ARC-AGI 2 panel, however if a model can solve 45% as was done by the latest version of Gemini 3 Deep Think, that is something far outside the curve.

But it is important to make it clear, the tasks you see in ARC AGI do not usually “happen” in real life. And the use of external tools also impacts the result.

GPQA Diamond bar chart with sixteen models, from Gemini 3 Pro Preview at 90.8% down to MiniMax-M2 at 77.7%.

Graduate-level reasoning — GPQA Diamond

GPQA is a multiple-choice exam in physics, chemistry, and biology at the graduate level.

The Diamond version is the elite of the elite:

  • difficult questions,
  • manually curated,
  • solved only by PhD specialists.

A good score means the model can deal with complex reasoning and advanced content—something like a “good graduate student.”

The problem is that the focus stays only on three areas (physics, chemistry, biology); nothing from the humanities, specific engineering fields, etc. And because it is multiple choice, it limits the model’s “creativity.”

MMMU chart tracking multimodal models from January 2023 to mid-2025, with separate trend lines for open-source and proprietary models and the human-expert bands marked at 76.2%, 82.6% and 88.6%.

Visual reasoning — MMMU

MMMU (Massive Multi-discipline Multimodal Understanding) is a multimodal benchmark that mixes image + text in multiple-choice questions. It has more than 11 thousand questions collected from exams, books, and quizzes at the university level in several areas (art, business, medicine, sciences, engineering, humanities) and several languages.

A good result indicates that the model can reason and retrieve knowledge in several languages, not just literally translate from English, but also understand cultural and contextual nuances.

The problem is that translations can introduce noise, ambiguities, or cultural differences, besides a possibility of data contamination.

Conclusion

Benchmarks are essential to measure the progress of LLMs and understand where each model is strongest.
But there is no single benchmark that indicates “proximity to AGI.”

Each test measures a different piece of intelligence:

  • some focus on practical capability,
  • others on abstract reasoning,
  • others on technical knowledge,
  • others on tool use.

The correct interpretation requires looking at the whole set, not an isolated number on the chart. And it is important to remember, even with all this, it still does not mean AGI is right there, but rather that every time a model gets a higher score than the previous ones, we are one step closer to it. And there are still many steps left, many before we get there.

Agentic coding — SWE-bench (and variants)

Agentic terminal coding — Terminal-Bench 2.0

Agentic tool use — τ²-bench (tau-squared bench)

Scaled tool use — MCP Atlas

Computer use — OSWorld

Novel problem solving — ARC-AGI / ARC Prize

Graduate-level reasoning — GPQA / GPQA Diamond

Visual reasoning — MMMU

Multilingual Q&A — MMMLU / Multilingual MMLU