What an LLM really is — and how to measure it with Artificial Analysis

Translated from the Spanish original. Read in Spanish

There’s a lot of cognitive vocabulary around these systems. Intelligence, reasoning, understanding. And yes, the results are impressive. But the vocabulary isn’t neutral: it sets expectations about how the system works, how it fails and what guarantees it offers. Then you take those expectations to production, and that’s where the problems show up.

Here’s how I understand the mechanism: the basic concepts (LLM, MCP, AGI), what really happens inside a transformer, and what I use to measure instead of argue. For that last part I turn to Artificial Analysis, which is the source I check most often.

Representación conceptual de un cerebro digital con el rótulo AGI LLM Powered, rodeado de nodos que enumeran capacidades como language understanding, reasoning, problem solving y world knowledge


Three terms that aren’t on the same level

LLM, MCP and AGI show up in the same paragraph all the time, and they’re not the same kind of thing. One is a type of model, another is a protocol, the third is a hypothesis with no operational definition.

LLM (Large Language Model). A model pre-trained on large volumes of text with a next-token prediction objective. That pre-training, however, is only part of the process: today’s models also go through supervised fine-tuning, preference optimisation and reinforcement learning. Inside there’s no queryable database of facts and, in the basic architecture, no symbolic inference engine. Knowledge is distributed across the parameters.

MCP (Model Context Protocol). An open standard for connecting models to tools and data sources. An MCP server exposes tools, resources and prompts through a uniform interface, and any compatible client consumes them without specific integration. It replaces one-off integrations with a common connector. What it is not: a requirement for building an agent. An agentic system can work just as well with direct APIs, proprietary function calling or browser automation. MCP brings interoperability, not agency.

AGI (Artificial General Intelligence). A system able to solve arbitrary intellectual tasks with transfer across domains. The problem is that there’s no agreed operational definition and no agreed test. Every organisation has its own, and several are conveniently tied to their own roadmap. Unless you specify the criteria up front, the term adds little.


The mechanism: from tokens to a distribution

Before the steps, what a transformer is. It’s the neural network architecture all these models are built on — it appeared in 2017, in Attention Is All You Need, and displaced the recurrent architectures that processed text position by position. Its contribution is the attention mechanism: each position in the sequence looks at all the others and weighs which ones matter for predicting what comes next. That made it parallelisable, and therefore trainable at today’s scale. The T in GPT stands for transformer.

A transformer’s inference can be described in five steps.

  1. Tokenisation — the text is split into tokens and each one becomes an integer index. The model never sees letters or words: it sees indices.
  2. Embeddings — each index is projected into a high-dimensional vector. That representation isn’t static: in the following layers it’s transformed contextually. Semantics doesn’t live in how close the initial embeddings are, but is distributed across the representation space.
  3. Attention layers — tens or hundreds of layers where each position combines information from the others, through matrix operations on queries, keys and values.
  4. Projection to the vocabulary — the last layer produces a vector of logits, one per token in the vocabulary. A softmax normalises it into a probability distribution.
  5. Decoding — the next token is chosen with a decoding strategy, which can be stochastic (sampling with temperature or top-p) or deterministic. It’s added to the context and the whole process repeats for the next position.

Training meant adjusting parameters by gradient descent to minimise prediction error over a corpus. In the base model there’s no step that checks a fact against a source. When an answer comes out right, it’s because the correct continuation had high probability. And here’s the distinction most often missed: the system you build around it can verify — search, code execution, querying a database. That’s a property of the system, not of the model.

What changes in frontier models

Those five steps are the basic scheme. Frontier models work on the same principles but with architectures considerably more advanced than the 2017 transformer: the dense feed-forward layers are often Mixture of Experts, where a router activates only a subset of parameters per token; classic multi-head attention has given way to variants such as Grouped-Query Attention or Multi-head Latent Attention, to cut the cost on long sequences and the size of the KV cache; and some models interleave state-space model blocks with a few full-attention layers.


What nobody programmed

The result that matters is that a procedure like this produces text that’s coherent at the level of discourse. Valid syntax. Arguments that hold up over several paragraphs. Code that compiles and does what it says.

Nobody encoded grammar rules or discourse-coherence constraints. That emerged as a by-product of optimising a statistical prediction at large scale. And it allows two readings: one about the capacity of a very large statistical approximator, and another about how much structural regularity natural language had — which turned out to be more than we assumed.


Why the vocabulary matters in production

The problem isn’t using functional terms to describe behaviour. The problem is taking for granted properties that fluency doesn’t guarantee: understanding in the human sense, truthfulness, causal traceability, or that a track record of correct answers implies reliable domain knowledge. None of that follows from the mechanism.

I find it more useful to think of them as highly complex probabilistic systems. That characterisation explains the behaviours that shape any production design:

  • Hallucinations are an inherent failure mode, not a bug — the distribution optimises for plausibility of the continuation, not truthfulness. That’s why the result can be convincing and factually wrong even when the relevant information is in the context.
  • Expressed confidence isn’t, on its own, a reliable measure of accuracy — an assertive tone is a stylistic pattern learned from the corpus. Models can be poorly calibrated, although calibration can be measured and improved.
  • Variability depends on decoding — with stochastic decoding, the same input can produce different outputs. In real systems it also varies with the state of the tools or the environment.
  • Performance degrades out of distribution — some kinds of distribution shift produce especially fragile behaviour, although the degradation isn’t always abrupt. This is where guardrails matter most.

Reasoning isn’t a new module

Reasoning models perform substantially better on hard tasks, and from there people jump to assuming they incorporate a different kind of inferential process. It helps to separate three levels: the model’s generation, its training, and the full system around it.

These models are optimised to spend more compute before answering, in several cases through reinforcement learning on reasoning chains. The intermediate sequence is generated with the same autoregressive procedure described above: no logic module or symbolic verifier appears in the model. The full system can incorporate tools, search or verifiers. What we observe is that generating that intermediate text conditions the model on its own output and shifts the final distribution towards regions where correct answers are more likely. In practical terms: more compute at inference time.

And here comes the uncomfortable part for anyone who needs traceability. There’s evidence that the intermediate chain doesn’t always reflect the process that determined the answer — the model can produce a plausible justification that doesn’t match the factors that actually influenced it. The chain is generated text. You shouldn’t automatically take it as a faithful causal record or as a complete execution trace.


Artificial Analysis: measure instead of argue

If these systems’ behaviour is statistical, choosing a model or a provider is a decision that needs data. Artificial Analysis is an independent site that runs its own evaluations on models, providers and agents and publishes the results in a comparable format. These are the sections I use most.

Intelligence Index

A composite index rather than a single benchmark. v4.3 aggregates ten evaluations: AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR. The advantage of a composite is that it reduces the risk of overfitting: optimising for ten heterogeneous evaluations is much harder than for one.

Cost per Task and Time per Task

This is the view that most changed how I compare. Instead of price per million tokens, they calculate how much it costs and how long it takes to solve a complete task from the index, including token consumption — reasoning tokens too, when the model exposes them. The difference is material: a model with a low price per token can have a high cost per task if it burns lots of reasoning tokens. The scatter of Intelligence Index against Cost per Task, with its Pareto frontier, lets you identify the non-dominated options for the quality level you need.

Agentic benchmarks

Beyond the index, they publish specialised evaluations of agentic work. Don’t confuse them with the ten components of v4.3 — some are part of the index and others are published separately. It’s the category that has grown the most, and the most relevant one if your system carries out tasks rather than answering questions:

  • AA-Briefcase — long-horizon knowledge work: producing spreadsheets, presentations and memos within business workflows.
  • GDPval-AA — tasks of real economic value across different occupations, with an Elo anchored to a human baseline of 1,000. Measuring the gap to a person is a demanding benchmark.
  • AutomationBench-AA and EnterpriseOps-Gym-AA — workflows on SaaS applications and business operations.
  • Terminal-Bench — agentic terminal use and end-to-end coding.
  • APEX-Agents-AA — long horizon, where errors accumulating across steps are the main failure factor.
  • ITBench-AA — root-cause analysis of Kubernetes incidents. Especially relevant if you work in IT ops.
  • AA-AnalystAgent — quantitative analysis on spreadsheets and documents, reported as pass^5: the share of tasks solved correctly in all five runs. That’s reproducibility, not peak capability. It’s useful for judging how robust an automation is, although on its own it doesn’t tell you whether you can remove human oversight.

Coding Agent Index

It’s kept separate from the model rankings, and rightly so: an agent is model + scaffold + tools, and the scaffold matters. v1.5 combines DeepSWE, Terminal-Bench and SWE-Atlas-QnA, and also reports cost per task and run time. There’s also a comparison table of general-work agents — Claude Cowork, ChatGPT Work, Microsoft Copilot Cowork, Manus, OpenClaw, Hermes Agent, Gemini Enterprise, among others — with platform, whether it’s open source, whether it supports bring-your-own-model, local file access, browser automation and price.

Maximiliano Díaz Doglia

AI Platform Engineer & Full-Stack Developer
Building Enterprise Integrations & Automations

Published in: AI