Skip to main content
Abstract AI evaluation course combining document, coding and agent task stations

Artificial Analysis v4.2 shifts AI benchmarking toward private agent and document tests

The updated composite index adds agentic knowledge work and long-document reasoning while doubling the weight of held-out test sets.

Published

06 Sep 2026

Reading Time

3 min read

Share this article:

Contents

What changed in version 4.2

Artificial Analysis announced Intelligence Index v4.2 on September 4, 2026, updating the composite benchmark it uses to compare language models across agentic work, coding, scientific reasoning and general knowledge. The release adds two evaluations, removes one saturated test and changes grading infrastructure.

The first addition is AA-Briefcase, a private evaluation of multi-step knowledge work. It asks models to work through linked tasks and large collections of source files, then combines rubric scoring with pairwise comparisons of task success, analysis and presentation. The second is GDP.pdf, created by Surge AI, which tests reasoning across 100 professional PDFs spanning ten domains and 4,592 pages.

Artificial Analysis removed GPQA Diamond because it considers that scientific reasoning benchmark saturated. It also updated answer keys and grading prompts for AA-LCR, re-anchored Elo ratings for two agent evaluations and changed SciCode sandboxes so slow but correct programs are less likely to be counted as failures.

More of the benchmark is kept private

The most important structural change is that 40 percent of the index weighting now comes from private held-out test sets, twice the share in version 4.1. The organization says this should reduce the opportunity for model developers to optimize directly against known questions.

Private tests address one familiar benchmark problem, but they introduce another trade-off: outsiders cannot inspect every item or reproduce every score. Users therefore have to evaluate both contamination resistance and transparency. Public methodology, stable test procedures and clear version labels become especially important when the underlying questions are not disclosed.

The index now combines ten evaluations. Its category weights are 30 percent for agents, 20 percent for coding, 20 percent for scientific reasoning and 30 percent for general capabilities. That weighting is an editorial and methodological choice; it makes the result more sensitive to agentic knowledge work than a benchmark aimed only at short-form question answering.

Why version changes complicate comparisons

A score from v4.2 should not be treated as directly interchangeable with a score from an earlier index. Adding tasks, changing weights and re-anchoring graders can move relative positions even if the underlying models have not changed. For procurement or research, the useful comparison is between models evaluated under the same version and settings.

Artificial Analysis reports that Anthropic and OpenAI lead the new composite index at launch, while four laboratories occupy its cost-per-task Pareto frontier. Those results are measurements by Artificial Analysis under its own methodology, not universal rankings for every use case. A model can perform well on the composite and still be a poor fit for a particular language, latency requirement, domain or tool environment.

What teams should do with the index

Use v4.2 as one signal in a broader evaluation. Check per-evaluation results, token use, cost and execution time instead of relying only on the aggregate score. Then run a smaller internal test set that reflects the actual documents, tools, failure costs and languages of the intended workload.

The update makes the index harder and more agent-focused. Its value will depend on whether those revised tasks predict success in the work teams genuinely need models to complete.

Sources

Tags:

#AI benchmarks #Artificial Analysis #AI agents #long context #model evaluation #held-out tests

23

views

0

shares

0

likes

Related Articles