Models · Benchmarks · Use selection

A useful landscape, not a podium.

Rankings age quickly. Model selection should begin with the task, unacceptable errors, data, required control and the total cost of use.

Dated viewNo universal best modelLocal testing required
What supports the whole
Thomas Cole’s The Titan’s Goblet, a landscape built in and around a giant goblet
Thomas Cole
The Titan’s Goblet · 1833

This world seems to rest entirely on a monumental support, yet its equilibrium depends on what carries it. The work reminds us that a model is never a sufficient foundation: reliable use also depends on data, tools, oversight and work organisation.

DATE

A ranking is a snapshot.

Model version, harness, tools and compute budget must accompany every result.

TASK

The best model is contextual.

Strong general performance does not guarantee quality in occupational health or French law.

CYCLE

An update requires re-testing.

A new version, prompt, corpus or tool means the earlier result is no longer enough.

01 / Three practical families

Begin with the system architecture.

The commercial name matters less than how the system is hosted, connected to data and authorised to act.

01 · HOSTED

Proprietary general-purpose model.

Simple access, broad capability and integrated tools, balanced against provider dependence, forced version changes and contractual conditions.

Examine: data, retention, location, subprocessors, stability and reversibility.
02 · OPEN WEIGHTS

Model operated under local control.

Greater control over hosting and updates, while the organisation takes on infrastructure, security, evaluation and maintenance.

Examine: licence, internal expertise, operating cost, patches and monitoring.
03 · COMPOSED

Specialised system, RAG or agent.

The model is connected to a corpus, rules or tools. Performance then depends as much on this chain as on the base model.

Examine: sources, permissions, logs, tool errors and stopping conditions.

02 / Selection criteria

Six questions before the headline score.

A useful comparison tests the exact system under the actual intended conditions.

01

Exact task.

Summarise, extract, draft, classify, search or act: requirements and critical errors differ.

02

Unacceptable errors.

Omissions, false sources, discrimination, disclosure and unauthorised action require separate tests.

03

Data.

Information type, legal basis, retention, reuse, transfers and access control.

04

Control.

Pinned version, override, traceability, tool permissions, human review and immediate stop.

05

Total cost.

Service price, integration, latency, human checking, incidents, maintenance and skill retention.

06

Effects on work.

Time actually saved, checking burden, autonomy, cooperation and allocation of accountability.

03 / Benchmarks

What they measure — and what they do not.

Results are comparable only when versions, tools, prompts, budgets and numbers of attempts are sufficiently similar.

ReferenceWhat it testsWorkplace limitation
SWE-bench Verified ↗Resolution of real software engineering issues in a tool-enabled environment.Does not assess clinical quality, law or effects on work.
GPQA Diamond ↗Difficult scientific questions designed to resist simple search.Knowledge scores do not establish reliability on internal documents.
MMMU ↗Multimodal reasoning over text, tables, diagrams and images.Does not guarantee correct reading of your forms, scans or procedures.
LM Arena ↗Human preferences between responses to varied prompts.Preference does not directly measure truthfulness, privacy or compliance.
METR ↗Agents completing software tasks of different durations.Longer technical autonomy does not show that workplace deployment is safe or desirable.

04 / Minimum protocol

Five items to preserve with every result.

Without this record, a score cannot be seriously reviewed or compared after an update.

01Test date
02Provider and exact version
03Prompt, corpus and tools
04Budget and number of attempts
05Errors and excluded cases

05 / Follow the landscape

Sources to consult with a date.

These sources help identify changes. They do not replace local evaluation on the intended task and data.