A ranking is a snapshot.
Model version, harness, tools and compute budget must accompany every result.
Models · Benchmarks · Use selection
Rankings age quickly. Model selection should begin with the task, unacceptable errors, data, required control and the total cost of use.

This world seems to rest entirely on a monumental support, yet its equilibrium depends on what carries it. The work reminds us that a model is never a sufficient foundation: reliable use also depends on data, tools, oversight and work organisation.
Model version, harness, tools and compute budget must accompany every result.
Strong general performance does not guarantee quality in occupational health or French law.
A new version, prompt, corpus or tool means the earlier result is no longer enough.
01 / Three practical families
The commercial name matters less than how the system is hosted, connected to data and authorised to act.
Simple access, broad capability and integrated tools, balanced against provider dependence, forced version changes and contractual conditions.
Examine: data, retention, location, subprocessors, stability and reversibility.Greater control over hosting and updates, while the organisation takes on infrastructure, security, evaluation and maintenance.
Examine: licence, internal expertise, operating cost, patches and monitoring.The model is connected to a corpus, rules or tools. Performance then depends as much on this chain as on the base model.
Examine: sources, permissions, logs, tool errors and stopping conditions.02 / Selection criteria
A useful comparison tests the exact system under the actual intended conditions.
Summarise, extract, draft, classify, search or act: requirements and critical errors differ.
Omissions, false sources, discrimination, disclosure and unauthorised action require separate tests.
Information type, legal basis, retention, reuse, transfers and access control.
Pinned version, override, traceability, tool permissions, human review and immediate stop.
Service price, integration, latency, human checking, incidents, maintenance and skill retention.
Time actually saved, checking burden, autonomy, cooperation and allocation of accountability.
03 / Benchmarks
Results are comparable only when versions, tools, prompts, budgets and numbers of attempts are sufficiently similar.
| Reference | What it tests | Workplace limitation |
|---|---|---|
| SWE-bench Verified ↗ | Resolution of real software engineering issues in a tool-enabled environment. | Does not assess clinical quality, law or effects on work. |
| GPQA Diamond ↗ | Difficult scientific questions designed to resist simple search. | Knowledge scores do not establish reliability on internal documents. |
| MMMU ↗ | Multimodal reasoning over text, tables, diagrams and images. | Does not guarantee correct reading of your forms, scans or procedures. |
| LM Arena ↗ | Human preferences between responses to varied prompts. | Preference does not directly measure truthfulness, privacy or compliance. |
| METR ↗ | Agents completing software tasks of different durations. | Longer technical autonomy does not show that workplace deployment is safe or desirable. |
04 / Minimum protocol
Without this record, a score cannot be seriously reviewed or compared after an update.
05 / Follow the landscape
These sources help identify changes. They do not replace local evaluation on the intended task and data.