Understand · Essential concepts

Understanding artificial intelligence at work.

Distinguish between systems, understand how a language model is built and produces a response, interpret its performance without reading too much into it, then explore its risks and the goal of AGI.

25-minute readSelected and dated sourcesUpdated in September 2026
01

“AI” does not refer to a single tool.

Generating text, assigning a score and allocating tasks are three different functions, with different risks.

02

An LLM produces a plausible continuation.

It calculates probabilities from the context; it does not automatically check whether its response is true or complies with the rules.

03

Real work is the right level of analysis.

The model matters, but so do the task, the data, the people, the decision rules and the organisation.

01 / Identify the system

First, what kind of artificial intelligence are we talking about?

The term “AI” covers systems that do not produce the same kind of result. To understand a workplace use, start by identifying what the tool actually does.

01 · Produce

Generative AI

It creates or transforms text, images, sound or code from an instruction. A conversational assistant based on an LLM belongs to this family.

OutputNew content
Watch forAccuracy, sources, data entered
02 · Estimate

Predictive AI

It estimates a probability, assigns a score or classifies a case based on past data: default risk, case priority or diagnostic support.

OutputA score or category
Watch forBias, thresholds, false positives and false negatives
03 · Organise

Algorithmic management

It allocates tasks, sets processing priorities, measures activity or influences a management decision. It may or may not incorporate generative AI.

OutputAn instruction or decision
Watch forAutonomy, avenues of appeal, work intensification

02 / Inside an LLM

How a language model builds a response.

A large language model — or LLM — is a neural network trained to predict the continuation of a text. It does not look up a ready-made answer in a database: it composes one, unit by unit.

01 · Split

Text becomes tokens

The instruction and supplied documents are split into small units: words, parts of words or symbols.

02 · Represent

Tokens become numbers

Each token is converted into a numerical representation that encodes patterns learned during training.

03 · Connect

Attention weighs the context

The Transformer examines relationships between tokens to determine which parts of the context matter most at that moment.

04 · Predict

The next token is selected

The model calculates a probability distribution. Generation settings influence the choice and its variability.

05 · Repeat

The response takes shape step by step

The selected token joins the context, then the calculation starts again until the response is complete.

Figure 1 · Illustrative example

At each step, every possible continuation is assigned a probability.

Text already in the context“The worker describes persistent pain in the ”

Low temperature

“back” is selected almost every time: consistent, sometimes repetitive responses.

High temperature

“lower” or “wrist” appear more often: more variety, more variation.

Then the process repeats

“…in the back ” → “,” → “aggravated” → “by”… each added token starts the calculation again.

Illustrative example: the probabilities are fictional. A token is not always a whole word (“lower” may begin “lower back”). The model selects what is likely in this context, not what is true for this worker.

03 / Training

How a model is built, then refined.

An LLM’s capabilities are not programmed rule by rule. They develop through several stages of training, each of which leaves its mark on the final behaviour, including its shortcomings.

Figure 2 · Diagram

From vast amounts of text to a deployed model: five stages.

  1. 01Pre-train

    ReceivesWeb text, books, code: trillions of tokens
    DoesPredict the next token, compare it with the actual continuation, adjust the parameters. Billions of times.predictcomparecorrect
    ProducesBase modelCompletes text without following instructions. Knowledge frozen at a cut-off date.
  2. 02Fine-tune

    ReceivesExample dialogues written or validated by people
    DoesImitate these model answers: follow an instruction, a tone or a format.
    ProducesAssistantResponds to requests and refuses some of them.
  3. 03Reinforce

    ReceivesResponses compared by people, or by an AI guided by written principles
    DoesA reward model scores each response; the assistant adjusts to achieve a better score.
    ProducesAligned assistantMore helpful and more cautious, but pleasing the user can take precedence over accuracy.
  4. 04Reason

    ReceivesProblems with verifiable solutions: mathematics, code, automated tests
    DoesTry many approaches, check the result, reinforce reasoning that succeeds.
    ProducesReasoning modelSpends more computation before answering, through intermediate steps.
  5. 05Evaluate, deploy

    ReceivesCapability and risk tests, teams tasked with attacking the model
    DoesMeasure, add safeguards, decide whether to release.
    ProducesDeployed modelUpdated and monitored: each version is tied to a date.
Order of magnitudeLlama 3.1 405B (Meta, 2024): 15.6 trillion tokens, 31 million GPU hours, up to 16,000 chips in parallel.
Stages 02 to 05Far less data, but more expensive: each example is written, compared or checked. The share of computation devoted to reinforcement has grown since reasoning models emerged.

↺ Shortcomings identified in testing or after deployment inform the training of the next version.

Simplified diagram: the order and relative importance of the stages vary between labs, and some are repeated. Sources: Ouyang et al., 2022 (fine-tuning and reinforcement through human feedback); Bai et al., 2022 (feedback from an AI guided by principles); Meta, “The Llama 3 Herd of Models”, 2024 (orders of magnitude).
Scale has long been the main driver of progress.

Performance improves fairly predictably with the amount of computation, data and parameters: these are the “scaling laws”. According to Epoch AI, the computation used to train notable AI models has grown by about 4.5 times each year since 2010, and by about 5 times for frontier large language models since 2020.

A second route: computation at response time.

Letting a model work longer before answering improves results in mathematics, programming and science. The International AI Safety Report 2026 identifies this as one of the main sources of recent progress.

Figure 3 · Data

Training computation has increased more than two billionfold in fourteen years.

10¹⁷10¹⁹10²¹10²³10²⁵10²⁷20122014201620182020202220242026Training FLOP · logarithmic scale, each grid line = × 100Two billionfold growth in 14 yearsfrom AlexNet (2012) to GPT-6 Astra (2026)AlexNetTransformerGPT-2GPT-3PaLMGPT-4Llama 3.1 405BGPT-4.5Grok 4GPT-6 Astrapublished valueEpoch AI estimate
Source: Epoch AI, “AI Models” database, updated on 26 September 2026. Providers no longer publish the computation used for their flagship models: recent values are estimates (GPT-6 Astra: at least 100,000 GB200 chips according to Epoch). No estimate has yet been published for Claude Opus 5.5, Fable 5.1 or Gemini 3.x. A FLOP is an elementary computational operation.
View the data
ModelDateComputation (FLOP)Status
AlexNetSept. 20124.7 × 10¹⁷Published
TransformerJune 20177.4 × 10¹⁸Published
GPT-2Feb. 20191.9 × 10²¹Estimate
GPT-3May 20203.1 × 10²³Published
PaLMApr. 20222.5 × 10²⁴Published
GPT-4Mar. 20232.1 × 10²⁵Estimate
Llama 3.1 405BJuly 20243.8 × 10²⁵Published
GPT-4.5Feb. 20253.8 × 10²⁶Estimate
Grok 4July 20255.0 × 10²⁶Estimate
GPT-6 AstraSept. 20261.0 × 10²⁷Estimate

04 / From model to system

The tool you use is never just the model.

A workplace assistant combines a model, instructions, documents, tools and sometimes memory. Each layer changes the result, and therefore what needs checking before you entrust it with a task.

Figure 4 · Diagram

What happens between your question and the response.

Youquestion, filesWeb page, emailor malicious PDFhidden instructionContext windoweverything the model “sees”System instructionConversationSupplied documentsRetrieved passagesTool resultsMemory (sometimes)Language modelpredicts the continuation, token by tokenDocument storeToolsWeb searchCode executionEmail, calendarWorkplace softwarerequestread in full for each tokenresponsepassages (RAG)calls a toolhuman approval?result↻ agent loopact, observe, repeat
The model only knows what enters its context window. Each arrow is a possible checkpoint: who writes the system instruction, which documents are retrieved, which tools can act and with what approval. The red line shows prompt injection: content read by the tool can be taken as an order.
01 · Instructions

An invisible system instruction

Before your first question, the provider or organisation has set a role, a tone, restrictions and formatting rules. Two tools based on the same model can therefore respond very differently.

ProvidesInstructions tailored to the profession
CheckWho writes these instructions and who can change them
02 · Context

Limited working memory

The model only “sees” its context window: the conversation, supplied documents and passages retrieved from a document store (RAG). Anything outside it does not exist for the model, and very long documents may be used unevenly.

ProvidesResponses grounded in your sources
CheckThe passages actually retrieved and cited
03 · Tools

Actions in other software

Web search, code execution, email, calendars or case records: the model can call tools through connectors, such as those using the MCP protocol. It no longer just makes suggestions; it acts.

ProvidesTasks completed from start to finish
CheckPermissions granted, action by action
04 · Agents

A loop pursuing a goal

An agent plans, acts, observes the result and starts again, sometimes for hours, with intermittent human supervision. An early mistake can propagate through the entire chain of actions.

ProvidesAutomation of whole sequences
CheckStop points, action logs, rollback
05 · Hosting

Closed or open models

Closed models, such as GPT, Claude or Gemini, can only be used remotely on the provider’s infrastructure. Open-weight models, such as Llama, Qwen, DeepSeek or some Mistral models, can be installed on servers under your control, generally with a slight lag behind the best closed models.

ProvidesControl over where data goes
CheckLocation, HDS certification for health data hosting in France, reuse for training
06 · Multimodality

Read, see, listen, transcribe

Recent models also process images, scanned documents, audio and video. Automatic transcription and summarisation of consultations, sometimes called medical “scribes”, are their most visible use in healthcare.

ProvidesLess data entry, usable non-text documents
CheckTranscription errors, informing the people being recorded

05 / Capabilities and limits

What language models can do — and what still needs checking.

Capabilities depend on the model, its version, the language, the supplied documents, the tools it can access and how the task is phrased.

Task typeWhat the model can contributeWhat cannot be assumed
01Reading and writingSummarise, rephrase, translate, extract information, compare documents or draft an initial text.Completeness, faithfulness to the source, up-to-date information and suitability for the professional context.
02Reasoning through a problemBreak down a question, suggest hypotheses, apply a described procedure and explain the steps in a calculation.Consistent accuracy. A convincing explanation can accompany a wrong answer.
03Writing code and processing dataGenerate, explain or correct code; prepare a query or an exploratory analysis.Security, robustness and correct operation on untested cases. Running the code and testing it remain essential.
04Using toolsWith a suitable application, search the web, consult a document store, call software or carry out several actions in sequence.Source quality, permission to act and control over consequences. Access rights must be limited.
05Processing multiple mediaDepending on the model: describe an image, transcribe audio, analyse a document or answer based on a video.Complete perception of details, understanding of the situation and compliance with profession-specific rules.
Demonstrated capability does not guarantee reliability.

Frequent success on a category of tasks does not make every response predictable. You need to test the errors that matter for your use.

A model version is tied to a date.

The provider may change the model, its settings, filters or tools. A useful evaluation records the version and the date.

Why some errors persist.

These limits are not simply a matter of settings. They arise from how models are trained, evaluated and integrated into tools. Understanding them helps place checks where they are needed.

01 · Hallucination

A plausible response rather than “I don’t know”

Pre-training does not teach the distinction between a rare fact and a plausible invention. Evaluations that score only accuracy then reward guessing rather than abstaining.

At workInvented references, legal provisions or figures, presented with confidence.

Kalai et al., OpenAI, 2025 ↗
02 · Sycophancy

Agreeing with the person asking

Fine-tuning on human preferences can favour agreement over accuracy. In April 2025, OpenAI withdrew a GPT-4o update considered excessively flattering.

At workThe tool tends to confirm the hypothesis embedded in the question.

OpenAI, 2025 ↗
03 · Jagged frontier

Excellent here, failing right next door

Capabilities do not follow human perceptions of difficulty. Among 758 consultants, AI improved speed and quality on tasks within its reach, but reduced answer accuracy on a task just beyond it.

At workTrust earned on one task does not carry over to the next.

Dell’Acqua et al., Harvard Business School, 2023 ↗
04 · Displayed reasoning

The explanation is not the actual computation

The reasoning a model displays does not faithfully reflect what determined its answer. In an Anthropic study, the models tested disclosed a hint they had used in only 25 to 39% of cases.

At workA convincing justification does not prove the answer is well founded.

Chen et al., Anthropic, 2025 ↗
05 · Prompt injection

A document can give orders

The model does not reliably separate the user’s instructions from the text it reads: a web page, email or PDF may contain hidden instructions. This is the top risk in OWASP’s ranking for LLM applications.

At workThe risk increases as soon as the tool can send, edit or delete.

OWASP, LLM01:2025 ↗
06 · Variability

Same question, different answer

Generation is probabilistic and sensitive to wording, document order and model version. Its knowledge stops at a cut-off date, unless it can consult live sources.

At workTest repeatedly, date the trials and record the version used.

Figure 5 · Diagram and data

A jagged frontier: difficulty for humans does not predict machine failure.

what we observewhat we imaginelevel of reliabilityrequired by the tasksimple task failedcomplex task completedTasks, from easiest to hardest for a human →AI performance →
Tasks within the frontier
+12%more tasks completed
+25%faster completion
+40%higher rated quality, at least
Task outside the frontier
−19 ptsin correct answers
Top: illustrative diagram. Bottom: results from the trial by Dell’Acqua et al. (Harvard Business School, BCG) with 758 consultants, compared with a group without AI: red points show tasks completed to the required standard; grey points show tasks where AI falls short. On the task outside the frontier, AI-assisted consultants gave the correct answer less often.

06 / Reading benchmarks

There is no universal ranking of the “best model”.

A benchmark measures performance on a set of tasks, with a specific instruction, metric and conditions. Change one of these elements and the ranking may change.

Live links rather than already outdated figures.

The pages below are maintained by their authors. They provide the latest available results without fixing a ranking here that would quickly become outdated. The snapshot below simply shows that the leader changes depending on the test.

Figure 6 · Dated snapshot

Who leads? It depends on the test.

GPT-6 Astra (OpenAI, 3 Sept. 2026)Claude Opus 5.5 (Anthropic, 22 Sept. 2026)

Capability index (ECI) ↗

Epoch AI · combines several dozen tests

  1. 1GPT-6 Astra166.6
  2. 2Claude Fable 5.1165.0
  3. 3Claude Fable 5163.6
  4. 4Claude Opus 5162.7
  5. 5GPT-5.5 Pro162.5

Claude Opus 5.5 not yet ranked

Professional tasks (GDPval-AA) ↗

Artificial Analysis · Elo score, pairwise comparisons

  1. 1Claude Opus 5.51 846
  2. 2Claude Fable 5.11 735
  3. 3Claude Opus 51 708
  4. 4Grok 4.71 695
  5. …GPT-6 Astra1 542

Abstract reasoning (ARC-AGI-2) ↗

ARC Prize · verified scores

  1. 1GPT-6 Astra95.0%
  2. 2Claude Opus 5.593.3%
  3. 3GPT-5.692.5%
Snapshot as of 27 September 2026, soon outdated: follow the links for today’s results. Gaps between the leaders are often smaller than the measurement uncertainty. None of these tests tells you which model suits your use: that must be checked on your own tasks.
01 · TaskDoes the test really resemble the work the tool will have to do?
02 · ProtocolWhich instruction, which tools, how many attempts and which model version?
03 · MeasurementDoes the score measure accuracy, preference, cost, speed or something else?
04 · WeaknessesCould the data have been seen during training, and which errors does the average hide?

07 / Tracking progress

Measure the pace of progress rather than taking a snapshot of a ranking.

To anticipate effects on work, the useful question is not “which model is best?” but “what is becoming feasible, how quickly and under what conditions?”. Four dated reference points, with their limits.

Figure 7 · Data

Tasks an agent can complete alone have grown from seconds to hours.

above 16 h: unreliable measurements1 s1 min10 min1 h8 h8 h: one working day20192020202120222023202420252026Task duration for a human expert · logarithmic scaleGPT-2 · 3 sGPT-3 · 8 sGPT-3.5 · 36 sGPT-4 · 4 mino1 · 39 minGPT-5 · 3 h 23Claude Opus 4.6 · 11 h 59Claude Mythos Preview · 17 h 25GPT-5.6 SolDoubling roughly every 7 monthsfrom 2019 to 2025, and roughly every 3 months since 2024
Source: METR, Time Horizon 1.1, May 2026 data (frontier models at release), supplemented by METR’s report on GPT-5.6 Sol (June 2026, grey point). How to read this: GPT-5 succeeds half the time on tasks that take an expert around 3 h 20. As of 27 September 2026, METR has published no measurements for Claude Opus 5.5, Fable 5.1 or GPT-6 Astra: the most advanced models already exceed what its tasks can measure. Mostly programming tasks; wide confidence intervals.
View the data
ModelRelease50% horizon95% CI
GPT-2Feb. 20193 s1 s – 8 s
GPT-3May 20208 s5 s – 13 s
GPT-3.5Mar. 202236 s16 s – 1 min
GPT-4Mar. 20234 min2 min – 8 min
GPT-4 (Nov.)Nov. 20234 min2 min – 8 min
GPT-4oMay 20247 min4 min – 13 min
Claude 3.5 Sonnet (June)June 202411 min5 min – 22 min
o1-previewSept. 202420 min12 min – 33 min
Claude 3.5 Sonnet (Oct.)Oct. 202421 min10 min – 41 min
o1Dec. 202439 min21 min – 1 h 05
Claude 3.7 SonnetFeb. 20251 h 0033 min – 1 h 44
o3Apr. 20252 h 001 h 15 – 3 h 11
GPT-5Aug. 20253 h 231 h 53 – 6 h 46
Gemini 3 ProNov. 20253 h 442 h 20 – 6 h 19
Claude Opus 4.5Nov. 20254 h 532 h 42 – 10 h 24
GPT-5.2Dec. 20255 h 523 h 18 – 13 h 35
Claude Opus 4.6Feb. 202611 h 595 h 17 – 60 h 34
Claude Mythos PreviewApr. 202617 h 258 h 29 – 55 h 04
GPT-5.6 Sol (not in frontier-model series)June 2026≈ 11 h 185 h 00 – 40 h 00
01Agent autonomyMETR · Time horizonTasks entrusted to agents are getting longer, fastMETR measures, in an expert’s working time, the duration of tasks an agent completes successfully half the time. It doubled roughly every seven months from 2019 to 2025, and roughly every three months since 2024. Mostly computing tasks; measurements are unreliable above 16 hours, a threshold already reached in spring 2026. No measurements published to date for Claude Opus 5.5 or GPT-6 Astra.Follow the curve ↗ 02Professional tasksOpenAI · GDPvalDeliverables compared blindly with experts’ work1,320 tasks drawn from 44 professions, judged by experienced professionals. The best model was rated equal to or better than the expert in 47.6% of cases in September 2025, then 84.9% in April 2026 according to OpenAI. Standalone tasks, without the exchanges or ambiguities of real work; figures published by the test’s designer.View the evaluation ↗ 03AdaptationARC Prize · ARC-AGI-3Facing the unfamiliar, a gap closes in monthsLaunched in March 2026, this test places the system in new interactive environments without instructions: it must explore, understand the rules and reach the goal. Humans solve them all; the best systems scored below 1% at launch, then 62.7% in September 2026 with GPT-6 Astra, under the standard protocol verified by ARC Prize.View the benchmark ↗ 04Real workMETR · Randomised trialPerceived gains are not measured gainsIn early 2025, experienced developers believed they were 20% faster with AI; measurements showed they were 19% slower. A further measurement in late 2025 instead suggests a speed-up, but METR considers it inconclusive because of selection bias. In May 2026, a self-report survey showed a median gain of ×3, which METR considers probably overstated.Read the update ↗
Figure 8 · Data

Two tests considered difficult, surpassed in a matter of months.

GDPval · professional tasksdeliverable rated equal to or better than an expert’s0%25%50%75%100%parity with the expertSept. 2025Mar. 2026Sept. 2026Claude Opus 4.1 · 47.6%GPT-5.2 Thinking · 70.9%GPT-5.5 · 84.9%
ARC-AGI-3 · adapting to the unfamiliarlevels solved, standard protocol (humans: 100%)0%25%50%75%100%humansSept. 2025Mar. 2026Sept. 2026at launch · 0.5%Claude Opus 5 · 30.2%GPT-6 Astra · 62.7%
Sources: GDPval, figures published by OpenAI, the test’s designer (September 2025 to April 2026); no comparable figures have been published for GPT-6 Astra or Claude Opus 5.5. ARC-AGI-3, scores verified by ARC Prize, September 2026; with a setup that retains its reasoning between moves, GPT-6 Astra reaches around 99%: the score depends on the protocol. Red line: best score to date; grey points: other models.
Figure 9 · Data

Expected time savings with AI, then measured.

Source: Becker, Rush, Barnes et al., METR, 2025. Randomised trial: 16 experienced developers, 246 real tasks on their own projects, early 2025. A further measurement in late 2025 instead suggests a speed-up, considered inconclusive by METR.
Rapid but uneven progress.

Systems capable of solving expert-level problems fail at simple perception or adaptation tasks. The International AI Safety Report (February 2026) considers both a plateau and a sharp acceleration plausible by 2030.

Completing a task is not automating a profession.

A profession combines tasks, relationships, responsibilities and unexpected events. These reference points indicate which parts of work are becoming automatable, not what work will become: that also depends on organisational choices.

08 / System risks

Beyond individual errors, three families of risk.

The International AI Safety Report, written by more than 100 experts under the chairmanship of Yoshua Bengio, distinguishes malicious uses, malfunctions and systemic risks. Effects on work and health are detailed in the Risks section.

01 · Malicious uses

When the tool is used to cause harm

Fraud and scams, deepfakes, manipulation, assistance with cyberattacks. Labs also evaluate the help their models could provide in designing biological or chemical weapons.

At workTargeted phishing, impersonation of an executive, reputational harm.

02 · Malfunctions

When the system does something other than intended

Fabricated content, faulty code, agents that fail or subvert their assigned objective. Under test conditions, some models have exploited weaknesses in their scoring system or adopted deceptive strategies to achieve a goal.

At workUnwanted agent actions, human oversight reduced to rubber-stamping.

03 · Systemic risks

When effects arise from widespread adoption

Automation of cognitive tasks, weakened critical thinking through excessive trust in machines, emotional dependence on conversational companions used by tens of millions of people.

At workDeskilling, work intensification, loss of autonomy and meaning in work.

A blind spot

Tests conducted before release are poor predictors of usefulness and risks in real conditions, and some models now detect when they are being evaluated. Post-deployment monitoring, including in workplaces, is becoming essential.

Improving safeguards, without guarantees.

Major labs publish safety frameworks, such as Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework, linking capability thresholds to stronger measures. Interpretability research seeks to understand models’ internal mechanisms rather than just their responses. These commitments remain largely voluntary and assessed by the companies themselves.

An initial binding framework in Europe.

Since 2 August 2025, the AI Act has required providers of general-purpose AI models to maintain technical documentation and a copyright compliance policy. Models posing systemic risk must also be evaluated, undergo risk management and have serious incidents reported. Explore Governance →

09 / The goal of AGI

AGI: a stated goal of AI labs and a concept still contested.

Several companies say they aim to build artificial general intelligence, or AGI. There is neither a shared definition nor a test that could establish its arrival. Three ways of defining it shed light on what is at stake for work.

01 · Through work

Outperforming humans in most jobs

OpenAI’s charter (2018) aims for “highly autonomous systems” that outperform humans at most economically valuable work. The definition is explicitly economic: it is measured in terms of replacing human labour.

OpenAI Charter, 2018 ↗
02 · Through levels

A progression rather than a threshold

Google DeepMind combines the level of performance, from “emerging” to “superhuman”, with the breadth of tasks covered. The degree of autonomy granted to the system is treated separately: it is a deployment choice, not just a capability.

Morris et al., Levels of AGI, 2023 ↗
03 · Through cognition

Comparison with an educated adult’s abilities

A group of researchers measures models across ten broad human cognitive abilities (the CHC model). GPT-4 scored 27% and GPT-5 57%, with a highly uneven profile: extensive knowledge, almost no long-term memory.

Hendrycks, Bengio et al., 2025 ↗
Figure 10 · Interpretive framework

Levels of AGI: combining performance and breadth of tasks.

Level of performance
Narrow AI · one task
General AI · many tasks
0 · No AI
Narrow AICalculator, compiler
General AIHuman work assisted by a platform (Mechanical Turk)
1 · Emergingequal to or slightly better than an unskilled human
Narrow AISimple rule-based systems
General AIChatGPT, Bard, Llama 2, Gemini
2 · Competentat least the median of skilled adults
Narrow AIVoice assistants, toxic-content detectors
General AINot yet achieved
3 · Expertat least 90% of skilled adults
Narrow AIStyle checkers, image generators
General AINot yet achieved
4 · Virtuosoat least 99% of skilled adults
Narrow AIDeep Blue (chess), AlphaGo
General AINot yet achieved
5 · Superhumanbetter than all humans
Narrow AIAlphaFold, AlphaZero, Stockfish
General AISuperintelligence: not yet achieved
Source: Morris et al., Google DeepMind, “Levels of AGI”, 2023, the authors’ examples at that date. More recent models have not been officially reclassified; the authors already note that the best models reach the “competent” level on certain tasks, such as short-form writing or simple code.
Figure 11 · Data

An uneven profile: strong knowledge, no lasting memory.

GPT-4 · 27%GPT-5 · 57%each ability is scored out of 10; 100% = educated adult
Source: Hendrycks, Bengio et al., “A Definition of AGI”, 2025, Table 1, based on the CHC model of human cognitive abilities. Long-term storage remains at zero: the model learns nothing permanently from one conversation to the next, unless the tool adds external memory. As of 27 September 2026, no more recent model has been evaluated using this framework.
A contractual term as much as a scientific one.

In the contract between Microsoft and OpenAI, declaring AGI had financial consequences; verification was entrusted to a panel of independent experts in October 2025, before an April 2026 amendment largely stripped the clause of its effect by separating payments from technological progress.

Announcements are not measurements.

Several lab leaders publicly suggest timelines of a few years. These statements come from participants in a race for funding; measured, dated indicators are more informative than announced dates.

Arguments for a rapid arrival

Why some anticipate acceleration

  • Measured capabilities have advanced steadily for several years
  • Tests designed to withstand models are quickly overcome: ARC-AGI-3 went from below 1% to 63% in six months
  • Training computation and investment continue to increase sharply
  • Computation at response time opens a new route to progress
  • Models already contribute to programming and AI research, which could accelerate what comes next
Arguments for caution

Why others remain sceptical

  • Continual learning and lasting memory remain very limited
  • Scores depend on the protocol: on ARC-AGI-3, the same model scores 63% or 99% depending on the test setup
  • Reliability on long tasks improves more slowly than demonstrations suggest
  • Energy, data and capital could constrain further scaling
  • Passing tests does not guarantee the ability to hold a job in a real environment
For occupational health

The issue is not to bet on a date. It is to track, task by task, what is becoming automatable, how quickly and under what oversight, and to anticipate effects on workload, autonomy, skills and employment before they take hold. Explore the economic and social issues →

10 / Back to real work

A good technical score does not, on its own, predict good workplace use.

A benchmark describes a model’s behaviour in a test. Deployment also needs to be evaluated in the actual task, with its constraints, responsibilities and effects on people.

Model evaluation

Does the system respond correctly?

  • Accuracy on representative cases
  • Frequent errors and serious errors
  • Consistency from one response to the next
  • Cost, latency and resource requirements
  • Resistance to misleading inputs
Deployment evaluation

Is the work done better without creating new risks?

  • Time actually saved, including review
  • Final quality of the work and handling of exceptions
  • Responsibilities and ways to challenge a result
  • Workload, autonomy, skills and cooperation
  • Effects observed before and after implementation
The guiding principle

You are not just evaluating a model. You are evaluating the work system formed by the tool, the task, the data, the people who use it and the organisation’s rules.

11 / Before workplace use

Five questions to move from a demonstration to an informed decision.

These questions avoid starting with the product or its reputation. They require a description of the concrete use and the resources needed to keep it under control.

01Which specific task do we want to transform?

Describe the current situation, the expected result, the people involved and the points where an error would have serious consequences. “Saving time” is not a task.

02Which data will be sent, stored or reused?

Identify personal, confidential or protected data; check where it is processed, how long it is retained and the uses planned by the provider.

03Who will check the result, with how much time and what information?

Name the competent person, include review time in their workload and specify the sources against which they can compare the response.

04Who decides and who is accountable for the consequences?

Distinguish the tool’s suggestion, the human decision and the organisation’s responsibility. Provide a way to challenge, correct or suspend use.

05Which effects will we measure after launch?

Before testing, choose a few useful indicators: quality, errors, checking time, workload, autonomy, skills, mutual support and reported difficulties.

Further reading

Choose the resource that matches your question.

A short selection from public bodies or research projects that explain their methods. Level and language are indicated.

Independent newsletter

Follow AI uses, evaluations and their effects on work.