“AI” does not refer to a single tool.
Generating text, assigning a score and allocating tasks are three different functions, with different risks.
Understand · Essential concepts
Distinguish between systems, understand how a language model is built and produces a response, interpret its performance without reading too much into it, then explore its risks and the goal of AGI.
Generating text, assigning a score and allocating tasks are three different functions, with different risks.
It calculates probabilities from the context; it does not automatically check whether its response is true or complies with the rules.
The model matters, but so do the task, the data, the people, the decision rules and the organisation.
01 / Identify the system
The term “AI” covers systems that do not produce the same kind of result. To understand a workplace use, start by identifying what the tool actually does.
It creates or transforms text, images, sound or code from an instruction. A conversational assistant based on an LLM belongs to this family.
It estimates a probability, assigns a score or classifies a case based on past data: default risk, case priority or diagnostic support.
It allocates tasks, sets processing priorities, measures activity or influences a management decision. It may or may not incorporate generative AI.
02 / Inside an LLM
A large language model — or LLM — is a neural network trained to predict the continuation of a text. It does not look up a ready-made answer in a database: it composes one, unit by unit.
The instruction and supplied documents are split into small units: words, parts of words or symbols.
Each token is converted into a numerical representation that encodes patterns learned during training.
The Transformer examines relationships between tokens to determine which parts of the context matter most at that moment.
The model calculates a probability distribution. Generation settings influence the choice and its variability.
The selected token joins the context, then the calculation starts again until the response is complete.
Text already in the context“The worker describes persistent pain in the ▯ ”
“back” is selected almost every time: consistent, sometimes repetitive responses.
“lower” or “wrist” appear more often: more variety, more variation.
“…in the back ” → “,” → “aggravated” → “by”… each added token starts the calculation again.
03 / Training
An LLM’s capabilities are not programmed rule by rule. They develop through several stages of training, each of which leaves its mark on the final behaviour, including its shortcomings.
01Pre-train
02Fine-tune
03Reinforce
04Reason
05Evaluate, deploy
↺ Shortcomings identified in testing or after deployment inform the training of the next version.
Performance improves fairly predictably with the amount of computation, data and parameters: these are the “scaling laws”. According to Epoch AI, the computation used to train notable AI models has grown by about 4.5 times each year since 2010, and by about 5 times for frontier large language models since 2020.
Letting a model work longer before answering improves results in mathematics, programming and science. The International AI Safety Report 2026 identifies this as one of the main sources of recent progress.
| Model | Date | Computation (FLOP) | Status |
|---|---|---|---|
| AlexNet | Sept. 2012 | 4.7 × 10¹⁷ | Published |
| Transformer | June 2017 | 7.4 × 10¹⁸ | Published |
| GPT-2 | Feb. 2019 | 1.9 × 10²¹ | Estimate |
| GPT-3 | May 2020 | 3.1 × 10²³ | Published |
| PaLM | Apr. 2022 | 2.5 × 10²⁴ | Published |
| GPT-4 | Mar. 2023 | 2.1 × 10²⁵ | Estimate |
| Llama 3.1 405B | July 2024 | 3.8 × 10²⁵ | Published |
| GPT-4.5 | Feb. 2025 | 3.8 × 10²⁶ | Estimate |
| Grok 4 | July 2025 | 5.0 × 10²⁶ | Estimate |
| GPT-6 Astra | Sept. 2026 | 1.0 × 10²⁷ | Estimate |
04 / From model to system
A workplace assistant combines a model, instructions, documents, tools and sometimes memory. Each layer changes the result, and therefore what needs checking before you entrust it with a task.
Before your first question, the provider or organisation has set a role, a tone, restrictions and formatting rules. Two tools based on the same model can therefore respond very differently.
The model only “sees” its context window: the conversation, supplied documents and passages retrieved from a document store (RAG). Anything outside it does not exist for the model, and very long documents may be used unevenly.
Web search, code execution, email, calendars or case records: the model can call tools through connectors, such as those using the MCP protocol. It no longer just makes suggestions; it acts.
An agent plans, acts, observes the result and starts again, sometimes for hours, with intermittent human supervision. An early mistake can propagate through the entire chain of actions.
Closed models, such as GPT, Claude or Gemini, can only be used remotely on the provider’s infrastructure. Open-weight models, such as Llama, Qwen, DeepSeek or some Mistral models, can be installed on servers under your control, generally with a slight lag behind the best closed models.
Recent models also process images, scanned documents, audio and video. Automatic transcription and summarisation of consultations, sometimes called medical “scribes”, are their most visible use in healthcare.
05 / Capabilities and limits
Capabilities depend on the model, its version, the language, the supplied documents, the tools it can access and how the task is phrased.
| Task type | What the model can contribute | What cannot be assumed |
|---|---|---|
| 01Reading and writing | Summarise, rephrase, translate, extract information, compare documents or draft an initial text. | Completeness, faithfulness to the source, up-to-date information and suitability for the professional context. |
| 02Reasoning through a problem | Break down a question, suggest hypotheses, apply a described procedure and explain the steps in a calculation. | Consistent accuracy. A convincing explanation can accompany a wrong answer. |
| 03Writing code and processing data | Generate, explain or correct code; prepare a query or an exploratory analysis. | Security, robustness and correct operation on untested cases. Running the code and testing it remain essential. |
| 04Using tools | With a suitable application, search the web, consult a document store, call software or carry out several actions in sequence. | Source quality, permission to act and control over consequences. Access rights must be limited. |
| 05Processing multiple media | Depending on the model: describe an image, transcribe audio, analyse a document or answer based on a video. | Complete perception of details, understanding of the situation and compliance with profession-specific rules. |
Frequent success on a category of tasks does not make every response predictable. You need to test the errors that matter for your use.
The provider may change the model, its settings, filters or tools. A useful evaluation records the version and the date.
These limits are not simply a matter of settings. They arise from how models are trained, evaluated and integrated into tools. Understanding them helps place checks where they are needed.
Pre-training does not teach the distinction between a rare fact and a plausible invention. Evaluations that score only accuracy then reward guessing rather than abstaining.
At workInvented references, legal provisions or figures, presented with confidence.
Kalai et al., OpenAI, 2025 ↗Fine-tuning on human preferences can favour agreement over accuracy. In April 2025, OpenAI withdrew a GPT-4o update considered excessively flattering.
At workThe tool tends to confirm the hypothesis embedded in the question.
OpenAI, 2025 ↗Capabilities do not follow human perceptions of difficulty. Among 758 consultants, AI improved speed and quality on tasks within its reach, but reduced answer accuracy on a task just beyond it.
At workTrust earned on one task does not carry over to the next.
Dell’Acqua et al., Harvard Business School, 2023 ↗The reasoning a model displays does not faithfully reflect what determined its answer. In an Anthropic study, the models tested disclosed a hint they had used in only 25 to 39% of cases.
At workA convincing justification does not prove the answer is well founded.
Chen et al., Anthropic, 2025 ↗The model does not reliably separate the user’s instructions from the text it reads: a web page, email or PDF may contain hidden instructions. This is the top risk in OWASP’s ranking for LLM applications.
At workThe risk increases as soon as the tool can send, edit or delete.
OWASP, LLM01:2025 ↗Generation is probabilistic and sensitive to wording, document order and model version. Its knowledge stops at a cut-off date, unless it can consult live sources.
At workTest repeatedly, date the trials and record the version used.
06 / Reading benchmarks
A benchmark measures performance on a set of tasks, with a specific instruction, metric and conditions. Change one of these elements and the ranking may change.
The pages below are maintained by their authors. They provide the latest available results without fixing a ranking here that would quickly become outdated. The snapshot below simply shows that the leader changes depending on the test.
Epoch AI · combines several dozen tests
Claude Opus 5.5 not yet ranked
Artificial Analysis · Elo score, pairwise comparisons
ARC Prize · verified scores
07 / Tracking progress
To anticipate effects on work, the useful question is not “which model is best?” but “what is becoming feasible, how quickly and under what conditions?”. Four dated reference points, with their limits.
| Model | Release | 50% horizon | 95% CI |
|---|---|---|---|
| GPT-2 | Feb. 2019 | 3 s | 1 s – 8 s |
| GPT-3 | May 2020 | 8 s | 5 s – 13 s |
| GPT-3.5 | Mar. 2022 | 36 s | 16 s – 1 min |
| GPT-4 | Mar. 2023 | 4 min | 2 min – 8 min |
| GPT-4 (Nov.) | Nov. 2023 | 4 min | 2 min – 8 min |
| GPT-4o | May 2024 | 7 min | 4 min – 13 min |
| Claude 3.5 Sonnet (June) | June 2024 | 11 min | 5 min – 22 min |
| o1-preview | Sept. 2024 | 20 min | 12 min – 33 min |
| Claude 3.5 Sonnet (Oct.) | Oct. 2024 | 21 min | 10 min – 41 min |
| o1 | Dec. 2024 | 39 min | 21 min – 1 h 05 |
| Claude 3.7 Sonnet | Feb. 2025 | 1 h 00 | 33 min – 1 h 44 |
| o3 | Apr. 2025 | 2 h 00 | 1 h 15 – 3 h 11 |
| GPT-5 | Aug. 2025 | 3 h 23 | 1 h 53 – 6 h 46 |
| Gemini 3 Pro | Nov. 2025 | 3 h 44 | 2 h 20 – 6 h 19 |
| Claude Opus 4.5 | Nov. 2025 | 4 h 53 | 2 h 42 – 10 h 24 |
| GPT-5.2 | Dec. 2025 | 5 h 52 | 3 h 18 – 13 h 35 |
| Claude Opus 4.6 | Feb. 2026 | 11 h 59 | 5 h 17 – 60 h 34 |
| Claude Mythos Preview | Apr. 2026 | 17 h 25 | 8 h 29 – 55 h 04 |
| GPT-5.6 Sol (not in frontier-model series) | June 2026 | ≈ 11 h 18 | 5 h 00 – 40 h 00 |
Systems capable of solving expert-level problems fail at simple perception or adaptation tasks. The International AI Safety Report (February 2026) considers both a plateau and a sharp acceleration plausible by 2030.
A profession combines tasks, relationships, responsibilities and unexpected events. These reference points indicate which parts of work are becoming automatable, not what work will become: that also depends on organisational choices.
08 / System risks
The International AI Safety Report, written by more than 100 experts under the chairmanship of Yoshua Bengio, distinguishes malicious uses, malfunctions and systemic risks. Effects on work and health are detailed in the Risks section.
Fraud and scams, deepfakes, manipulation, assistance with cyberattacks. Labs also evaluate the help their models could provide in designing biological or chemical weapons.
At workTargeted phishing, impersonation of an executive, reputational harm.
Fabricated content, faulty code, agents that fail or subvert their assigned objective. Under test conditions, some models have exploited weaknesses in their scoring system or adopted deceptive strategies to achieve a goal.
At workUnwanted agent actions, human oversight reduced to rubber-stamping.
Automation of cognitive tasks, weakened critical thinking through excessive trust in machines, emotional dependence on conversational companions used by tens of millions of people.
At workDeskilling, work intensification, loss of autonomy and meaning in work.
Tests conducted before release are poor predictors of usefulness and risks in real conditions, and some models now detect when they are being evaluated. Post-deployment monitoring, including in workplaces, is becoming essential.
Major labs publish safety frameworks, such as Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework, linking capability thresholds to stronger measures. Interpretability research seeks to understand models’ internal mechanisms rather than just their responses. These commitments remain largely voluntary and assessed by the companies themselves.
Since 2 August 2025, the AI Act has required providers of general-purpose AI models to maintain technical documentation and a copyright compliance policy. Models posing systemic risk must also be evaluated, undergo risk management and have serious incidents reported. Explore Governance →
09 / The goal of AGI
Several companies say they aim to build artificial general intelligence, or AGI. There is neither a shared definition nor a test that could establish its arrival. Three ways of defining it shed light on what is at stake for work.
OpenAI’s charter (2018) aims for “highly autonomous systems” that outperform humans at most economically valuable work. The definition is explicitly economic: it is measured in terms of replacing human labour.
OpenAI Charter, 2018 ↗Google DeepMind combines the level of performance, from “emerging” to “superhuman”, with the breadth of tasks covered. The degree of autonomy granted to the system is treated separately: it is a deployment choice, not just a capability.
Morris et al., Levels of AGI, 2023 ↗A group of researchers measures models across ten broad human cognitive abilities (the CHC model). GPT-4 scored 27% and GPT-5 57%, with a highly uneven profile: extensive knowledge, almost no long-term memory.
Hendrycks, Bengio et al., 2025 ↗In the contract between Microsoft and OpenAI, declaring AGI had financial consequences; verification was entrusted to a panel of independent experts in October 2025, before an April 2026 amendment largely stripped the clause of its effect by separating payments from technological progress.
Several lab leaders publicly suggest timelines of a few years. These statements come from participants in a race for funding; measured, dated indicators are more informative than announced dates.
The issue is not to bet on a date. It is to track, task by task, what is becoming automatable, how quickly and under what oversight, and to anticipate effects on workload, autonomy, skills and employment before they take hold. Explore the economic and social issues →
10 / Back to real work
A benchmark describes a model’s behaviour in a test. Deployment also needs to be evaluated in the actual task, with its constraints, responsibilities and effects on people.
You are not just evaluating a model. You are evaluating the work system formed by the tool, the task, the data, the people who use it and the organisation’s rules.
11 / Before workplace use
These questions avoid starting with the product or its reputation. They require a description of the concrete use and the resources needed to keep it under control.
Describe the current situation, the expected result, the people involved and the points where an error would have serious consequences. “Saving time” is not a task.
Identify personal, confidential or protected data; check where it is processed, how long it is retained and the uses planned by the provider.
Name the competent person, include review time in their workload and specify the sources against which they can compare the response.
Distinguish the tool’s suggestion, the human decision and the organisation’s responsibility. Provide a way to challenge, correct or suspend use.
Before testing, choose a few useful indicators: quality, errors, checking time, workload, autonomy, skills, mutual support and reported difficulties.
Further reading
A short selection from public bodies or research projects that explain their methods. Level and language are indicated.
A short definition to establish the vocabulary without going into technical detail.
CNIL ↗Seven short videos on how it works, strengths, limits, instructions and tool use.
Inria ↗Tokens, context, parameters, attention, limitations and model adaptation.
Google for Developers ↗Benefits, inaccurate results, personal data and initial deployment precautions.
CNIL ↗Framework, resources and a profile addressing risks specific to generative AI.
NIST ↗A synthesis by more than 100 experts on the capabilities, risks and safeguards of general-purpose AI systems.
International report ↗Connect an evaluation protocol with professional knowledge and the errors that really matter.
AI & Occupational Health ↗Connect an understanding of models with risks, uses in occupational health services, evaluation, prevention and law.
AI & Occupational Health →