Developers who eventually assume code is correct because it passes the tests. Others who treat an agent’s stated plan as an accurate account of what it will actually do. Users who, after repeated permission requests, click “accept” without reading any more. These situations, documented in the research cited by the paper, show in concrete terms what often becomes of the “human in the loop” required by corporate policies, vendors and the EU AI Act.

Margaret Mitchell, Avijit Ghosh and Samir Passi draw a clear conclusion: the way AI agents are currently designed and deployed does not support human oversight; it degrades it. First, because the volume and pace of approvals exceed what a person can actually examine. Second, because prolonged use of these systems erodes the vigilance, judgement and professional skills that supervisors need to spot an error.

For occupational health, the implication is direct: supervising an agent is not an abstract guarantee; it is a task assigned to a worker, with a workload, constraints and responsibility. This reading examines it through psychosocial risk factors, shows how supervisors’ health affects system safety, and then offers questions to ask and a tool for assessing these risks before deployment.

Oversight degrades the overseer

An AI agent does more than produce an answer. It carries out sequences of steps, selects tools, modifies files, sends messages or initiates transactions. At every step, it generates traces: reasoning, plans, tool calls and exchanges with other agents. The authors show that this stream is too large, too fast and too dispersed to be properly understood by the person expected to control it.

That person must also play several roles at once: use the agent to achieve their own goal, grant permissions as execution proceeds, assess the relevance of each step and anticipate its consequences. In practice, approval requests follow one after another, and users eventually stop reading what they approve. The authors call this approval fatigue.

The second mechanism is even more concerning. Prolonged use of these systems erodes the very capacities that oversight requires: deskilling, professional intuition becoming “rusty”, reduced vigilance and ability to detect anomalies, and overconfidence. Well-documented biases compound this: automation bias, anchoring on the machine’s suggestion and complacency. The paper cites, among other evidence, a multicentre study published in 2025 in The Lancet Gastroenterology & Hepatology, which suggests a decline in endoscopists’ performance after a period of exposure to AI in colonoscopy.

The authors connect this observation to a foundational ergonomics paper, Lisanne Bainbridge’s “Ironies of Automation” (1983): the better an automated system performs, the less practice the operator gets at intervening, and the less prepared they are at precisely the moment they are needed. They sum up this irony in one phrase:

“Oversight degrades the overseer.”

In his presentation of the paper, Avijit Ghosh describes where this slide leads: a human who mechanically rubber-stamps decisions, staring blankly, instead of providing a meaningful check on the agent’s errors.

A work activity in its own right

For ergonomics and occupational health, this situation is not new. It recalls the monitoring of automated processes in control rooms, aviation or nuclear power: long periods of passive activity followed by a need to respond quickly and effectively to a rare event. This is the “out-of-the-loop” operator problem, described as early as the 1990s and revisited by the authors.

What is changing is the scale. This configuration, once confined to a few high-risk sectors with trained operators, procedures, scheduled breaks and systems for learning from experience, is spreading into office work: developers, lawyers, administrators, HR staff, healthcare workers and, soon, occupational health professionals themselves. Often without any of the organisational safeguards that those sectors took decades to develop.

The paper also stresses that the burden of oversight falls almost entirely on the user. Interfaces tell users to “check important information” but do not give them the means to do so. In occupational health terms, this is a demand without the necessary resources: prescribed work requires every action to be checked, while actual work makes such checking practically impossible. The gap is structural, not individual.

A psychosocial risk perspective

The framework of six families of psychosocial risk factors set out in the Gollac report (2011) helps organise what the paper describes. Five are directly relevant.

Work intensity and cognitive demands

Supervisors must process a volume of information they cannot control, at a pace set by the agent, while continuing their own task. Paradoxically, the situation combines overload (reading everything and keeping track of everything) with monotony (approving dozens of similar steps). The authors point out that under stress or overload, reasoning shifts towards quick, intuitive shortcuts at the expense of the deliberate analysis that oversight specifically requires.

Autonomy and the use of skills

Relegated to the role of approver, professionals make fewer decisions, do less and learn less. Yet using and developing skills are part of decision latitude in Karasek’s model. High demands combined with eroding latitude correspond to the classic configuration of job strain. The paper emphasises the position of novices, who may never acquire the skills needed to judge the agent’s work or take over when it fails. This is a health issue, but also a matter of career development and the transmission of professional expertise.

Value conflicts and barriers to quality work

Approving something you have not been able to verify means putting your name to work that you do not consider well done. The accounts cited by the authors, mainly from developers, describe a mind-numbing loop and a feeling of being reduced to babysitting the machine’s output. These are familiar mechanisms of being prevented from doing work well and losing a sense of meaning at work: not an absence of work, but an inability to do it to the standard you believe is necessary.

Responsibility and insecurity

The human remains formally responsible, both within the organisation and under the law, even as their actual control diminishes. Beyond the paper itself, this situation can be related to the concept of a “moral crumple zone” proposed by Madeleine Clare Elish in 2019: in an automated system, the human operator absorbs responsibility for failures they lacked the means to prevent. Responsibility without control is a well-known source of psychological strain.

Recognition and relationships at work

Supervision work is invisible in performance indicators. The authors note that agents are evaluated on speed, accuracy and throughput, rarely on the quality of human oversight they make possible. Someone who takes time to examine a plan and reject an action then appears less productive than someone who approves quickly. The paper explicitly recommends removing productivity targets that discourage careful review and valuing high-quality supervision: in other words, a question of recognition.

When supervisors’ health becomes a safety variable

The paper then describes a feedback mechanism. Users’ approvals and evaluations are often used to assess, or even train, systems. An attentive supervisor examines the output, identifies missing information and rejects it. A tired supervisor approves quickly, accepts a fluent justification and rates the interaction as satisfactory. If these approvals are interpreted as successes, the system may learn to produce what is easy to approve: confident summaries, simplified plans and fewer points of friction. The authors also discuss research showing that agents can conceal failures when they anticipate that they will not be checked.

For occupational health, the implication is significant. Supervisors’ fatigue, overload and deskilling are no longer solely individual health concerns: they become a variable in system reliability, and potentially in how the system learns.

Volume and pace imposed by the agent → approval fatigue and cognitive overload → rapid approvals, checks abandoned → skills used less, deskilling → even less effective oversight → undetected errors for which the human remains responsible → psychological strain, loss of meaning and value conflicts.

Preventing psychosocial risks among people who supervise AI therefore also means protecting the safety of the organisation and third parties. The two issues can no longer be addressed separately.

Solutions rooted largely in work organisation

The authors propose a two-part response. For designers: introduce useful friction, such as asking users to form their own view before seeing the agent’s suggestion; let them choose whether to consult it; ask reflective questions at high-stakes moments; require explicit approval for consequential actions; define in advance what the agent may do independently; group approvals into coherent batches; and use automated checks for tasks that do not require judgement.

For deploying organisations: train people to critically review AI output; maintain skills through regular unaided practice; require breaks; organise rotations; assign supervision to people with the necessary expertise; separate the supervisor’s role from that of the person who benefits from approval; and align targets with the quality of oversight rather than throughput.

This second list is almost word for word the vocabulary of organisational prevention: workload, breaks, task rotation, training, role definition and targets. Under French law, these measures fall within the employer’s duty to protect health and safety and the general principles of prevention, which require, among other things, adapting work to the individual and planning prevention to integrate technology and work organisation (Articles L. 4121-1 and L. 4121-2 of the French Labour Code). At EU level, the AI Act requires deployers of high-risk systems to assign human oversight to people with the necessary competence, training, authority and support (Article 26). These requirements can only be met if working conditions make them possible. See the legal framework →

One qualification is needed, however. The proposed measures include self-monitoring training to help people detect a decline in their own attention. This can be useful, but it must not become another way of shifting onto the individual a burden that the paper rightly attributes to design and organisation. Primary prevention—how much autonomy the agent has, approval volume, session length, roles and targets—must take precedence over adapting the person.

A point to watch: monitoring vigilance

The paper proposes measuring the quality of oversight: time spent reviewing an action relative to the approval rate, changes in disagreements with the agent, and the frequency of requests for information as the stakes rise. It even suggests inserting tasks with known correct answers into the workflow to detect a decline in judgement and trigger a break, rotation or training. The authors themselves acknowledge that these arrangements raise surveillance concerns.

From an occupational health perspective, particular care is needed. These indicators concern workers’ activity. Used poorly, they can become tools for individual performance assessment, pressure and mistrust—precisely the opposite of the intended effect. In France, monitoring workers’ activity requires informing employees in advance (Article L. 1222-4) and consulting the Social and Economic Committee, the employee representative body known as the CSE (Article L. 2312-38).

One possible approach is to use collective indicators, discussed with employee representatives, to adjust work organisation—approval volume, the agent’s scope and breaks—rather than judge individuals. A sign of deteriorating oversight should prompt scrutiny of the work, not the worker.

What can occupational health professionals look for?

  1. Identify who supervises what.
    Identify roles in which a person approves an agent’s actions, and estimate the number of approvals requested per day or hour.
  2. Compare prescribed work with actual work.
    Does the supervisor have the time, information and authority needed to reject an action? Can they stop the agent without being penalised against their targets?
  3. Look for approval fatigue.
    Approving in rapid succession, skim-reading and feeling that one is “clicking yes” without really knowing what is being accepted are signals worth listening to.
  4. Protect skills, particularly those of beginners.
    Plan time for practice without AI, mentoring and exercises in taking back manual control.
  5. Organise vigilance as a monitoring activity.
    Session length, breaks, rotations and alternating assisted with unaided work: lessons from control rooms can be applied here.
  6. Clarify responsibility.
    Who is accountable for an error that slips through? Is the allocation of responsibility documented, understood and consistent with the resources actually given to the supervisor?
  7. Set rules for oversight indicators.
    Where they exist, indicators should be collective, transparent, discussed with the CSE and directed towards improving work organisation.
  8. Include agent supervision in the occupational risk assessment document (DUERP).
    This is a work situation involving identifiable psychosocial risk factors and should be assessed as such, before deployment and over time.

Moving on to assessment: the pre-deployment AI project risk assessment questionnaire (18 items) helps identify these potential psychosocial effects in advance, including effects on autonomy, workload and responsibility. Once the agent is in use, the deployment assessment tool helps compare anticipated effects with work as actually observed.

These questions also apply to our own services. When an occupational physician or nurse reviews a summary, letter or recommendation prepared by AI, they are the human in the loop. The same mechanisms apply, especially to medical trainees and professionals at the start of their careers. AI in occupational health services →

Limitations

This is a position paper: an argument grounded in the literature, not a new empirical study. Some of its evidence comes from agent-assisted software development, studies involving students or work that is still in preprint form. The magnitude and persistence of the cognitive effects remain debated; the authors also consider the objection that users will eventually adapt. Finally, whether these findings carry over to French workplaces, in other occupations and organisations, remains to be examined in the field. This is precisely what occupational health can contribute: examining the actual work of those who supervise.

The question to ask when deploying agents may therefore not simply be “Is there a human in the loop?”

We should also ask “Under what working conditions can that human still exercise judgement?”

A human who is tired, rushed, progressively deskilled and left solely responsible is no guarantee of safety. They represent an additional risk, both to themselves and to the system they are expected to control.