Publications · Detailed summary

AI and medical recommendations

An experimental study of AI-assisted review of medical recommendations, published by INRS in June 2026. This page presents the methods, results and limitations.

INRS · Experimental study · June 2026

Artificial intelligence in occupational health: assessing its potential to improve the precision and clarity of medical recommendations.

Why study the review of medical recommendations?

An imprecise recommendation may be misunderstood or difficult to implement. Earlier work by Dr N’Guessan collected 4,217 recommendations from twelve intercompany occupational health services in the Hauts-de-France region. Of the 3,776 assessed through multidisciplinary consensus, 2,947 (78%) had at least one defect according to the criteria used. This finding concerns that corpus; it is not a national estimate.

The study asks whether the o1 reasoning model can help identify these wording defects by comparing its assessments with this human reference.

Methods

The model’s instructions were refined over five iterations using approximately 100 recommendations, then retested on 50 new examples. The final prompt was applied to a new random sample of 385 recommendations from the database. The information submitted did not identify workers, companies, employers, physicians or occupational health services.

Each recommendation was assessed against five criteria:

  1. Imprecision or difficulty understanding and implementing the recommendation.
  2. Uncertainty about whether the recommendation is binding.
  3. Information that does not belong in the recommendation document.
  4. A disguised change of job or finding of unfitness for work.
  5. Breach of medical confidentiality or privacy.

Disagreements were reassessed by the author with the assistance of an occupational health professor, distinguishing relevant alerts, excessive flagging, missed defects and hallucinations.

Results

Table II reports the following distribution across the five criteria. These are assessments by criterion, not the percentage of recommendations that were entirely correct.

Distribution of assessments — 385 recommendations, five criteria
OutcomeProportionInterpretation
Agreement74.60%Agreement with the multidisciplinary consensus.
Justified disagreement10.54%Disagreement judged relevant after reassessment.
Excessive flagging13.20%A finding judged too severe.
Missed defect1.45%A defect identified by the reference but missed by the model.
Observed hallucination0%None detected within this experiment.

Performance varied by criterion. Agreement reached 96.1% for medical confidentiality, with a Cohen’s kappa of 0.82, compared with 45.5% for uncertainty about the binding nature of recommendations and 61.3% for imprecision. The model tended to flag more defects than human assessors, particularly where interpretation was involved.

What the study supports

The findings support the value of a second reader that draws attention to wording requiring further review. Some disagreements revealed defects initially missed by the consensus. However, a high agreement rate does not make the model an autonomous validator: its alerts must be considered in context and assessed by the professional.

Limitations and next steps

Neither the model nor the assessors had access to job details or all relevant constraints. The ratings involve interpretation. The findings concern a specific model, prompt and corpus, and cannot automatically be generalised to other systems. The absence of observed hallucinations in this study does not guarantee their absence in other uses.

The study assesses the detection of wording defects, not a demonstrated improvement in work retention or health. It opens the way to occupational health evaluations that should be repeated as models and their conditions of use evolve.

Key takeaway

A review aid that can strengthen attention to medical recommendations, with the final decision remaining the occupational physician’s responsibility.

Source: full article TF 335, particularly methods, Tables I–II and limitations, pp. 55–58 (French).