AGI has no universal definition.
Different frameworks propose different thresholds; there is no single recognised test that can declare a system to be AGI.
Technological frontier · Critical perspective
“Artificial general intelligence”, benchmark performance, autonomy and reliability are not interchangeable. This page separates the concepts and explains why safety measures must grow with a system’s capabilities and access.
In the fourth painting of The Course of Empire, a city that appeared all-powerful collapses into chaos. The work is not used here as a catastrophic prediction about AI, but as a reminder: when technical power advances faster than our ability to control it, failures can be amplified rather than contained. AI safety is about identifying those weaknesses, limiting their consequences and preserving the ability to intervene before a crisis. View the work ↗
Different frameworks propose different thresholds; there is no single recognised test that can declare a system to be AGI.
They describe measured trends as data, model size and compute increase. They do not guarantee general intelligence.
Each benchmark measures particular tasks under particular conditions. It does not summarise reliability, usefulness or safety.
The more a system can act, the more explicit its access limits, approvals, monitoring, shutdown and accountability must be.
01 / Define without oversimplifying
AGI is a research and forecasting concept, not a settled technical category. Definitions differ on the human reference level, the range of tasks and the degree of autonomy required.
OpenAI’s Charter defines AGI as highly autonomous systems that outperform humans at most economically valuable work. This framing emphasises broad usefulness and autonomy.
Read the OpenAI Charter ↗Google DeepMind researchers propose levels based on the range of tasks a system can perform and its level of performance. Autonomy is considered separately because it changes the risk of deployment.
Read “Levels of AGI” ↗A fluent conversation, an excellent mathematics score or success on one autonomous task does not establish general intelligence, consciousness or human-like understanding of the world.
02 / Scaling laws
Scaling laws describe empirical relationships: on some training metrics, error decreases fairly regularly when model size, data and compute increase together.

03 / Capabilities
There is no universally “best model”. Suitability depends on the task, allowed tools, cost, latency, error tolerance and the degree of human control required.
Decompose questions, compare hypotheses, calculate and explain a process across many domains.
Draft, modify and test code; summarise long records; prepare analyses and professional documents.
Depending on the model, analyse text, images, documents, audio or video within one task.
Search, run code, call software or operate an interface when an application grants access.
Plan several steps, maintain context and correct some mistakes without becoming consistently reliable.
Produce excellent work and then fail on a simple step, invent a source or pursue the wrong interpretation of a goal.
04 / Current benchmark families
The links below point to maintained project pages and dashboards, so that results can be read with their current protocols rather than frozen into a ranking that will quickly become obsolete.
05 / Why AI safety matters
AI safety covers work to prevent, detect and reduce harm from AI systems, from routine reliability failures to severe risks enabled by greater capability and autonomy.
Code, scientific assistance and automation can support legitimate work or make attacks easier. Access and high-risk use therefore need specific controls.
A system may misunderstand an instruction, fabricate information or use a tool incorrectly. Automation can propagate a small error before it is noticed.
When an agent plans and acts for longer, it becomes harder to anticipate each action and detect when the objective has been misread.
Personalised content at scale can amplify fraud, manipulation and disinformation, while polished language may invite excessive trust.
Provider dependence, version changes, outages and gradual deskilling can create systemic organisational risk.
A technically safe model can still intensify work, reduce autonomy, shift responsibility or impose continuous monitoring. These effects belong in AI safety.
Safety is not about predicting the date of AGI. It is about checking, before each increase in capability or autonomy, that dangerous behaviour can be detected, consequences limited and decisions challenged.
06 / From power to control
Model evaluations are necessary but insufficient. They must be combined with technical, organisational and human measures suited to the real context of use.
Capabilities, severe errors, misuse, robustness, cybersecurity and effects on work.
Bound data, tools, budget, duration, number of actions and operational scope.
An identified person approves sensitive, irreversible or consequential decisions.
Logs, alerts, worker feedback, version tracking and an accessible reporting channel.
A shutdown process, fallback route, incident analysis and reassessment after major changes.
A voluntary framework for governing, mapping, measuring and managing AI risks throughout the system lifecycle.
Open the framework ↗Use the short questionnaire to structure a discussion about psychosocial risks, opacity, oversight burden and skill retention before a pilot.
Open the short questionnaire →