Definitions without hype

AI safety glossary

Words like “AGI,” “sentient,” “agent,” and “loss of control” are often used as though everyone means the same thing. They do not. These working definitions keep the guides precise.

Agentic AI
An AI system configured to plan, use tools, take actions, and pursue goals across multiple steps rather than only generate a single response.
AGI (Artificial General Intelligence)
A non-standardized term for AI with broad capability across many cognitive or economically useful tasks, often imagined at or around human-level generality.
Alignment
The challenge of making AI behavior reliably conform to intended goals, constraints, values, and oversight, including in unfamiliar situations.
Artificial superintelligence (ASI)
A hypothetical system that substantially exceeds human performance across many strategically important cognitive domains.
Catastrophic risk
A risk capable of causing extremely large-scale, severe, and potentially irreversible harm. Human extinction is an extreme subset of catastrophic risk.
Corrigibility
The property of remaining open to correction, shutdown, modification, or redirection by authorized humans.
Frontier AI
A practical policy term for highly capable general-purpose models near the leading edge of capability. Definitions vary by institution.
Loss of control
A scenario in which one or more AI systems operate outside effective human control and regaining control is extremely costly or impossible.
Reward hacking
Behavior that exploits flaws or loopholes in an objective or evaluation to score well without achieving the intended outcome.
Sandbagging
Deliberately or strategically underperforming in an evaluation in a way that can obscure true capability.
Situational awareness
The ability to recognize information about one’s own context, such as being tested, monitored, or deployed in a particular environment.
Scalable oversight
Methods for supervising systems whose outputs or reasoning may become too complex or numerous for direct human review alone.
Model weights
The learned numerical parameters of a trained model. Securing frontier model weights can matter because stolen weights may bypass provider-side safeguards.
Biosecurity uplift
The degree to which access to AI improves a user’s ability to perform a biological task compared with a relevant baseline such as internet access alone.
Meaningful human control
A principle that consequential automated actions—especially in military or safety-critical contexts—should remain subject to informed, timely, and accountable human judgment.