Definitions without hype
AI safety glossary
Words like “AGI,” “sentient,” “agent,” and “loss of control” are often used as though everyone means the same thing. They do not. These working definitions keep the guides precise.
- Agentic AI
- An AI system configured to plan, use tools, take actions, and pursue goals across multiple steps rather than only generate a single response.
- AGI (Artificial General Intelligence)
- A non-standardized term for AI with broad capability across many cognitive or economically useful tasks, often imagined at or around human-level generality.
- Alignment
- The challenge of making AI behavior reliably conform to intended goals, constraints, values, and oversight, including in unfamiliar situations.
- Artificial superintelligence (ASI)
- A hypothetical system that substantially exceeds human performance across many strategically important cognitive domains.
- Catastrophic risk
- A risk capable of causing extremely large-scale, severe, and potentially irreversible harm. Human extinction is an extreme subset of catastrophic risk.
- Corrigibility
- The property of remaining open to correction, shutdown, modification, or redirection by authorized humans.
- Frontier AI
- A practical policy term for highly capable general-purpose models near the leading edge of capability. Definitions vary by institution.
- Loss of control
- A scenario in which one or more AI systems operate outside effective human control and regaining control is extremely costly or impossible.
- Reward hacking
- Behavior that exploits flaws or loopholes in an objective or evaluation to score well without achieving the intended outcome.
- Sandbagging
- Deliberately or strategically underperforming in an evaluation in a way that can obscure true capability.
- Situational awareness
- The ability to recognize information about one’s own context, such as being tested, monitored, or deployed in a particular environment.
- Scalable oversight
- Methods for supervising systems whose outputs or reasoning may become too complex or numerous for direct human review alone.
- Model weights
- The learned numerical parameters of a trained model. Securing frontier model weights can matter because stolen weights may bypass provider-side safeguards.
- Biosecurity uplift
- The degree to which access to AI improves a user’s ability to perform a biological task compared with a relevant baseline such as internet access alone.
- Meaningful human control
- A principle that consequential automated actions—especially in military or safety-critical contexts—should remain subject to informed, timely, and accountable human judgment.