What can we do now without pretending to know the future?
The best safety strategy does not depend on winning the argument about a precise extinction probability.
Many safeguards are valuable across a wide range of beliefs.
Evaluate dangerous capability before deployment
Labs can test models for high-consequence capabilities such as sophisticated cyber operations, biological knowledge uplift, autonomous planning, deception, and the ability to circumvent oversight. OpenAI’s Preparedness Framework, Anthropic’s Responsible Scaling Policy, and Google DeepMind’s Frontier Safety Framework are different attempts to connect capability thresholds to stronger safeguards.
These are voluntary company frameworks and should be evaluated critically, but their existence is itself evidence that leading developers treat some capability thresholds as safety-relevant.
Restrict access to high-consequence tools
An AI assistant that can draft text is different from an agent holding cloud administrator credentials, financial authority, laboratory control, or military permissions. Least-privilege design—giving a system only the permissions required for the task—can sharply reduce the damage a failure can cause.
Secure the infrastructure and model weights
A well-aligned API is not sufficient if an attacker can steal powerful model weights, bypass provider controls, or compromise the servers on which the system runs. Cybersecurity is part of AI safety.
Keep humans in genuinely meaningful control
“Human in the loop” is useless if the human is overwhelmed, has two seconds to approve a machine recommendation, or is culturally trained to click yes. Meaningful control requires time, information, authority, and the practical ability to disagree.
Stage deployment
Releasing to small, monitored environments first creates opportunities to discover failures before scaling. High-risk capabilities should face stricter deployment gates than ordinary convenience features.
Prepare for incidents
Organizations should know who can revoke credentials, isolate a model, shut down integrations, preserve logs, contact infrastructure providers, and communicate with authorities. Safety improves when emergency powers are defined before the emergency.
Read the full guide: How Do We Prevent Catastrophic AI Risk?.
This is one section of a comprehensive guide.
Read the Full Article →