Can AI Safety Keep Up With AI?
AI is increasingly helping build, test and monitor the safeguards around AI systems. That can strengthen security — but it also creates a new question: who independently verifies the verifier, and do we still have enough time to prove the safeguards work before the next capability arrives?
AI is taking a larger role in building, testing and monitoring digital systems. That includes the security around AI itself. This is fundamentally positive: better analysis, clearer logs and faster detection can make safeguards more useful. But it opens a harder question: Do we still have enough time to verify safeguards before capabilities change again?
Security really is becoming more sophisticated
It would be wrong to portray the AI industry as indifferent to security. Several companies now publish concrete control approaches. NVIDIA, for example, describes isolated execution, policy controls and monitoring outside the agent. Microsoft describes separate agent identities, least privilege and deterministic human review for high-risk actions.
These are important directions. An agent should not receive broader access than its task requires, and a consequential action should not pass merely because a model judges it reasonable. But these are company descriptions of their systems, not proof that every system is secure.
Trust cannot simply move one layer deeper
We do not fully trust the agent, so we build a policy or control layer around it. Then the next question is: who verifies the policy layer? We ground an agent in evidence and history. But who verifies the evidence chain, selection and assumptions?
Trust cannot simply move one layer deeper and stop there. Every time control moves to a new layer, that layer needs its own independent verification mechanism.
Independence means more than one AI model checking another. The reviewer needs to be able to challenge both the result and its frame: which data are used, which errors may be missed and what actions the system is actually allowed to take. Organizational separation between the builder and the evaluator can matter, especially when a positive review affects whether a product proceeds.
The control methods themselves also need variety. Several verification methods built on the same assumptions or failure modes are not necessarily independent. For critical controls, deterministic or other non-AI checks can be useful where they fit. Human review is not automatically sufficient either: the reviewer needs relevant information, time and real authority to challenge or stop a result. AI-based verification can be part of the answer, but it should not alone define what counts as adequate security.
When AI helps secure AI
AI can help security teams find patterns in logs, review code, prioritize alerts and propose actions. CrowdStrike describes continuous identity and risk-aware authorization for AI agents. OpenAI describes sandboxing, network policies, approvals and agent-native telemetry for Codex.
That can be highly valuable. But AI-assisted security is not the same thing as AI-trusted security. When a model does work and another model participates in review, the core question is not whether AI may be used in the control loop. It is whether there is still a verifiable, independent authority with enough information, time and mandate to say no.
Then comes the time problem
Security is not only a technical question. It is also a time question. A control may be well designed for today’s system and still be insufficiently tested for a system with new tools, wider authority or more integrations. Google DeepMind has presented both a threshold-based safety framework and a pilot for double-blind external evaluation. Anthropic has described external and embedded evaluation as ways to identify blind spots.
Google DeepMind explicitly tries to address the timing problem through safety buffers and recurring evaluations intended to provide warning before critical capability thresholds are reached. The open question is whether those margins remain sufficient as capabilities and use cases change ever more quickly.
Anthropic also notes that there are not yet established standards for what information such evaluators should be given access to or how results should be reported, nor an established system for funding independent evaluations.
This suggests that independent testing is being taken seriously. Yet it is not automatic that an evaluation can remain meaningful before the next product, model or capability step changes the conditions. When AI helps verify AI, who verifies the verifier?
Security has to be designed in from the start
Essential safeguards should primarily sit in the product, with the people who build it. Users have responsibilities in how they use a system, but customers should not have to become cybersecurity experts simply to compensate for safeguards that could reasonably have been engineered in from the start.
A car maker cannot reasonably sell a faster car and promise that the brakes will arrive in a later update. But it is not enough for the brakes to arrive with the car. We must also have enough time to test that the brakes work before the next model gets an even more powerful engine.
What could help?
No single safeguard eliminates risk. A defensible starting point is multiple independent controls: least privilege, bounded authority, default deny, clear agent identities, auditable logs and provenance. Other measures can include external testing, separation between builder and safety evaluator, red teaming, adversarial testing, staged deployment, blast-radius limits, circuit breakers, rollback and human approval for irreversible or high-impact actions.
The important question is not whether these mechanisms look complete in a presentation. It is whether they can be tested, whether they work under pressure and whether critical controls depend only on the same AI they are meant to constrain.
The hardest security problem may be time
The system being secured cannot also become the final authority deciding that its security is good enough. That does not mean AI should be excluded from security work. On the contrary, it can strengthen it. But final control needs to be independent and intelligible, with real authority to stop the system.
Perhaps the biggest AI security problem is not whether safeguards exist. Perhaps it is whether we still have enough time to test the brakes before the next engine arrives.
💬 What do you think?
What controls would make you trust that AI safety can keep pace with new capabilities?
Please share your thoughts in the comments.
📚 Related articles
Vad är din reaktion?
Gilla
0
Ogilla
0
Kärlek
0
Rolig
0
Wow
0
Ledsen
0
Arg
0
Kommentarer (0)