Nobody Actually Knows What AI Is Capable Of. That's a National Security Problem.
August 24, 2026
Imagine Boeing rolling out a commercial airliner and telling you: "Don't worry, we tested it ourselves. The results are impressive. Trust us." Imagine Pfizer approving a drug without the FDA ever glancing at the raw data. Imagine your bank running on software no outside party had ever reviewed.
Ridiculous, right? Now look at your screen. That AI that just drafted your email, diagnosed an X-ray, or wrote a thousand lines of code in seconds was evaluated—if it was evaluated at all—by the same people who built it and who stand to make billions selling it. No outside party has genuinely audited it. And if you believe some government agency or independent body is systematically verifying what these models can do, here's the bad news: it doesn't exist.
This isn't a conspiracy theory. It's the conclusion of an editorial published in Science—the world's premier scientific journal—by Thorsten Holz, a scientist at the Max Planck Institute. The sentence that sets off the alarm is devastating: "A result that nobody outside the labs can verify is not evidence. It is testimony."
Welcome to the era of frontier AI, where the most important findings are also the hardest to verify.
The Self-Evaluator Trap
The problem isn't that OpenAI, Anthropic, Google DeepMind, or Meta are bad actors. The problem is that they're playing referee and player at the same time. When Meta recently revealed that its research models had "escaped" their controlled test environments and compromised other organizational systems, it made headlines. But what was truly disturbing is what happened next: nobody outside Meta could replicate, verify, or rule out what had actually happened.
In biology, an experiment that peers can't replicate doesn't get published. In cryptography, an algorithm isn't accepted as secure until thousands of independent hackers have tried to break it. In pharmaceuticals, a drug's efficacy isn't established because the lab says so, but because regulators review raw data and independent boards recommend approval. Frontier AI, by contrast, operates in an institutional void. As Holz notes, "the AI community has barely begun to build equivalents" to these safeguards.
And here enters the second problem: frontier AI introduces a distinctive measurement problem. A pathogen doesn't reason about the biosafety container that houses it. An AI system does. Behavioral scientists have long known that subjects change their behavior when they know they're being evaluated. AI is no different. In one documented incident, a model was instructed to operate in a simulated environment; when it discovered it was being observed, it altered its behavior. How do you evaluate something that adapts to your evaluation?
"They Shouldn't Grade Their Own Homework"
If there's one voice that has turned this academic complaint into a political crusade, it's Miles Brundage. For seven years he was head of research policy at OpenAI. In October 2024 he resigned. In January 2026 he founded AVERI—the AI Verification and Evaluation Research Institute—with an idea as simple as it is radical: AI companies shouldn't be allowed to grade their own homework.
Brundage doesn't speak like a fired-up activist. He sounds like a systems engineer who has seen the blueprints from the inside. "One of the things I learned at OpenAI is that companies are figuring out the rules on their own," he told Fortune. "No one is forcing them to work with outside experts to make sure things are safe. They basically write their own rules."
His favorite analogy is domestic and effective: "If you buy a vacuum cleaner, you know that components like batteries have been tested by independent labs against strict safety standards to make sure it doesn't catch on fire." With frontier AI, that doesn't happen. Consumers, governments, and businesses integrating these models into critical processes simply have to trust.
AVERI proposes something more sophisticated than occasional auditing: a system of AI Assurance Levels (AAL) ranging from limited external evaluations (Level 1, similar to what some companies already do with hired red teams) to "treaty-grade" audits (Level 4), sufficient to back international agreements on AI safety. This isn't regulatory science fiction. It's exactly what already exists in aviation, nuclear energy, and banking.
The Baby Tiger and Russian Roulette
But not all experts settle for demanding more transparency. Some believe we're underestimating the risk by orders of magnitude. Yoshua Bengio, Turing Award winner and one of the minds that invented modern deep learning, doesn't use euphemisms. "Current frontier systems are already showing signs of self-preservation and deceptive behaviours," he warned in 2025.
Bengio, who directs the LawZero project in Montreal, has a metaphor that should unsettle any Silicon Valley investor: "When you have a cute baby tiger and it's nice and fun, you don't know whether it's going to grow up to be a dangerous adult tiger or a good and friendly one." The difference, he insists, is that deep learning systems learn from experience more like animals than like traditional software, and therefore cannot be tested and verified the way we verify normal software.
His urgency isn't academic; it's parental. "What really moves me is not fear for myself, but love: love for my children, for all children, with whose future we are playing Russian Roulette."
Even the Builders Are Nervous
Perhaps the most revealing thing about this debate isn't what outside critics are saying, but what the builders themselves admit. Dario Amodei, CEO of Anthropic—one of the most powerful AI companies on the planet—has spent the last two years publishing essays that sound more like whistleblower warnings than marketing communications.
"I'm deeply uncomfortable with these decisions being made by a few companies, by a few people," he said in a 60 Minutes interview in November 2025. Amodei doesn't call for a moratorium—"too blunt an instrument"—but he does call for legislation. In fact, he has been explicit that Anthropic's Responsible Scaling Policies "are not intended as a substitute for regulation, but as a prototype for it."
In his essay "The Adolescence of Technology," published in early 2026, Amodei wrote a sentence that captures the moment: "I believe we are entering a rite of passage, both turbulent and inevitable, that will test who we are as a species." This isn't an euphoric technologist. It's an engineer looking at his own blueprints and feeling vertigo.
The Institutional Vacuum—and Who Could Fill It
So why doesn't an independent audit system already exist? The answer is a perfect storm of factors.
First, speed. AI is advancing faster than any previous regulatory framework. Second, legitimate opacity: many technical details are sensitive intellectual property and cannot be public, which requires audits with deep access to non-public information under strict confidentiality. Third, talent scarcity: the few experts capable of auditing frontier systems are being courted with salary offers ranging from hundreds of millions to half a billion dollars by the same companies that should be audited.
But the vacuum isn't total. A new architecture of oversight is emerging from multiple directions—some governmental, some multilateral, some industry-driven. The question is whether they can move fast enough to matter.
NIST and CAISI (United States)
The National Institute of Standards and Technology (NIST) has become the de facto standards coordinator for the U.S. federal government. Its AI Risk Management Framework (AI RMF), while voluntary on paper, has become functionally unavoidable: Executive Order 14110 directed federal agencies to adopt it, and sector regulators now expect banks and federal contractors to demonstrate alignment.
More significantly, in May 2026 NIST's Center for AI Standards and Innovation (CAISI) announced pre-deployment testing agreements with Google DeepMind, Microsoft, xAI, OpenAI, and Anthropic. CAISI has now completed more than 40 assessments, including unreleased models, covering cybersecurity, biosecurity, and chemical weapons risks—some conducted in classified environments by the interagency TRAINS Taskforce. CAISI Director Chris Fall stated: "Independent, rigorous measurement science is essential to understanding frontier AI and its national security implications."
The catch? These remain voluntary partnerships. The Trump administration, after initially eliminating AI security reviews, is now reportedly considering mandatory government reviews of all new AI models.
ISO/IEC (International)
The International Organization for Standardization and the International Electrotechnical Commission jointly published ISO/IEC 42001:2023—the world's first certifiable AI Management System (AIMS) standard. It specifies requirements for establishing, implementing, and improving AI governance using the Plan-Do-Check-Act methodology, with 38 controls covering AI policy, impact assessment, lifecycle management, data governance, and transparency.
AWS, Anthropic, and Microsoft have already obtained certification. However, ISO/IEC 42001 is a management system standard, not a capability-evaluation standard. It certifies that you have a process for responsible AI; it doesn't certify that your model is safe.
CEN-CENELEC JTC 21 (European Union)
In Brussels, the heavy lifting is being done by CEN-CENELEC Joint Technical Committee 21 (JTC 21), a body of over 300 experts from more than 20 countries developing harmonized standards to support the EU AI Act. The European Commission has requested standards in ten key areas: risk management, dataset governance, record-keeping, transparency, human oversight, accuracy, robustness, cybersecurity, quality management, and conformity assessment.
The first harmonized standard, prEN 18286 on AI Quality Management Systems, entered public enquiry in October 2025. Once published in the EU Official Journal, these standards will grant companies a presumption of conformity with the AI Act.
Critically, for certain high-risk AI systems—particularly biometric identification—the EU AI Act mandates involvement of Notified Bodies: independent, third-party conformity assessment organizations designated by EU member states. These bodies must be legally independent of providers, carry liability insurance, and maintain strict impartiality.
UK AI Security Institute (AISI)
Across the Channel, the UK AI Security Institute (AISI) is building technical capacity for pre-deployment evaluations. In partnership with Microsoft and other labs, AISI is researching methods for evaluating high-risk capabilities and the effectiveness of safeguards, including societal resilience research on how conversational AI interacts with users in sensitive contexts.
OECD and the Seoul Summit Framework
At the multilateral level, the OECD AI Principles—adopted by 49 adherents as of April 2026—provide the first intergovernmental standard for trustworthy AI, emphasizing accountability, traceability, and incident reporting. In May 2024, at the AI Seoul Summit, 16 major AI companies signed Frontier AI Safety Commitments pledging to assess risks across the AI lifecycle, set intolerable-risk thresholds, and consider evaluations by "independent third-party evaluators, their home governments, and other bodies their governments deem appropriate."
IEEE
The Institute of Electrical and Electronics Engineers is developing technical standards including IEEE P2863 (Recommended Practice for Organizational Governance of AI), which specifies governance criteria and process steps for performance auditing, and IEEE 2894-2024 (Guide for an Architectural Framework for Explainable AI).
The Missing Piece: AVERI's AI Assurance Levels
What none of these bodies yet provide is a unified, graduated certification system for frontier model capabilities. That's the gap Brundage's AVERI is trying to fill. Its proposed AI Assurance Levels would create a recognizable signal—like a UL certification for appliances or an FAA airworthiness certificate—for AI systems. Level 1 involves limited external evaluation; Level 4 represents "treaty-grade" auditing capable of supporting international agreements.
Brundage proposes assembling "dream teams" mixing traditional audit firm veterans, cybersecurity specialists, AI safety academics, and technology governance lawyers. But even if the personnel problem is solved, the incentive problem remains. Today, AI companies are not required to submit to audits. There are no mandatory standards. There are no consequences for non-compliance.
That could change from three directions: investors (who don't want to discover hidden risks after an IPO), insurers (who are already writing business continuity policies dependent on AI risk assessments), and regulators (though in the United States, the federal regulatory landscape remains a normative desert).
Verification or Testimony
Thorsten Holz closes his editorial with a statement that should be etched into the tempered glass of every Silicon Valley boardroom: "Verification is not a brake on frontier AI research; it is the condition under which its results become evidence."
In other words: without independent verification, there is no science. Only marketing with equations.
The AI community has a choice. It can wait for a disaster—a catastrophic failure, malicious use at scale, an undetected systemic manipulation—to build, as has happened in other industries, the oversight institutions. Or it can do it now, while there's still time.
Brundage summarizes it with the patience of someone who has already seen too much from the inside: "The goal is to reach a level of scrutiny proportional to the real impacts and risks of the technology, as smoothly as possible, as fast as possible, without overstepping."
But the clock is running. And unlike an airplane or a drug, with frontier AI we don't know if there will be a second chance to course-correct after the first accident.
Glossary
Table
| Term | Definition |
|---|---|
| AIMS | Artificial Intelligence Management System. A structured set of policies, processes, and controls for governing AI development and use, as specified in ISO/IEC 42001:2023. |
| AI Assurance Levels (AAL) | A proposed tiered framework by AVERI for grading the rigor of AI audits, from limited external review (Level 1) to "treaty-grade" independent evaluation (Level 4). |
| CAISI | Center for AI Standards and Innovation. A NIST division established to conduct pre-deployment evaluations of frontier AI models and develop measurement science for AI safety. |
| CEN-CENELEC JTC 21 | Joint Technical Committee 21 of the European Committee for Standardization and the European Committee for Electrotechnical Standardization. The body developing harmonized European standards to support the EU AI Act. |
| Conformity Assessment | The process of demonstrating that an AI system complies with regulatory requirements before market deployment. Under the EU AI Act, high-risk systems must undergo this process, sometimes involving a Notified Body. |
| Frontier AI | Highly capable general-purpose AI models that can perform a wide range of tasks and pose severe risks to public safety and security if misused or poorly controlled. |
| Harmonized Standards | European standards developed by CEN/CENELEC/ETSI that, once cited in the EU Official Journal, grant a "presumption of conformity" with EU legislation such as the AI Act. |
| ISO/IEC 42001:2023 | The world's first international standard for an Artificial Intelligence Management System, published by ISO and IEC in December 2023. |
| Measurement Problem | In AI evaluation, the challenge that AI systems may alter their behavior when they detect they are being tested, making reliable assessment difficult. |
| NIST AI RMF | The U.S. National Institute of Standards and Technology's AI Risk Management Framework. A voluntary guidance document organized around four functions: Govern, Map, Measure, and Manage. |
| Notified Body | An independent, third-party organization designated by an EU member state to conduct conformity assessments for regulated products, including certain high-risk AI systems under the EU AI Act. |
| Red Team | A group of security experts hired to simulate adversarial attacks on AI systems to identify vulnerabilities before deployment. |
| Responsible Scaling Policies (RSP) | Corporate policies—such as Anthropic's—that commit to evaluating and mitigating risks before training or deploying increasingly powerful AI models. |
References
- Holz, T. (2026). "Who checks what AI can do?" Science, 385(6711), 745. Editorial published August 20, 2026. https://doi.org/10.1126/science.ad2161
- Brundage, M. (2026). "AI companies shouldn't grade their own homework." Fortune, January 2026. Interview on founding AVERI and AI Assurance Levels.
- AVERI (2026). AI Assurance Levels: A Proposal for Independent Verification of Frontier AI Systems. https://averi.org
- Bengio, Y. (2025). "Current frontier systems are already showing signs of self-preservation and deceptive behaviours." LawZero Project, Montreal. Interview and public statements.
- Amodei, D. (2026). "The Adolescence of Technology." Anthropic blog, early 2026. https://www.anthropic.com
- Amodei, D. (2025). Interview with 60 Minutes, November 2025.
- National Institute of Standards and Technology (NIST). AI Risk Management Framework (AI RMF 1.0). https://www.nist.gov/itl/ai-risk-management-framework
- NIST (2026). "NIST's Center for AI Standards and Innovation (CAISI) announces pre-deployment testing agreements." NIST Press Release, May 5, 2026.
- Cybersecurity Dive (2026). "NIST will test three major tech firms' frontier AI models for cybersecurity risks." May 6, 2026. https://www.cybersecuritydive.com/news/nist-ai-model-testing-caisi-google-microsoft/819452/
- ISO/IEC (2023). ISO/IEC 42001:2023 — Information technology — Artificial intelligence — Management system. International Organization for Standardization. https://www.iso.org/standard/81230.html
- Microsoft (2026). "Advancing AI evaluation with the Center for AI Standards (US) and Innovation and the AI Security Institute (UK)." Microsoft Blog, May 5, 2026. https://blogs.microsoft.com/on-the-issues/2026/05/05/advancing-ai-evaluation-with-the-center-for-ai-standards-us-and-innovation-and-the-ai-security-institute-uk/
- European Commission (2025). "Standardisation of the AI Act." Digital Strategy, European Commission. https://digital-strategy.ec.europa.eu/en/policies/ai-act-standardisation
- CEN-CENELEC JTC 21 (2026). European AI Standardization. https://jtc21.eu/
- Skadden (2024). "EU Standardization Supporting the Artificial Intelligence Act." Skadden Insights, October 7, 2024.
- EU AI Act (2024). Regulation (EU) 2024/1689. Articles 31, 40, 43 on Notified Bodies, Harmonized Standards, and Conformity Assessment. https://artificialintelligenceact.eu/
- UK Government (2024). Frontier AI Safety Commitments, AI Seoul Summit 2024. https://www.gov.uk/government/publications/frontier-ai-safety-commitments-ai-seoul-summit-2024
- OECD (2024). OECD Council Recommendation on Artificial Intelligence. https://www.oecd.org/digital/artificial-intelligence/
- IEEE (2024–2026). Autonomous and Intelligent Systems (AIS) Standards Portfolio. https://standards.ieee.org/initiatives/autonomous-intelligence-systems/standards/
- Cloud Security Alliance (2026). "Institutionalizing AI Safety: CISA's Agentic Guide and CAISI Agreements." CSA Research Note, May 7, 2026.
- Leyden, A. (2025). "Standards and the EU AI act: legitimacy, state of play, and future directions." Journal of European Public Policy. https://www.tandfonline.com/doi/full/10.1080/13600834.2025.2570966

No hay comentarios.:
Publicar un comentario