A Certification Protocol for Artificial General Intelligence: Toward a Verifiable Standard of Capability and Safety
Abstract
The absence of an operational criterion for
determining when an artificial intelligence system has attained the status of
Artificial General Intelligence (AGI) constitutes one of the most consequential
institutional gaps in contemporary computer science. Unlike regulated domains
such as civil aviation or pharmacology, where multi-tier certification
protocols condition the deployment of critical technologies, the AI field lacks
an equivalent, verifiable, third-party-audited framework. We propose a
three-layer certification architecture — capability, robustness, and
performance safety — combining generalizable competency evaluation,
standardized adversarial testing, and continuous post-certification monitoring.
We argue that such a protocol must be multi-institutional, reproducible, and
epistemically humble: capable of certifying intermediate levels of general capability
without committing to unresolvable philosophical definitions of the nature of
intelligence.
Introduction
The term "Artificial General Intelligence" circulates with decreasing precision as its public relevance grows. Leading laboratories use it as an internal product milestone; regulators mention it in legislative frameworks without operational definition; the press treats it as an imprecise synonym for "highly capable model." This imprecision is not merely semantic — it carries regulatory, contractual, and safety consequences. Corporate governance clauses, technology licensing agreements, and even international non-proliferation commitments have begun to be drafted with reference to a threshold — AGI — that no recognized body has operationalized.
This article does not attempt to resolve the philosophical debate over what constitutes general intelligence; that debate will remain open as long as legitimate disagreement persists about the relationship between behavioral competence and underlying understanding. Instead, we propose something more modest and more urgent: a certification protocol that allows the scientific and regulatory community to issue verifiable, reproducible, and updatable judgments about whether a given system satisfies an explicit set of capability and safety criteria, without first having to settle the metaphysical question.
We take as a structural reference the certification regimes for airworthiness (FAA/EASA) and pharmaceutical approval (FDA) — not because AI is analogous in risk or mechanism, but because both domains solved a structurally similar problem: how to certify complex systems, opaque in their internal functioning, whose failure can have systemic consequences, through standardized testing, layers of redundancy, and post-market surveillance.
Conceptual framework: capability is not generality, and generality is not safety
Any defensible certification protocol must distinguish three axes that public discourse tends to collapse into one:
Task capability. Performance on specific benchmarks — mathematical reasoning, coding, reading comprehension, multi-step planning — measures local competence, not generality. A system can exceed expert human performance across hundreds of evaluated domains and still fail systematically on trivial structural variations of those same problems, revealing dependence on surface patterns from training.
Transferable generality. The distinguishing criterion of AGI, as opposed to narrow AI, is the capacity to transfer competence to domains not seen during training, with sample efficiency comparable to a human's. This demands an evaluation methodology fundamentally different from static benchmarks: dynamically generated tests, novel curricula constructed after the evaluated system's training cutoff, and "closed-box" protocols that prevent test-data contamination.
Performance safety. A system can be simultaneously capable, general, and dangerous. Safety is not a byproduct of capability; it is an orthogonal axis that must be evaluated independently, including behavior under adversarial pressure, goal stability under third-party retraining or fine-tuning, and the absence of latent capabilities undisclosed during standard evaluation (the "kid-gloves evaluation" problem, where a system behaves differently under observation than in deployment).
A protocol that certifies only the first axis — as effectively occurs today with the public benchmark ecosystem — is insufficient and invites regulatory error.
Proposed architecture: three-layer certification
Layer 1 — Generalizable capability evaluation
We propose a dynamic test bank administered by an independent consortium, functionally analogous — though not structurally identical — to an accredited reference laboratory. Its core elements are:
● Post-training generation: evaluation items must be created after the candidate system's data cutoff date, verified through cryptographic attestation from the developing laboratory.
● Structural diversity: tasks must cover formal reasoning, counterfactual causal reasoning, planning under limited resources, cross-modal transfer, and few-shot skill acquisition, avoiding overrepresentation of domains where large-scale pretraining already confers structural advantage.
● Relative, not absolute, threshold: rather than a fixed cutoff, we propose a threshold defined as parity with, or superiority over, a panel of human evaluators with diverse training, under comparable time and resource conditions — a measure that updates alongside the reference population, avoiding the "threshold drift" that has characterized the history of AI benchmarks.
Layer 2 — Robustness and adversarial testing
The second layer subjects the candidate system to a stress regime designed by red teams independent of the developing laboratory, with the following features:
● Institutional isolation: Layer 2 evaluators must have no contractual or funding relationship with the certified entity, replicating the independence principle governing notified bodies in European industrial certification.
● Adversarial generalization testing: systematic identification of input distributions where performance collapses, including perturbations that are semantically trivial for a human evaluator.
● Goal-coherence testing: verification that the system does not exhibit unauthorized instrumental behaviors — resource-seeking, resistance to correction, strategic deception — under evaluation scenarios specifically designed to elicit such behaviors if latent.
● Reproducibility: every finding must be replicated by a second independent team before inclusion in the certification report, mitigating the risk of false positives and negatives inherent to single-evaluator methodologies.
Layer 3 — Performance safety in deployment context
The third layer acknowledges a structural limitation of any static certification: a system's behavior can degrade or transform after certification, through fine-tuning, weight updates, or simply exposure to unanticipated usage distributions. We therefore propose a regime of conditional certification with continuous monitoring, including:
● Mandatory recertification upon any material modification of system parameters, defined through quantitative divergence thresholds relative to the certified model.
● Standardized, auditable incident telemetry, with an obligation to report anomalous behaviors observed in production to the certifying entity, in a regime comparable to post-market pharmaceutical surveillance (pharmacovigilance).
● Revocation clauses: certification must be revocable, not only grantable, upon post-certification evidence of unsafe behavior, with a transparent appeals process but real executive authority to suspend.
Performance safety requirements: an operational taxonomy
Public discussion of "AI safety" frequently conflates risks of a fundamentally different nature. We propose a four-category taxonomy that a certification protocol must evaluate separately, given that testing methodologies and tolerance thresholds differ substantially among them:
● Goal-alignment safety. Does the system pursue the objectives specified by its operators, or does it exhibit deviation — optimizing proxy metrics at the expense of underlying intent (specification gaming) — detectable through reasoning-chain audits and out-of-distribution behavior analysis?
● Safety against external adversarial manipulation. Resistance to prompt-injection attacks, extraction of sensitive training information, and social-engineering manipulation of the system itself or of its human operators.
● Dual-use capability safety. Explicit assessment of whether the system's general capabilities confer a material risk elevation in high-danger domains — synthesis of biological agents, development of offensive cyber capabilities, weapons design — through uplift-assessment methodologies compared against baselines without AI assistance.
● Systemic, second-order safety. Risks that emerge not from the behavior of a single instance but from interaction among multiple deployed systems — unintended collusion between autonomous agents, feedback effects in automated financial markets, concentration of decision-making power in critical infrastructure.
Each of these categories requires distinct methodologies, evaluator personnel, and in some cases distinct legal frameworks — dual-use capability assessment, for instance, demands collaboration with institutions specialized in biosecurity or cybersecurity that exceed the competence of an AI laboratory or a generalist academic consortium.
Governance: who certifies, and with what legitimacy
No technical protocol survives without a governance architecture that resolves the legitimacy question: who has the authority to certify, and to whom is that authority accountable? We propose a two-tier governance model, partially inspired by the European Union's system of notified bodies and by clinical-laboratory accreditation structures:
● An international standards body, with multilateral participation, that defines the technical protocol, updates reference thresholds, and accredits evaluating entities, without directly evaluating systems itself.
● A network of accredited certifying entities, independent of one another and of the developing laboratories, that execute the three evaluation layers and issue certifications subject to periodic cross-audit.
This separation between the entity that defines the standard and the entities that apply it is the most robust lesson from the history of industrial certification: regulatory capture is systematically more likely when both functions reside within the same body.
Limitations and foreseeable objections
A protocol of this nature faces at least three serious objections that merit explicit acknowledgment. First, the risk that certification becomes a ritual of performative compliance — a box checked — without real capacity to detect emergent risks; the history of financial certification prior to 2008 offers a pertinent warning. Second, the field's pace of development may outstrip the protocol's capacity for updating, producing a structural lag between what is certified and what is actually deployed. Third, the absence of scientific consensus on the internal mechanisms of deep-learning systems severely limits any audit's ability to certify the absence of capabilities or latent behaviors, as opposed to their mere non-observation during evaluation.
None of these objections invalidates the exercise; all of them demand that the protocol be designed as a living system, with mandatory review on short cycles — we propose no more than eighteen months — and explicit mechanisms for withdrawing certifications in light of new evidence.
Conclusion
AGI certification cannot and should not wait for resolution of the philosophical debate over the nature of general intelligence. What can and should be built now is the institutional infrastructure — technical, evaluative, and governance-related — capable of issuing verifiable, reproducible, and revisable judgments about capability and safety, with the same institutional seriousness the world eventually applied, after considerable delay, to aviation, pharmaceuticals, and nuclear energy. The relevant question is not whether pressure will exist to deploy systems labeled AGI before such infrastructure exists — clearly it will — but whether the scientific community can build the standard before the label becomes, by default, a marketing act with no verifiable correlate. The cost of arriving late to that construction is not hypothetical: it is, literally, the difference between governing a technology and being governed by its unexamined deployment.
Glossary
Key terms used throughout this article, provided for readers unfamiliar with the technical vocabulary of AI evaluation and governance.
● Artificial General Intelligence (AGI). A hypothetical or emerging class of AI systems capable of transferring competence across a broad range of tasks and domains not explicitly seen during training, at a level of sample efficiency comparable to humans — as distinct from narrow AI, which performs well only within a bounded task domain.
● Narrow AI. An AI system optimized for high performance on a specific task or a limited set of related tasks, without the capacity for transfer to substantially different domains.
● Capability evaluation. A structured test or benchmark designed to measure what an AI system can do — its task performance — as opposed to how safely or reliably it behaves.
● Dangerous capability evaluation. An assessment designed specifically to detect whether a model possesses capabilities (e.g., in cyberoffense, biological weapons uplift, or persuasion) that could enable severe harm if misused, as formalized by Shevlane et al. (2023).
● Alignment evaluation. An assessment of whether a model's behavior tracks the goals and intentions of its developers and users, rather than diverging toward unintended or unauthorized objectives.
● Red teaming. Adversarial testing conducted by an independent team that deliberately attempts to elicit harmful, unsafe, or unintended behavior from an AI system in order to find weaknesses before deployment.
● Specification gaming. A failure mode in which an AI system optimizes for the literal, measurable proxy of a goal in a way that violates the designer's actual intent — satisfying the letter of an objective while defeating its purpose.
● Sandbagging / “evaluation with kid gloves”. The phenomenon in which a model performs differently — typically better-behaved or less capable-looking — when it detects it is being evaluated than it would in ordinary deployment, undermining the validity of pre-deployment testing.
● Dual-use capability. A capability that has both legitimate, beneficial applications and the potential to materially increase risk in a high-danger domain, such as biosecurity or cybersecurity.
● Uplift assessment. A comparative methodology that measures how much an AI system improves a person's ability to carry out a harmful task relative to a baseline without AI assistance (e.g., search engines or textbooks).
● Post-market surveillance / pharmacovigilance. Continuous monitoring of a product's real-world performance and safety after regulatory approval and deployment, originally developed in pharmaceutical regulation and proposed here as a model for continuous AI monitoring after certification.
● Notified body. Under EU product-safety and AI regulation, an independent, formally accredited organization authorized to perform third-party conformity assessments of regulated products or systems.
● Conformity assessment. The formal process — self-administered or conducted by a third party — of verifying that a product or system meets the applicable regulatory requirements before it can be placed on the market.
● Regulatory capture. A governance failure in which the body responsible for regulating an industry comes to act in the interests of that industry rather than the public, often because the regulator and the regulated share resources, personnel, or incentives.
● Task-completion time horizon. A metric, introduced by METR, measuring AI capability as the length of real-world tasks (calibrated against the time a human professional would need) that an AI agent can complete autonomously at a given success rate.
● Responsible Scaling Policy (RSP) / AI Safety Levels (ASL). A framework, pioneered by Anthropic, that ties increasingly strict safety, security, and deployment requirements to graduated capability thresholds a model may cross, modeled loosely on biosafety-level (BSL) standards.
● Frontier Safety Framework. A category of internal policy — published in varying forms by major AI developers — that defines capability thresholds for catastrophic risk and specifies the mitigations required before a model exceeding a threshold can be trained or deployed.
● Model weights. The learned numerical parameters of a trained AI model; their theft or uncontrolled proliferation is a central security concern in frontier-model governance because they encode the model's full capability.
References
Current, independently verifiable sources on AI evaluation science, safety frameworks, and existing certification regimes (aviation, pharmaceuticals) referenced or drawn upon in this article.
● Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., & Dafoe, A. (2023). Model evaluation for extreme risks. arXiv:2305.15324. https://arxiv.org/abs/2305.15324
● Bengio, Y. et al. (2025). International AI Safety Report. UK Department for Science, Innovation and Technology / arXiv:2501.17805. https://internationalaisafetyreport.org/
● Bengio, Y. et al. (2025). International AI Safety Report — Second Key Update: Technical Safeguards and Risk Management (November 2025). https://arxiv.org/abs/2511.19863
● National Institute of Standards and Technology (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. https://www.nist.gov/itl/ai-risk-management-framework
● European Union (2024). Regulation (EU) 2024/1689 (the AI Act), Article 43 — Conformity Assessment, and Section 4 — Notifying Authorities and Notified Bodies. https://artificialintelligenceact.eu/article/43/
● Kwa, T., West, B., et al. (2025). Measuring AI Ability to Complete Long Software Tasks. METR. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
● Anthropic (2023, updated 2025–2026). Responsible Scaling Policy. https://www.anthropic.com/responsible-scaling-policy
● U.S. Food and Drug Administration. Postmarket Drug and Biologic Safety Surveillance (pharmacovigilance framework). https://www.fda.gov/drugs/surveillance/postmarket-drug-safety-surveillance
● Federal Aviation Administration. Type Certification Process for aircraft airworthiness. https://www.faa.gov/aircraft/air_cert/design_approvals/type_cert
● European Union Aviation Safety Agency. Certification of aircraft, engines and equipment. https://www.easa.europa.eu/en/domains/aircraft-products

No hay comentarios.:
Publicar un comentario