Trustworthy AI
AI with behavior that can be verified, understood, and relied upon.
Trustworthy AI is AI that deserves the trust we place in it, which should be earned through evidence. This type of AI system must provide assurances that scale with capability and risk.
An Astrobee flies through the International Space System using cameras, sensors, and an electric fan propulsion system. These robots are designed to autonomously complete various tasks including taking inventory, documenting, moving cargo, and carrying out experiments. (Image by Suni Williams/NASA.)
The Core Definition
Right now, people can't really trust AI systems, and they shouldn't. Current systems confidently state falsehoods, behave unpredictably, resist explanation, and change their behavior in ways users can't anticipate. The companies building them can't fully explain what the systems will do. Benchmarks measure narrow performance, not reliability in the real world.
Trustworthy AI is different. It means:
AI whose behavior can be verified, understood, and relied upon.
This requires evidence: technical verification, institutional oversight, and legal accountability. The goal is to make AI systems trustworthy so that trust is warranted.
Assurances must scale with capability and risk. A text summarizer needs minimal verification. A medical diagnostic system needs rigorous clinical validation. An AI controlling physical systems in the real world needs extensive safety certification. The higher the capability and the greater the stakes, the stronger the evidence required.
What Trustworthiness Requires
Trustworthiness has several components. A system may be strong on some and weak on others; highly robust trustworthiness requires them all.
Safe and secure.
The system will not cause harm through misuse, accident, or loss of control. It resists attacks and cannot be easily subverted.
Reliable.
The system does what it claims to do, consistently and predictably. Performance is validated for specified use cases, with clear communication about what has been tested and where competence ends. When things go wrong, they go wrong in known and understandable ways.
Transparent.
The system is clear about its capabilities, limitations, operation, and failure modes. It always identifies itself as an AI system and doesn't claim to be something it is not. Users can obtain meaningful explanations appropriate to their needs.
Fiduciary (where applicable).
When the system acts on behalf of a user, it serves that user's interest. This includes loyalty, conflict transparency, no manipulation, and user-inspectable models of user preferences.
Private.
Sound data governance. Users have confidence in how their data is secured and meaningful control over how it is retained and used.
Why Tool AI Is More Trustworthy
Current AI systems are not trustworthy, and their architecture makes trustworthiness difficult to achieve.
The dominant approach starts with an extremely general system trained to predict text, then applies feedback to shape its behavior: rewarding outputs that seem good, penalizing outputs that seem bad. This changes the system's tendencies but cannot produce reliable guarantees.
The feedback process involves real tradeoffs (making the system more cautious in one area may make it less helpful in another, or push problematic behavior into forms harder to detect). The result is a system whose behavior can be nudged but not specified, influenced but not verified.
Purpose-driven tools are different. They're verifiable in ways that general-purpose systems cannot be.
Verification requires specification. To verify a system, you need to know what it's supposed to do. A system with defined scope and objectives can be tested against that specification. A system with open-ended goals and general capability cannot be validated against all possible behaviors.
Narrow scope enables testing. A medical diagnostic tool can be validated against clinical outcomes across defined conditions and populations. A general-purpose agent that might do anything cannot be validated against everything it might do.
Defined objectives enable formal verification. Mathematical proofs of safety properties are possible when objectives are formally specifiable. Open-ended goal-pursuit resists formalization.
Modularity enables component-level assurance. When systems are composed of discrete modules with clean interfaces, each can be verified independently. Monolithic systems resist analysis.
Purpose enables accountability. When something goes wrong with a tool, there's a clear question: did it do what it was supposed to do? With a general-purpose agent, what it was "supposed to do" is much harder to answer.
The current paradigm tries to have it both ways: push general AI into many domains while avoiding the responsibility appropriate to systems in those domains. Trustworthy AI keeps responsibility with humans by giving them tools they can actually verify and understand.
Safety, Security and Robustness as Systemic Properties
Crucially, trustworthiness is a property of the whole system and context, not an intrinsic property of the AI alone.
The current paradigm puts enormous weight on "alignment," treating it as the solution to all issues that aren't raw capability. But alignment as currently practiced means training a general system to behave well through feedback. The method is: show it examples, reward good behavior, penalize bad behavior, and hope the lessons generalize.
This produces shallow compliance. The system learns to avoid triggering penalties in contexts similar to training, but the underlying capability remains, and behavior in novel contexts is unpredictable. Alignment researchers acknowledge these techniques are not robust, can be circumvented, and may produce systems that appear aligned while pursuing other objectives.
Betting everything on getting alignment right is a bad bet.
Tool AI reframes the problem. The AI doesn't have to be trustworthy enough to be safe on its own. The human-AI system provides ongoing correction through multiple overlapping layers: human oversight, architectural constraints, policy controls, scope limitations, and alignment techniques working together.
This approach has several advantages:
- Narrower scope makes risk assessment tractable
- Evaluation depth can scale with capability and risk
- Defense in depth means no single point of failure
- Humans remain in the loop, providing ongoing correction
The most dangerous region is high autonomy combined with high generality and high intelligence, where none of these advantages apply. Tool AI stays out of that region by design.
Prohibited Systems
Some systems should not be built regardless of how trustworthy they might appear.
Superintelligence.
There are very strong reasons to believe that a system substantially more intelligent than humans across all domains cannot be under meaningful human control. The core assumption of oversight (that the overseer can understand and correct the system) breaks down when the system vastly exceeds human capability. Superintelligence is not a capability to regulate; it is a capability to prevent. See Control Inversion for detailed analysis.
AGI as currently pursued.
The race to human-level-and-beyond capability across all domains leads directly to superintelligence and is inherently replacement-oriented. Even "aligned" AGI poses massive concentration risk and disruption.
Intrinsically dangerous architectures.
Regardless of capability level, some designs must be categorically prohibited because the risks hugely outweigh any plausible good reason to pursue them:
- Recursive self-improvement without human authorization
- Autonomous self-replication
- Systems designed to resist shutdown or correction
- Systems designed to escape containment or oversight
- Systems designed to deceive evaluators
Intrinsically dangerous applications.
Some uses are anathema to functioning of a psychologically and socially free society, and thus unacceptable regardless of system capability:
- Mass surveillance infrastructure
- Fully autonomous weapons with lethal authority
- Social scoring systems
- Psychological manipulation systems
- Deceptive AI relationships
These prohibitions are not about current systems being unsafe. They are about categories of systems that cannot be made safe, and uses that are unacceptable regardless of safety.
CT Technician, Amber Kendrick, uses NAETOM Alpha at Birmingham VA. This AI-driven imaging technology helps radiologists perform safer, more accurate scans that reduce patient exposure to radiation and advances early diagnosis. (Image by Birmingham VA.)
Tiered Standards
Assurance requirements should scale with the system's capability profile and risk context. The A+G+I framework provides a useful basis for such tiered standards.
Low capability across all dimensions: Minimal requirements. Standard product liability applies.
High intelligence, low generality, low autonomy (expert tools): Domain-specific certification. Professional standards in the relevant field. Human expert remains accountable for decisions; use restrictions for (rare) high-stakes contexts.
High autonomy, low generality (narrow autonomous systems): Scope verification to ensure boundaries hold. Shutdown reliability. Defined operational envelope. Monitoring for out-of-scope behavior.
High generality, low autonomy (general-purpose tools): Broad evaluation across contexts. Transparency about capabilities and limitations.
Elevated on multiple dimensions: Requirements compound. As systems become more capable across dimensions, assurance burden increases nonlinearly.
Near the dangerous region: Most demanding tier. Formal proofs of controllability where feasible; otherwise, deployment prohibition. Mandatory reporting and regulatory pre-approval. Limited or no safe harbors.
Example: The Verified Medical Diagnostic
Although AI even now has powerful capabilities for medical diagnosis, it cannot and should not be trusted, dramatically limiting its benefit. What would it look like if it were trustworthy?
Clinical validation. Performance validated against clinical outcomes across diverse patient populations. Clear documentation of what conditions it can assess, what populations it's validated for, and where its competence ends.
Uncertainty quantification. Calibrated confidence intervals. The system knows what it doesn't know and communicates uncertainty accurately.
Audit trail. Every recommendation traceable to the data and reasoning that produced it.
Failure mode analysis. Known failure modes documented. Monitoring systems detect when the system may be operating outside its reliable range.
Human-in-loop design. Designed to inform clinical judgment, not replace it. The physician remains accountable; the system is a tool. The part of the physician's job is to understand this tool, as they understand many other tools that they use.
Ongoing monitoring. Post-deployment performance tracking. Automatic alerts when performance diverges from validation benchmarks.
This is what trustworthy AI looks like: verifiable assurance at every level, with trustworthiness emerging from the whole system, not from creating something inherently untrustworthy, slapping a "this system may mess up, trust at your own risk" label on it, and shipping it at a low enough cost that professionals feel pressure to use it.
How Trustworthiness Gets Verified
The framework includes an assurance ecosystem: accredited certifiers, tiered requirements, and ongoing monitoring. Developers demonstrate trustworthiness through formal assurance cases covering safety, reliability, control, and compliance.
This isn't self-certification. Independence requirements, anti-capture mechanisms, and dual certification for high-risk systems ensure the process has integrity.
Deserving Trust
Benchmark scores are not trustworthiness. Marketing claims are not trustworthiness. Impressive demos are not trustworthiness.
Trustworthiness is verified performance, transparent operation, accountable deployment and clear legal responsibility. It's earned through evidence.
A comprehensive approach to AI that keeps humans in charge
The three components
Tool AI
Keeping AI under meaningful human control.
Human-Empowering AI
The principles for AI that makes humans flourish.
Trustworthy AI
How verification and assurance work.
Frequently asked questions
Looking for more context? Start with our FAQs. Can't find what you're looking for? Contact us.