Trustworthy AI
Trustworthy AI is a broad term for artificial intelligence systems designed to meet a set of quality and ethics-related expectations, such as being reliable, safe, secure, transparent, explainable, and fair. Rather than pointing to a single fixed rule, it describes a collection of characteristics that organizations aim for so that people can reasonably rely on an AI system. The specific characteristics emphasized vary depending on which framework or organization is describing the term.
Trustworthy AI is an umbrella concept encompassing a set of characteristics that AI systems are expected to exhibit; there is no single authoritative definition, and the constituent characteristics differ across frameworks and organizations. As articulated in the NIST AI Risk Management Framework, trustworthy AI characteristics commonly include being valid and reliable, safe, secure and resilient, accountable and transparent, and explainable and interpretable (NIST). Other formulations, such as those from private-sector and academic sources, add or reframe characteristics including fairness and robustness (IBM; Wikipedia). Practitioners should treat these characteristic sets as complementary but non-identical, note that some listed properties (for example explainability versus interpretability, or bias-related fairness) are distinct concepts that should not be conflated, and recognize that 'trustworthy AI' as commonly defined is an aspirational governance framing rather than a binding legal standard. This entry does not resolve the differing definitions or map them to any specific regulatory obligation.
Why it matters
"Trustworthy AI" has become a central organizing concept in AI governance because it gives organizations, regulators, and the public a shared vocabulary for the qualities an AI system should exhibit before people rely on it. Rather than treating reliability, safety, transparency, and fairness as isolated goals, the term bundles them into a set of expectations that can be discussed together at the level of policy, procurement, and oversight. This matters because AI systems are increasingly used in consequential settings where a failure to meet any one of these expectations can undermine confidence in the whole system.
A practical consequence is that the term means different things depending on who is using it. The NIST AI Risk Management Framework lists characteristics such as valid and reliable, safe, secure and resilient, accountable and transparent, and explainable and interpretable. Other formulations from private-sector and academic sources emphasize or add characteristics such as fairness and robustness. Because these characteristic sets are complementary but not identical, an organization citing "trustworthy AI" in one context may be invoking a materially different set of properties than another. Professionals should confirm which framework or source is being referenced rather than assuming a single canonical meaning.
It is equally important to recognize what "trustworthy AI" is not. As commonly defined, it is an aspirational governance framing rather than a binding legal standard, and describing a system as trustworthy does not by itself demonstrate compliance with any specific regulatory obligation. Some of the properties bundled under the term are also distinct concepts that should not be conflated in practice: explainability and interpretability are related but different, and fairness is not the same thing as the absence of bias. Treating the label as a substitute for the underlying, separately assessed characteristics is a common pitfall.
Who it's relevant to
Inside Trustworthy AI
Common questions
Answers to the questions practitioners most commonly ask about Trustworthy AI.