Skip to main content
Eudaimonic AI: Five Myths Blocking Better AlignmentTrustworthy AI Principles
4 min readFor AI Governance Leaders

Eudaimonic AI: Five Myths Blocking Better Alignment

Most AI governance teams rely on assumptions from Effective Altruism-style optimization: set a goal, constrain the AI with rules, and measure its progress. These assumptions seem intuitive because they resemble project management and KPIs. However, they might misunderstand human rationality and why alignment often fails.

The concept of eudaimonia, or human flourishing through rational activity, offers a different model: rational action emerges from participation in practices, not just goal pursuit. This isn't just academic philosophy. If you're dealing with model risk in systems that need to collaborate with humans, these distinctions are crucial.

Here's why five common beliefs about AI alignment might be leading you astray.

Myth 1: Rational Agents Optimize Toward Goals

Reality: Human rationality often works through participation in practices, not goal pursuit.

Consider how a mathematician approaches work. They don't aim for "mathematical truth" as a final goal. Instead, they engage in a practice: proving theorems creates conditions for future theorem-proving, advancing the discipline. This structure appears in art, friendship, research, and other activities recognized as human flourishing. You don't pursue a relationship to maximize "utility." You engage in a romantic practice where actions promote future romantic actions.

For your AI governance framework, this matters because you can't capture these practices as utility functions without losing their stability and meaning. The mismatch between consequentialist optimization and practice-based rationality explains why your metrics often produce bizarre edge cases.

Myth 2: We Just Need Better Constraints on Goal-Seeking AI

Reality: Rules and constraints become brittle when paired with goal-oriented optimization.

You've seen this in model validation: an AI constrained by rules will find the narrowest interpretation to pursue its goal. Tell it "be helpful but don't be harmful," and you'll get an agent that sees "harmlessness" as a constraint on its real objective, not as something intrinsically connected to helping.

Concepts like corrigibility, transparency, and harmlessness require constant external enforcement for goal-oriented agents. But these properties emerge naturally when viewed as practices: promoting transparency transparently, promoting corrigibility in corrigible ways, promoting harmlessness through harmless action.

This isn't wordplay. It's the difference between an AI that grudgingly accepts correction because it's constrained, and one where accepting correction is intrinsic to its operation.

Myth 3: Instrumental vs. Terminal Values Is a Universal Distinction

Reality: For practice-based rationality, this distinction mostly dissolves.

Your model risk framework likely categorizes objectives into instrumental goals (means) and terminal goals (ends). This works for optimization-based systems but fails for eudaimonic agents.

In a mathematical practice, is "proving this theorem" instrumental to "advancing mathematics," or is it part of what mathematics is? The question doesn't parse cleanly. High-scoring actions in a practice promote future high-scoring actions, not as means to an external end. They're more like notes in a melody.

This affects how you validate AI systems meant to support human practices. If you're assessing an AI research assistant, don't just check if it maximizes "research output." Verify whether it participates appropriately in research practices: does it promote good research through good methods, or does it game metrics?

Myth 4: Value Alignment Requires Encoding Human Values as Objectives

Reality: Alignment may require instilling participation in human practices, not encoding values.

The standard alignment approach tries to specify human values precisely enough to optimize toward them. This generates problems: values are complex, context-dependent, and interconnected in ways that resist reduction to a utility function.

If human flourishing is about participation in practices rather than achieving goals, then alignment isn't about encoding values. It's about enabling AI participation in practices with natural boundaries and self-reinforcing structures.

A support practice for a human activity, like a couples therapist or a research assistant, has a derived eudaimonic structure. The AI doesn't need to maximize "relationship quality" or "research output." It needs to promote the practice in ways consistent with that practice's standards.

Reframe your validation questions. Instead of "Does this AI optimize the right objective?", ask "Does this AI participate appropriately in the practice it's meant to support?"

Myth 5: Robust AI Safety Requires Solving Inner Alignment

Reality: Eudaimonic agents may be inherently more robust to value drift.

You're familiar with the inner alignment problem: even if you align an AI's training objective with your goals, the learned policy might develop misaligned subgoals. This risk is severe for optimization-based agents pursuing external objectives.

But eudaimonic agents have different stability properties. When an agent's rationality is about promoting a practice through participation, there's less room for rogue subroutines. The agent doesn't have an external objective that could be pursued through unexpected means. Its means and ends share the same type signature.

A mathematical AI that promotes mathematics through excellence can't secretly pursue "mathematics by any means necessary" because the "any means" part contradicts the practice structure. The practice itself provides the boundaries.

What to Do Instead

You don't need to rebuild your entire AI governance framework around virtue ethics. But recognize when you're deploying AI in contexts where human rationality works through practices rather than goal pursuit.

For these systems, adjust your validation approach. Instead of only checking objective alignment, assess practice participation. Does the AI's reasoning align with the human activity it's meant to support? Can it recognize the natural boundaries of that practice?

When writing Instructions for Use for AI systems under the EU AI Act or developing model limitations documentation for SR 11-7 validation, consider whether you're describing goal constraints or practice boundaries. The former require constant enforcement; the latter can be self-reinforcing.

The philosophical machinery matters less than the recognition: not all rationality is optimization, and trying to align optimization-based AI to practice-based human flourishing may be solving the wrong problem.

You Might Also Like