UnitedHealth's AI Failure: A Governance Lesson in Defining and Measuring Real-World Outcomes
This article highlights the critical distinction between having AI governance structures and achieving true accountability for AI system outcomes, using UnitedHealth's nH Predict tool as a stark example. Internal audit and assurance professionals should note that formal governance alone is insufficient; organizations must proactively define what constitutes AI failure in outcome terms, align incentives, and establish robust monitoring frameworks that account for all affected populations, not just those who can appeal.
The Illusion of AI Governance: UnitedHealth's nH Predict Case
The UnitedHealth case involving its nH Predict AI tool serves as a cautionary tale for organizations deploying AI in high-stakes decision-making. Despite having formal AI oversight structures, including a board committee and a "clinically validated" label, the tool led to a ninefold increase in skilled nursing home denials and a reported 90% error rate. This resulted in significant harm to Medicare patients, a class-action lawsuit, and ultimately, a criminal probe by the Department of Justice. The core issue wasn't a lack of governance frameworks, but rather the absence of a clear, pre-defined understanding of what constitutes failure in real-world outcomes and a mechanism to hold the system accountable to those standards.
Beyond Technical Validation: The Need for Outcome-Based Accountability
The article emphasizes that technical validation of an AI model's performance on benchmark datasets is not equivalent to ensuring acceptable outcomes in the real world. UnitedHealth's internal documents even revealed efforts to predict which denials were likely to be appealed, rather than focusing on reducing errors. This highlights a critical gap: governance frameworks often lack triggers because they don't define in advance what specific, measurable outcomes would necessitate a response. For internal auditors, this means moving beyond simply verifying the existence of controls to evaluating their effectiveness in achieving desired, ethical, and compliant real-world results.
Key Audit Considerations for High-Stakes AI Systems
The UnitedHealth scenario underscores three crucial questions for audit and assurance professionals:
- Defining Failure and Tolerances: Has the board formally defined what constitutes failure for the AI system in outcome terms, not just technical metrics, and documented acceptable tolerances? This includes acceptable error rates and adverse outcome distributions, with clear thresholds for automatic review or suspension.
- Aligning Incentives: Are the incentives of those operating the AI system aligned with the intended outcomes, or do they inadvertently encourage behaviors that compromise those outcomes? Internal audit must scrutinize compensation structures and performance metrics to ensure they don't create pressure to disregard AI outputs that conflict with financial targets.
- Comprehensive Monitoring and Feedback Loops: Does the monitoring framework account for all affected populations, especially those least likely to appeal, and assess structural barriers to feedback? Relying solely on appeal rates can mask systemic errors, particularly in vulnerable populations. Internal audit needs direct access to real-world outcome data, not just summaries, to evaluate whether outcomes align with board-approved tolerances.
Regulatory Landscape and Internal Audit's Evolving Role
The evolving regulatory landscape, including ISO 42001, the EU AI Act, and updated guidance like SR 11-7, increasingly demands demonstrable evidence of effective controls and accountability for AI system outcomes. The Global Internal Audit Standards (GIAS), effective January 2025, further position internal audit as a critical function for evaluating the reliability and appropriateness of AI-assisted processes. The UnitedHealth case serves as a powerful reminder that internal audit must proactively ask the difficult questions about AI's real-world impact, ensuring that organizations are accountable for the outcomes of their AI systems before external regulators or litigation force the issue.
Read more