Navigating AI Deployment: Technical Risks and Audit Considerations for Live Systems
This article highlights critical technical risks that emerge when AI models transition from controlled development to real-world deployment. For internal audit and assurance professionals, understanding these risks—such as training-serving skew, lack of human oversight, unchecked generative AI outputs, and adversarial inputs—is crucial for evaluating the robustness of AI systems and ensuring organizational accountability. The Air Canada chatbot case serves as a stark reminder that organizations are fully responsible for their AI's actions, underscoring the need for robust controls and clear accountability frameworks in AI deployment.
The Chasm Between Development and Deployment
The transition of an AI model from a controlled development environment to live deployment introduces a host of technical risks that often go unaddressed. In development, data is curated, inputs are standardized, and the system operates under ideal conditions. However, real-world deployment exposes the model to messy, unpredictable inputs, adversarial users, and unforeseen circumstances. This fundamental shift can lead to significant performance degradation and unexpected behaviors, challenging the assumption that a model proven in development will perform identically in production. Internal auditors must recognize this inherent gap and scrutinize the processes in place to manage the 'shock' of real-world interaction.
Key Technical Risks in AI Deployment
- Training-Serving Skew: A primary risk is the divergence between the data used for training and the data encountered in production. Subtle differences in data formats, sources, or characteristics can silently degrade model performance. Auditors should look for evidence of continuous monitoring and comparison between production and validation environments to ensure the model operates on inputs it was designed for.
- Lack of Human Oversight and Fallback Mechanisms: Systems deployed without human review, override capabilities, or safe default states are inherently risky. When an AI system fails, it should ideally degrade gracefully, perhaps by deferring to a human or indicating uncertainty. The absence of such mechanisms means every error can directly lead to harmful consequences. Audit teams should assess whether human intervention points are proportionate to the stakes of the AI's decisions and if clear fallback protocols are established.
- Uncontrolled Generative Output (Hallucination): Generative AI models can produce confident, yet entirely false, outputs. When these outputs reach customers without validation or guardrails, organizations face significant reputational and legal liabilities. The Air Canada chatbot case exemplifies this, where the airline was held responsible for its chatbot's erroneous advice. Auditors must verify the presence of validation layers between generative AI outputs and end-users, ensuring the organization explicitly owns and stands behind the information provided by its AI systems.
- Exposure to Adversarial Input: Unlike controlled testing, deployed AI systems are vulnerable to malicious attacks, such as prompt injection in language models. These attacks exploit the model's inability to distinguish trusted instructions from untrusted data, potentially leading to data exfiltration or system manipulation. Audit procedures should include assessing the system's exposure to such manipulations and the controls in place to constrain the actions of a compromised model.
Audit Implications and Organizational Accountability
Each of these technical risks stems from the inability of a controlled environment to replicate the complexities and adversities of the real world. The Air Canada case serves as a critical precedent, demonstrating that organizations are unequivocally accountable for the actions and outputs of their deployed AI systems. The notion of an AI acting as a separate legal entity is not viable. Therefore, internal audit and assurance professionals must ensure that robust controls, continuous monitoring, and clear lines of accountability are established for all AI systems in deployment. This includes verifying that guardrails are in place to prevent erroneous or malicious outputs from reaching users and that the organization has a defined strategy for managing and mitigating these technical risks.
Read more