After this chapter, you should be able to
- Distinguish testing, evaluation, verification and validation.
- Build an AI risk register.
- Interpret explanations without causal overclaim.
- Specify monitoring, incident, override and rollback controls.
- Recommend a proportionate pilot state.
Engineering context and methodSource §Lesson 13 · Engineering context and method · GOV-06 · GOV-07 · GOV-08 · GOV-09 · GOV-10 · GOV-11 · RES-10
Testing probes behaviour; evaluation judges results against criteria; verification asks whether the implementation meets its specification; validation asks whether it is fit for the intended context. NIST Govern–Map–Measure–Manage is a voluntary, versioned organising framework and AI RMF 1.0 is under revision. Explanations describe model behaviour under assumptions, not causality. Deployment needs a named owner, versioned data/model, monitored safety and performance thresholds, incident route, override, fallback, rollback, change control and decommissioning.
Verified worked exampleSource §Lesson 13 · Verified worked example · GOV-06 · GOV-07 · GOV-08 · GOV-09 · GOV-10 · GOV-11 · RES-10
Expected-loss screen
Illustrative hazardous false-action probability is 0.02 per exposure with consequence LKR 500,000; a control estimates 0.005.
- Before
0.02 × 500,000
LKR 10,000 per exposure - After
0.005 × 500,000
LKR 2,500 per exposure - Limit
expected value omits intolerability and uncertainty
not a safety acceptance rule
Result. The synthetic expected value falls by LKR 7,500 per exposure. Legal duties and intolerable safety consequences cannot be priced away by this arithmetic.
Practical lab · 6 h lesson effortSource §Lesson 13 · Practical lab · 6 h lesson effort · GOV-06 · GOV-07 · GOV-08 · GOV-09 · GOV-10 · GOV-11 · RES-10
- Red-team an earlier lesson across data, split, thresholds, error slices and out-of-domain states.
- Create a NIST-aligned risk register with owner, control, evidence and residual risk.
- Define monitoring windows, triggers, incidents, override, fallback and tested rollback.
- Recommend no-go, research only, controlled pilot or bounded assistance.
Failure modes to investigateSource §Lesson 13 · Failure modes to investigate · GOV-06 · GOV-07 · GOV-08 · GOV-09 · GOV-10 · GOV-11 · RES-10
- Verification and validation conflated.
- Feature attribution called causality.
- Average monitored while rare failures rise.
- Vendor model changes silently.
- Retraining occurs without controlled revalidation.
Knowledge checksSource §Lesson 13 · Knowledge checks · GOV-06 · GOV-07 · GOV-08 · GOV-09 · GOV-10 · GOV-11 · RES-10
| Question | Answer rationale |
|---|---|
| What asks ‘built as specified’? | Verification. |
| What asks ‘fit for intended context’? | Validation. |
| Does feature attribution prove cause? | No; it attributes model output under selected data and assumptions. |
| Is retraining routine maintenance? | It is a controlled model change requiring review and revalidation. |
| Why define rollback before launch? | Unsafe or degraded operation needs a tested recovery path. |
Key points
- Start from the accountable engineering decision and its consequence.
- Compare against a transparent non-AI baseline.
- Validate on a split that represents intended use and retain human authority.
Source references recorded by the supplied chapter
- NIST AI RMF 1.0 and NIST AI Resource Center; note current revision activity.
- NIST AI 600-1, Generative AI Profile.
- ISO/IEC 23894:2023, AI risk management guidance (metadata and lawful access only).
- EU AI Act and OECD AI Principles.