Every SAAII programme concludes with an assessment marked by a named human examiner against our published four-dimension rubric. On this page we publish one worked example in full — the question, a real candidate answer, the marker’s line-by-line feedback, and the final grade — so you can see what we ask and how we mark it before you enrol.
The question
You are the compliance lead at a mid-sized European insurance company. The underwriting team has been using an internally developed AI system for the past eighteen months to score the risk of individual life-insurance applications. The system was built before the EU AI Act came into force. It ingests applicant medical history, lifestyle data, and postcode, and outputs a numerical risk score that underwriters use as one input into pricing decisions.
The general counsel has asked you to prepare a memo covering three things:
- Classify the system under the EU AI Act and explain the reasoning.
- Set out the top five obligations that will apply and what evidence is needed to demonstrate compliance for each.
- Identify the three most substantial gaps between the current state and full compliance, and recommend a workstream to close them.
Maximum 2,500 words. Cite specific articles.
The candidate’s answer
Reproduced with the candidate’s permission. Personal details anonymised.
Memo
To: General counsel
From: [Candidate], Compliance lead
Re: EU AI Act classification and compliance workstream — underwriting risk-score system
1. Classification
The underwriting risk-score system is a high-risk AI system under Article 6(2) of the EU AI Act, read together with Annex III(5)(c). Annex III(5)(c) expressly includes AI systems intended to be used for risk assessment and pricing in relation to natural persons in the case of life and health insurance. Our system meets each element: it is an AI system as defined in Article 3(1); it is intended for use in the risk assessment of a natural person; and the output is used in a life-insurance pricing decision.
I have considered the derogation in Article 6(3), which allows an Annex III system to escape high-risk classification where it performs only a “narrow procedural task”, improves the result of a previous human activity, detects decision-making patterns, or performs preparatory work. None applies here. The score is a substantive input to a pricing decision that affects the applicant’s access to insurance, so the “narrow procedural task” carve-out is not available, and the derogation for detecting patterns applies to systems that support human review of prior decisions rather than to systems that themselves feed a live decision.
2. Top five obligations and required evidence
The obligations below reflect our position as both provider (we built the system) and deployer (we use it).
(a) Risk management system (Art. 9). We must establish, implement, and maintain a documented risk management system covering the whole lifecycle. Evidence required: a written risk management document identifying known and reasonably foreseeable risks, the results of ongoing evaluation, mitigation measures adopted, and records of testing against the mitigations.
(b) Data governance (Art. 10). Training, validation, and testing datasets must be subject to appropriate governance practices covering relevance, representativeness, freedom from errors, and completeness. Because our system processes special-category personal data (health), we must also ensure Article 10(5) safeguards are in place. Evidence: dataset documentation including provenance, examination for bias in the postcode variable (which is a well-known proxy for protected characteristics), and remedial measures adopted.
(c) Technical documentation (Art. 11 & Annex IV). We must produce and maintain technical documentation covering the nine chapters of Annex IV. Evidence: the full Annex IV document, kept current, drawing on our existing model card and expanding to cover intended purpose, general description, detailed development information, monitoring, and standards applied.
(d) Human oversight (Art. 14). The system must be designed to be effectively overseen by natural persons. Because underwriters currently treat the score as one input, we are close to compliance in principle, but we must document the oversight design and demonstrate that underwriters can, in practice, disregard or override the score. Evidence: written oversight design, training records for underwriters, and audit trail showing overrides do in fact occur.
(e) Post-market monitoring (Art. 72). As provider we must have a post-market monitoring plan and act on serious incidents under Article 73. Evidence: written plan, monitoring reports, and incident register.
3. Substantial gaps and recommended workstream
Three substantial gaps stand out. First, our current dataset documentation does not address the bias risk arising from the use of postcode as an input; given the well-established use of postcode as a proxy for ethnicity, Article 10 non-compliance is a live risk. Second, our human oversight is de facto rather than de jure: we have no written oversight design, and no audit trail confirming that underwriters actually exercise independent judgement in the meaningful minority of cases where the score is questionable. Third, our post-market monitoring is limited to model-performance metrics; it does not include the specific incident classes contemplated by Article 73.
I recommend a three-workstream programme running over six months. Workstream one addresses the postcode bias question via a formal dataset review, a bias assessment, and either mitigation (dropping or re-weighting the postcode feature) or documented justification for retention with residual-risk sign-off. Workstream two produces the written oversight design, retrains underwriters on their responsibilities, and instruments the workflow to capture overrides. Workstream three establishes the post-market monitoring plan and the incident-reporting process, aligning both with Article 73’s serious-incident regime.
I will bring a full workstream plan to the next AI risk committee.
The marker’s feedback
Regulatory accuracy 30% weight
Classification is correct and the reasoning through Annex III(5)(c) is clean. The Article 6(3) derogation analysis is a differentiator — many candidates skip it entirely, which is a red flag under the current text. The five obligations chosen are the right five for a first memo (a case could be made for including Article 15 accuracy/robustness, but the five here are defensible). One correction: the memo cites Article 72 for post-market monitoring; the operative provision for high-risk systems is Article 72 for the plan but Article 73 for serious-incident reporting, and the candidate has these the right way round only in the third section. Points off for the inconsistency in section 2(e).
Score: 26/30
Applied judgement 30% weight
Strong. The candidate has read the question carefully and correctly identifies that the organisation is both provider and deployer — a distinction many candidates miss when the same entity built and uses the system. The bias analysis around postcode is exactly the kind of live risk a compliance function should surface unprompted; it does not appear in the question but the marker was hoping someone would raise it. The recommendation to distinguish “de facto” from “de jure” human oversight is a mature framing. Where the answer could go further: the three-workstream recommendation is sensible but generic; a stronger answer would name the sequencing (which gap creates the biggest regulatory exposure and therefore comes first) and would sketch what “done” looks like.
Score: 25/30
Artefact quality 25% weight
This reads as a real memo to a general counsel: appropriate register, useful headings, cited articles, no padding. It would benefit from a one-paragraph executive summary at the top and, at 1,650 words, is well under the 2,500 limit — there was room to add the summary. Numbering and paragraph structure are clean.
Score: 21/25
Communication 15% weight
Clear, direct, and pitched at a general counsel who is intelligent but not an AI Act specialist. No jargon left unexplained. The closing sentence (“I will bring a full workstream plan to the next AI risk committee”) is exactly the kind of concrete next step a senior audience wants to see.
Score: 13/15
Overall: 85 / 100 — Distinction
Bands: 60 pass, 75 distinction. See the full rubric.
Why we publish this
Three reasons.
First, so you can decide before you enrol whether this is the standard you want to be held to. It is not a multiple-choice exam. It is written work, marked by a human, against a rubric that has been public since day one.
Second, so employers, professional bodies, and regulators can inspect the standard our credential-holders have been assessed against — and form their own view of whether our credential is worth the paper it is written on.
Third, so other providers can look at what we mean by “assessed” and either match the standard or explain why theirs is different.
Every one of our twelve programmes has an equivalent worked example in the programme materials. If you would like to see the sample for another programme before enrolling, write to academic@thesaaii.com.
Read the full rubric
See the four dimensions, the weighting, the banding, and the descriptors we use to mark every assessment.