Data and Machine Learning Primer

A deeper technical primer for professionals who want to understand data, models, and evaluation well enough to challenge and govern them substantively.

Fee
£395
CPD hours
28
Duration
5–7 weeks
Level
Foundational

Who this is for

  • Compliance and audit professionals moving into AI governance
  • Risk professionals leading AI risk workstreams
  • Consultants and advisers building AI practices
  • Anyone who found the non-technical foundations too high-level

What you will be able to do

  • Read and reason about datasets, distributions, and data quality
  • Understand the main model families (linear, tree-based, neural, transformer)
  • Distinguish training, validation, testing, and holdout — and why the distinction matters
  • Interpret common evaluation metrics and their trade-offs
  • Identify overfitting, drift, and bias in a model’s reported behaviour
  • Read a technical paper or model card at working depth

Syllabus

Each module comprises a mix of structured reading, worked examples, and applied exercises. Every programme concludes with an integrated written assessment marked against the published rubric.

Module 1. Data

Structured and unstructured. Features and labels. Distributions, sampling, and representativeness. Data quality and its downstream effects.

Module 2. Statistics you actually need

Probability basics, distributions, correlation, hypothesis testing. Just enough to read a paper.

Module 3. The main model families

Linear and logistic regression, trees and forests, gradient boosting, neural networks, transformers. When each is used and why.

Module 4. Training, validation, testing

The workflow. Cross-validation. Data leakage. Why the holdout matters.

Module 5. Evaluation metrics

Classification, regression, ranking, generative. Confusion matrices, ROC/AUC, F1, calibration. Metric choice as a governance question.

Module 6. Overfitting, drift, bias

The three failure modes you will meet most often. Diagnosis and remediation at the practitioner level.

Module 7. Foundation models and LLMs

What is different when the model is generative. Prompt-time evaluation. Hallucination, grounding, retrieval augmentation.

Assessment

Brief

Choose a public dataset or model with sufficient documentation. Produce a 2,000-word technical review covering: the data’s characteristics and limitations, the model or evaluation methodology, the metrics reported and their appropriateness, and the risks a governance function would raise.

Sample question

A model card reports 92% accuracy on the test set and 91% on the validation set. What further questions do you need to answer before trusting these numbers?

Assessments are marked by a named human examiner against the four-dimension rubric: regulatory accuracy (30%), applied judgement (30%), artefact quality (25%), communication (15%). Pass at 60, distinction at 75.

Prerequisites

Foundations of AI for Non-Technical Professionals is helpful but not required. Comfort reading structured technical documentation is assumed.

Certification

On successful completion (pass mark 60), you receive a SAAII Certified Practitioner (CP) — Data and Machine Learning Primer credential. The credential is CPD-accredited, verifiable at thesaaii.com/verify, and forms one component toward higher-tier credentials. See the certification ladder for how it stacks.

Ready to enrol?

Data and Machine Learning Primer runs continuously with rolling enrolment. Founding-cohort discount (25%) applies to the first 100 enrolments across the whole programme portfolio.

Enrol or ask a question →

Cohort licensing available from £395/seat (10+). Public sector, education, and registered charity: 20% discount. Instalment plans available for programmes at £495 and above. See For organisations and the FAQ for detail.