Conducting: evidence over impressions
Lesson 3 of 5 in Lead Auditor Track: Auditing an AIMS with ISO 19011 Discipline.
Fieldwork begins with the opening meeting, and running it well is choreography, not ceremony. In ten minutes you: introduce the team (including any technical expert and their role), reconfirm objectives, scope, and criteria, walk the schedule and flag any changes, agree communication channels and guides, confirm confidentiality and any safety rules, state how findings will be classified and communicated, remind everyone that the audit works by sampling — so a clean audit is not a guarantee of perfection — and set the closing meeting time. Two sentences matter more than the rest: findings will be shared as they emerge, so nothing at the closing meeting will be a surprise, and the audit examines the system, not individual performance. The first buys you dispute-free closings; the second buys you honest interviewees.
Then the real work: interviews, records, observation — triangulated. A procedure says what should happen; an interview says what people believe happens; a record says what actually happened. One source is an anecdote. Two is a lead. Three that agree is evidence.
Ask open, then chase the evidence
Closed questions get you compliance theatre (‘Do you assess risk before deployment?’ — ‘Yes.’). Open questions get you the system as lived: ‘Walk me through what happened the last time a model went to production.’ Then chase: Show me that risk assessment. Who signed it? Where is the sign-off recorded? What did you do about the third risk on the list? Every claim earns a ‘show me’. The chase ends at a record or at an admission — both are evidence.
Audit up and down the organisation
Auditing only the governance team audits the paperwork’s authors. Go up: interview top management on the AI policy, resourcing, and management-review decisions — clause 5 commitments are their commitments, and an executive who cannot say what the AI policy commits to is itself evidence. Go down: ask the ML engineer what happens when a drift alert fires, ask the data annotator how labelling disputes get resolved. Discrepancy between floors is where paper systems come apart.
Nervous auditees
Most interviewees have never been audited and assume they personally are on trial. Lower the temperature: explain you are testing the system, ask them to show you their normal work, start with easy ground. A nervous engineer who relaxes will show you the real workflow, including the unofficial spreadsheet that bypasses the approved pipeline — the most valuable find of the week.
Obstruction and the over-helpful guide
Delays that consume interview slots, ‘the record owner is on leave’, a guide who answers every question addressed to the operator — treat all three as flow-control, and manage them identically: note it factually, restate the request in writing with a deadline, escalate to the auditee’s management contact. If access is still refused, say plainly what follows: evidence not provided is evidence absent — the conclusion is drawn on what was seen, and a scope limitation may itself end the audit. Never negotiate evidence away; never lose your temper. The notebook wins the argument later.
Note discipline
Your notes are the audit’s legal spine. Record what you saw, verbatim where it matters: document titles and versions, record identifiers, dates, system names, who said what, sample sizes (‘requested 5 of 23 impact assessments; received 3’). Not impressions (‘seemed disorganised’), not adjectives — retrievable facts. When a finding is disputed three weeks later, the auditor with record IDs wins; the auditor with vibes retracts.
Sampling with intent. You cannot read everything, so every choice of what to read is a bet — make the bets consciously. Risk-based sampling means the sample leans toward: the highest-impact AI systems, whatever changed recently (new model, new supplier, retraining), whatever the last audit flagged, and the records the auditee did not volunteer. Always record the sampling frame (‘23 deployed systems; sampled 4: the two highest-risk per the auditee’s own register, one recent acquisition, one selected at random’) — it is what makes your conclusion reproducible, and what protects you when someone asks why you missed system number 19. And keep the humility the evidence-based principle demands: a clean sample supports ‘no nonconformity found’, never ‘no nonconformity exists’.
Auditing an impact assessment
The claim: ‘we assess AI system impacts per 6.1.4’. The chase: pick one real, high-risk system and pull its assessment. Check the lens — does it analyse consequences for individuals, groups, and societies, or is it a relabelled security risk assessment (the failure the previous module taught you to spot)? Check timing — performed before deployment, and re-performed after the significant change the change log shows in March? Check consequence — did any identified impact change a decision: a mitigation, a rollout condition, an entry in the risk register? An assessment that never altered anything downstream is a filing exercise, and you should say so with the record IDs to prove it.
Auditing a bias-testing claim
The claim: ‘the hiring model is bias-tested’. The chase: protocol, metrics, results, thresholds, consequences. Which fairness metrics were chosen, and is the choice justified against the use case — or was the metric chosen after the fact because the model passed it? Are results disaggregated by the relevant subgroups? What threshold triggers action, who owns the decision, and — the question that kills most claims — was the test re-run after the last retraining? A bias test performed once, on a model version three retrainings ago, is evidence of a control that existed, past tense. Ask for the test artefacts of the current production version.
Auditing logging
The claim: ‘all model decisions are logged’ (A.6.2.8 territory). The chase is beautifully concrete: pick a specific decision on a specific date — ‘show me the log entry for the credit decision on applicant file 8841, 14 May’ — and watch. Can they retrieve it at all? Does it capture what their own procedure promises (inputs, model version, output, timestamp)? Then test the edges: retention (‘show me one from 13 months ago’), coverage (‘now the same for the other model’), and access control (‘who can edit these?’). Logging claims fail at the edges, never in the demo.
Key terms: audit evidence, sampling, objective evidence, triangulation
Interactive checkpoint quiz (2 questions) — open this page in a browser to take it.