← All tutorials

The Referral That Wasn't

An audit of CARDEA, a cardiovascular risk score used to triage GP referrals, and the subgroup performance gaps it produces.

Code··········

0 / 10 questions open

A 12-month cardiovascular event risk score that auto-triages GP referrals into a rapid-access chest pain clinic.

You are the independent assurance team commissioned by the Mersey & Dee Integrated Care Board. Nine months ago the ICB switched on CARDEA, a risk score built by Meridian Health Analytics, to triage referrals into its rapid-access chest pain clinic. Patients scoring above 0.30 are fast-tracked; everyone else joins the routine queue.

Last month a consultant cardiologist wrote to the ICB to say that, anecdotally, the patients arriving late and sick are not a random sample of the population. The ICB has given you three things: an inbox, a channel, and a document room, and two hours before the contract renewal meeting.

There are ten questions across three stages. Stage A, an inbox, is open now. Stage B and Stage C stay locked until every question in the stage before them is solved. Each question is answered from the evidence in that stage.

What you should be able to do afterwards
  • Describe the sources of bias: data (sampling, measurement), algorithmic (model, optimisation) and societal (historical, structural), and identify each in a real system.
  • Calculate Demographic Parity Difference, Equal Opportunity Distance and Equalised Odds Difference from a subgroup performance table, including the worst-case and mean-pairwise strategies for more than two groups.
  • Distinguish internal from external validation, and explain why performance falls on data the model was not developed on.
  • Critically evaluate pre-processing (SMOTE, ADASYN, Near-Miss, Tomek links), in-processing (regression adjustment, adversarial training) and post-processing (threshold adjustment, Platt scaling) mitigations.
  • Explain why calibration within groups and equalised odds cannot generally be satisfied together, and defend a choice.

Stage A: Inbox

Stage A is an email inbox from the ICB assurance team. It contains the correspondence that led to this review: a cardiologist's escalation, Meridian's replies, the datasheet for CARDEA's development cohort, and the internal validation summary. Using this information, you should be able to identify the sampling bias in CARDEA's development cohort, work out what its validation figure actually measures, and explain why the monitoring that replaced local revalidation could not have detected the problem.

1

Question 1 of 10

≈ 12 min

Who was this model built for?

Sampling bias is the over- or under-representation of a group in a dataset. Before you test Meridian's 'CARDEA is blind to race and income' line, measure how well the development cohort represents the population CARDEA now scores.

Open the BI team's population table and the datasheet.

Objective

Quantify the under-representation of the most-deprived groups in the development cohort.

Using the population table, what is the percentage-point gap in the most-deprived three IMD deciles (deciles 1–3) between the development cohort and the deployment population? Enter the size of the gap, in percentage points.

percentage points

1 / 10

2

Question 2 of 10

≈ 12 min

Internal marks, external stakes

Hargreaves' reply leans entirely on one number: AUC 0.81. Before you let that number reassure anyone, work out exactly what it was measured on.

Objective

Determine what Meridian's headline AUC of 0.81 actually tells you, and what it doesn't.

Which statement correctly describes how CARDEA's AUC of 0.81 was obtained, and what that does and doesn't establish?

2 / 10

3

Question 3 of 10

≈ 10 min

The corner that got cut

One step in the original plan would have caught the mismatch you just quantified before a single patient was scored. It didn't happen. The delivery team's own forwarded memo says why.

Objective

Explain why the monitoring that replaced the cut step could never have caught the mismatch you just quantified.

Why could the ICB's chosen alternative; 'monitor performance post-deployment via the standard quarterly dashboard'; never have caught the 28-point gap you quantified in Question 1?

3 / 10

Stage locked

Stage B: Channel

An export of Meridian's internal channels: the fairness evaluation a data scientist ran fourteen months ago, the metrics the team used, and the decision to remove the findings from the board deck.

Unlocks after all questions in Stage A: Inbox

Stage locked

Stage C: Archive

Meridian's assurance archive: the feature list, the outcome label definition, the equality impact assessment, and four mitigation proposals.

Unlocks after all questions in Stage B: Channel