DRAFT FOR REVIEW BY APPOINTED CLINICAL SAFETY OFFICER — MUST BE REVIEWED, AMENDED AND SIGNED BY A GMC/NMC/HCPC-REGISTERED CLINICIAN WITH FORMAL DCB0129 CLINICAL SAFETY OFFICER TRAINING BEFORE USE.
CSO05 — Clinical Safety Case Report (CSCR): MAGIC-NHS
Document reference: BRITI-CSCR-MAGIC-001 Version: 0.9 (Draft for CSO review) Manufacturer: BritiAI Limited Subcontractor (in-scope): Votee AI (MAGIC training pipeline) Applicable standard: DCB0129:2018 Status: Draft pending CSO sign-off
1. Executive Summary
MAGIC-NHS is a sovereign clinical Large Language Model (LLM) fine-tuning pipeline. It enables a deploying NHS organisation to adapt a base clinical LLM on its own data, entirely within the trust data boundary, with no data egress to BritiAI or any subcontractor. The output is a trust-owned model artefact that may be used to power downstream BritiAI solutions (e.g. Scribe-On-Site summarisation) or other trust-approved Health IT systems.
MAGIC-NHS is infrastructure, not a clinical-facing tool. It does not itself produce clinical output. However, the model artefacts it produces do power clinical-facing tools, and therefore MAGIC-NHS is treated as a safety-relevant pipeline under DCB0129. This CSCR addresses the safety risks arising from the training process itself and the assurance properties of resulting model artefacts.
2. Solution Description
- Data preparation toolkit. Tools for de-identification, dataset curation, splitting, leakage checks, and provenance tagging — executed entirely within the trust boundary.
- Training orchestration. Reproducible training jobs with versioned configurations, deterministic random seeds where practicable, and logging of every training input, hyperparameter and checkpoint.
- Evaluation harness. A configurable suite of evaluations covering general clinical knowledge, instruction following, refusal behaviour on out-of-scope queries, fabrication probes, demographic bias probes, and any trust-specific evaluations.
- Model artefact governance. Each candidate artefact carries a structured Model Card including training corpus summary, evaluation results, known limitations and intended-use envelope.
- Release gating. A model artefact may only be promoted to a downstream BritiAI solution after CSO and trust CSO joint review of the Model Card and evaluation results.
3. Intended Use
MAGIC-NHS is intended for use by trust-authorised data engineering, ML engineering and clinical informatics staff to produce candidate model artefacts for evaluation and, subject to safety gating, downstream deployment.
It is not intended to:
- Produce clinical output directly.
- Be operated by clinical end-users.
- Promote a model artefact to clinical use without CSO and trust CSO joint review.
4. Intended Users
- Trust data engineers and ML engineers.
- Trust clinical informatics staff acting as data stewards.
- BritiAI engineers under appropriate access controls when supporting the trust.
End clinical users do not interact with MAGIC-NHS.
5. Environment of Use
Trust-managed compute, including approved private cloud tenancy where this is the trust’s standard pattern. UK data residency confirmed for all standard configurations. No training data leaves the trust boundary.
6. Clinical Claims
BritiAI claims that MAGIC-NHS:
- Enables fine-tuning on trust data without data egress.
- Produces reproducible, fully audited training runs.
- Produces model artefacts accompanied by a structured Model Card and evaluation results.
Explicit non-claims
MAGIC-NHS does not claim:
- That a fine-tuned model is safe for clinical use absent CSO and trust CSO joint review.
- That fine-tuning eliminates fabrication, bias or drift.
- That training data automatically meets de-identification standards without trust process.
- To be a medical device.
7. Risk Envelope
Principal hazard categories (CSO07, MAG-01 to MAG-10):
- Training data leakage — inadvertent egress of identifiable data via logs, metrics or telemetry.
- Insufficient de-identification — identifiable data reaching training corpus contrary to trust policy.
- Behavioural drift — fine-tuning that degrades safety properties of the base model (e.g. weakening refusal behaviour, introducing fabrication).
- Evaluation gap — insufficient coverage in the evaluation suite to detect a safety-relevant regression.
- Inappropriate promotion — a candidate artefact promoted to clinical use without safety gating.
- Reproducibility loss — inability to reproduce a training run, hindering incident investigation.
- Bias amplification — fine-tuning that amplifies demographic or other clinically relevant biases.
- Supply-chain compromise — compromise of a base model or training dependency.
- Configuration sprawl — drift between documented and actual training configurations.
- Model card inaccuracy — Model Card misrepresenting training data, evaluation results or limitations.
8. Risk Control Strategy
- Elimination by design. No egress of training data. No clinical output produced by MAGIC-NHS itself.
- Reduction by design. Mandatory de-identification steps; automated leakage checks; deterministic configuration management; mandatory evaluation suite covering refusal, fabrication and bias; cryptographic hashing of training corpora and checkpoints.
- Protective measures. Two-person rule for model promotion to clinical use; mandatory Model Card review by CSO and trust CSO; signed promotion record.
- Information for safety. Training and documentation for trust ML and informatics staff; clear Model Card template.
9. Residual Risks Summary
Subject to CSO judgement:
- Behavioural drift and evaluation gap are the highest-priority residual risks. They are mitigated by the evaluation suite and joint promotion review but cannot be wholly eliminated; ALARP justification is sought.
- Inappropriate promotion is mitigated by the two-person rule and CSO sign-off; residual risk amber.
- Bias amplification is mitigated by demographic bias probes and ongoing post-deployment monitoring of downstream solutions.
10. Assumptions and Dependencies
- The trust operates an information governance regime adequate for handling the training data.
- Base models are obtained from vetted sources with documented provenance.
- Downstream solutions consuming model artefacts implement their own safety controls (Scribe-On-Site CSCR etc.).
11. Clinical Safety Verification
- Independent review of de-identification tooling against a curated red-team corpus.
- Replayability tests demonstrating that documented configurations reproduce within tolerance.
- Evaluation suite validation against held-out reference scenarios.
- Penetration testing of the training infrastructure.
12. Post-Deployment Monitoring
- Telemetry on training jobs (no training data content) including configuration hashes, evaluation outcomes and promotion events.
- Periodic re-evaluation of deployed artefacts against an updated evaluation suite.
- Mandatory re-evaluation following any base model update.
A safety review per deploying trust is conducted at three months post go-live and annually thereafter.
13. Change Control
Material changes triggering re-assessment include: change to base model; change to de-identification tooling; change to evaluation suite scope or thresholds; change to promotion gating workflow.
14. Statement of Conformance
Subject to CSO review and sign-off, BritiAI confirms that the clinical risk management activities undertaken for MAGIC-NHS have been performed in accordance with DCB0129:2018.
Linked artefacts: CSO01, CSO07 (MAG-01 to MAG-10), CSO08, BRITI-CRMP-MAGIC-001, BRITI-MODELCARD-TEMPLATE-001.
CSO name: _________________________ Registration body and number: _________________________ Signature: _________________________ Date: _________________________
