Worked Example  ·  Trinzo AI Operating System

Post-Market Complaint Handling

Assembling the evidence pack without touching the reportability decision

Status: illustrative

This is a worked example, not an account of a named client engagement. It shows how the AI and LLM Implementation Playbook applies to a representative workflow. All figures are illustrative and are not measured outcomes. Any organisation applying this method must generate its own evidence for its own workflow, sources, configuration and people.

Complaint handling is a good test of the method because it looks like one workflow and is actually two. Most of it is assembling evidence: pulling device history, finding prior similar events, extracting what the reporter actually said. That part is slow, and it is where handlers lose time they would rather spend on judgement. Inside it sits a reportability determination, which is a regulatory decision with patient-safety consequences. The temptation is to treat the whole thing as one candidate for automation. The routing questions in Part 2 exist to stop that.

One currency note, since it changes where the requirement is read from. Since the Quality Management System Regulation took effect on 2 February 2026, complaint record content requirements sit at 820.35(a) together with clause 8.2.2 of ISO 13485 as incorporated, rather than at the former 820.198. The practical addition is that unique device identifiers now have to appear in the complaint record where the device carries one, which is a retrieval and extraction requirement before it is a documentation one.

Section 1The example at a glance

WorkflowComplaint intake, evidence assembly, event coding and regulatory report drafting for a single product family
Front Door dispositionFD2 Controlled Regulated Workflow. Output may be examined by a regulator or inspector.
Regulatory relationshipR1 supports regulated work, after redesign. The reportability determination was R2 as originally proposed.
OverlaysOverlay A (US medical devices) and Overlay B (EU medical devices and IVDs), with the EU horizontal layer.
Amplification intentExtend specialist capacity for judgement by reducing time spent assembling the evidence pack.
Configuration baselineModel, prompt and retrieval index versions are recorded at evaluation. A change to any of them re-opens evaluation at a defined scope.
Pilot boundaryOne product family, one site, complaints excluding death and serious injury, five named handlers.

Section 2Front Door and regulatory routing

The first pass through the seven routing questions returned R2. As originally described, the workflow generated a candidate reportability recommendation which a handler then confirmed. That is a determination the machine proposes and the human ratifies, and the R1/R2 boundary in Part 2 is explicit that editing or accepting an AI-generated conclusion does not convert it into independent human reasoning.

The redesign removed the recommendation entirely. The workflow now assembles and presents evidence, and the handler reaches the reportability decision by applying the criteria to that evidence themselves. This is slower than the original proposal and it is the reason the route is defensible. The second pass returned R1.

Removing the recommendation does not remove all influence, and it would be dishonest to claim otherwise. Evidence the machine fails to surface cannot be weighed, and a prior similar event that was never retrieved biases quietly toward not reportable. That is the reason the Retrieve class is evaluated for silence and false negatives rather than for precision, and why the source gate requires the handler to record what could not be resolved rather than only what was found.

One further consequence followed. Because the machine no longer produces a candidate answer, the handler cannot anchor on one, which is the failure mode the second whitepaper in the series calls the reviewer trap.

Section 3Material-step map

#Material activityTask classAllocationConsequence
1Retrieve device history, prior complaints and applicable proceduresRetrieveCollaborativeTier 2
2Extract reported event detail from unstructured intake recordsExtractCollaborativeTier 2
3Normalise dates, lot identifiers, device references and unitsTransformCollaborativeTier 2
4Propose candidate event and device problem codesClassifyCollaborativeTier 2
5Confirm coding and evidence completenessHuman judgementUnassisted gateTier 2
6Determine reportability against applicable criteriaHuman judgementUnassisted gateTier 3
7Draft complaint narrative and regulatory report contentDraftCollaborativeTier 3
8Approve report content and make regulatory submission decisionHuman judgementUnassisted gateTier 3

Consequence tiers are assessed before crediting any control. Tier 3 means an unchallenged error could affect patient safety, product quality, release or a regulatory position.

Section 4Substantive human gates

GatePositionRequired human contributionArtefact
Source gateBefore retrievalHandler confirms the device history, procedure versions and prior-event set are current and complete, and records any source that could not be resolved.Approved source and version list, unresolved-source note.
Evidence gateAfter extractionHandler accepts, corrects or rejects each extracted event element against the original intake record, item by item.Element-by-element disposition with source reference.
Coding gateBefore determinationHandler confirms or replaces each proposed code and records the rationale for boundary and ambiguous cases.Coding decision and rationale for non-obvious assignments.
Reportability gateBefore report draftingHandler applies the reportability criteria to the assembled evidence and reaches the determination without a machine-proposed conclusion in view.Determination, criteria applied, and rationale.
Report content gateBefore submissionHandler verifies every factual statement, date, identifier and quantity in the drafted report against the verified evidence set.Verified content set and correction record.

A gate is substantive only when passing it requires the person to produce something the machine did not. Approval alone is not a gate.

Section 5Prohibited use

The configured workflow must not:

  • Generate, recommend or rank a reportability conclusion.
  • Close, downgrade or de-duplicate a complaint.
  • Assign a final event code without handler disposition.
  • Draft regulatory report content before the determination has been made.
  • Operate on complaints involving death or serious injury during the pilot.

The exclusion is enforced at intake rather than by handler discretion, because severity is frequently not known when a complaint is first logged. Every complaint enters an unassisted triage step that asks only whether death or serious injury is reported, alleged or cannot be excluded. Only complaints clearing that step are released to the assisted path, and any complaint whose severity is revised upward at any later point is withdrawn from it and re-handled unassisted from the beginning. A complaint that has already passed a gate is not exempt from that.

Section 6Evaluation approach

The evaluation described here produces the evidence that feeds validation for intended use. It does not replace it.

Evaluation followed the task classes rather than the workflow as a whole, because each class fails differently. Retrieve was tested for provenance, currency and corpus silence, including cases where the relevant prior complaint did not exist and the correct output was to say so. Extract was tested against complete-set comparison, with seeded omissions and misattributions. Classify was tested almost entirely on minority and boundary cases, since majority-class drift is the characteristic failure. Draft was tested by claim decomposition, tracing every statement in the narrative back to a verified evidence item.

Elapsed time was measured as well as effort, because the reporting clock is what this work is actually judged against. US Medical Device Reporting timeframes and EU vigilance timelines both run from awareness, not from the point a handler picks the file up, so a design that adds verification without extending elapsed time is the result that matters.

The evaluator was independent of the configuration owner. Expected results came from previously approved complaint records with independent adjudication where the historical record was itself ambiguous.

Section 7Illustrative evaluation results

MeasureUnassisted baselineAI-assisted result
Evaluation set and period24 representative complaints plus 18 constructed boundary and failure casesSame set, eight-week window
Median handler effort per complaint2.6 hours unassisted1.7 hours AI-assisted
Verification Effort RatioNot applicable0.63
Seeded extraction omissions detected at evidence gateNot applicable14 of 14
Seeded misattributions detectedNot applicable9 of 10
Boundary-case coding errors surfaced at coding gateBaseline 4 across set2 corrected, 0 propagated
Unsupported statements reaching report content gateNot applicable3, all corrected
Elapsed time from intake to reportability determinationMedian 6.5 working daysMedian 5.0 working days
Reportability determinations altered by AI-assisted pathTarget: noneNone observed

These figures are constructed to illustrate the shape of a defensible result. They are not measured outcomes and must not be cited as evidence of performance.

Section 8Capability and authorisation

RoleMinimum levelRequired capability
Complaint handlers operating the workflowLevel 2Prepare inputs, operate the configured workflow, recognise boundaries and stop conditions.
Gate operatorsLevel 3Retrieve, Extract and Classify verification competencies within a Tier 3 consequence ceiling.
Quality and release authorityLevel 4Challenge evidence, restrict scope, suspend or retire the workflow.
ManagersRole dutyProtect verification time and monitor whether gates remain substantive under volume.

Section 9Disposition

Progression decision

Acceptable with restrictions, for a bounded operational pilot.

The restrictions are the point. Complaints involving death or serious injury stay entirely on the unassisted path, because the pilot cannot generate enough of those cases to characterise failure in them. The one undetected misattribution is recorded as a known residual, with a monitoring trigger rather than a claim that it has been solved. Scope extension to a second product family requires a delta assessment, not a repeat of the whole evaluation.

The result worth noticing is not the time saving. It is that the reportability determination did not move. A workflow that had cut handler effort by a third while quietly shifting who decides what gets reported would have been a worse outcome than doing nothing, and it would have looked identical on a dashboard.

Records and retention, supplier and platform qualification, audit trail and signature controls, and the linkage to corrective action are all in scope for the method and are handled in the playbook itself. They are deliberately out of scope here so this remains readable in ten minutes.

The method described here is set out in full in the Trinzo AI and LLM Implementation Playbook. The reasoning behind it is developed across the three whitepapers in The AI Capability Series, particularly the third, which deals with verification and gate design.

← All projects