A full batch delivery, worked end to end.
Synthetic vignettes, scripted outputs, simulated raters — shows the shape of a delivery, not clinical evidence.
Download the sample report (PDF)Not an assertion — a published formula, versioned per protocol, with a sample delivery you can inspect before you ever get on a call.
reliability_score = 100 × ( 0.40 × gold_pass_rate
+ 0.35 × disposition_agreement
+ 0.25 × (1 − critical_miss_rate) )Weights are protocol-versioned per task family — this composite is uc1-rubric-v0 under protocol v0.1. Every delivery states the protocol version and ships component metrics, not just the composite.
Synthetic vignettes, scripted outputs, simulated raters — shows the shape of a delivery, not clinical evidence.
Download the sample report (PDF)Disposition, red-flag coverage, and dangerous omission/commission — the scales raters actually use.
Download the rubric excerpt (PDF)Validating against published JSON Schemas — vignette and rating.
Gold IDs and provenance, so every score card traces back to its source cases.
gold_pass_rate, disposition_agreement, and critical_miss_rate shipped alongside the composite — not just the single number.
Schemas: uc1-vignette.schema.json and uc1-rating.schema.json. Publishing the schema files themselves is a planned follow-up.
Bring your task family and we'll show you exactly how the score would be computed — protocol, gold set, and weights.