Research
No publications yet, and no performance claims
This page exists to describe how validation would be done and what would have to be true before any claim is made. When there is evidence, it will appear here with its methods. Until then, the honest content of this page is a plan.
Current evidence status
VerityRT has no peer-reviewed validation, no external-institution evaluation, no prospective pilot data and no clinical performance claim. The prototype runs on synthetic fixtures. Any figure shown anywhere on this site is illustrative and labelled as such.
Model development lifecycle
Thirteen steps, in this order, with no shortcuts between them.
- 01Define intended use, users, environment, inputs, outputs and the harms that follow from being wrong.
- 02Create a data specification and a label handbook.
- 03Obtain lawful, representative data with documented rights.
- 04Split by patient, and preferably by institution and time, to prevent leakage.
- 05Create locked development, validation and test sets.
- 06Train a baseline before reaching for a complex architecture.
- 07Evaluate overall and by clinically meaningful subgroup.
- 08Perform external validation at a separate institution.
- 09Run prospective silent mode, influencing no care.
- 10Conduct human-factors and workflow testing.
- 11Release one immutable version.
- 12Monitor drift, overrides, incidents and outcomes.
- 13Retrain offline and repeat change control.
How performance would be reported
- Imaging and contours
- Sensitivity and specificity for the stated finding; Dice, surface Dice, 95th-percentile Hausdorff distance, mean surface distance and volume difference; added path length as a measure of editing effort; the clinically material error rate; and the dose impact of accepted versus corrected contours.
- Planning and dose
- Target and organ metrics defined by the approved protocol; predicted-versus-calculated error; DVH and 3D distribution error; feasibility and deliverability; critical-check sensitivity and false-negative rate; and time saved measured in manual actions, not adjectives.
- Knowledge and language models
- Claim-level citation precision, source faithfulness, contradiction detection, unsupported-claim rate, correct abstention rate, evidence freshness, numeric transcription error rate and resistance to adversarial prompt injection.
- Operations
- Planning turnaround, review time, rework rate, near misses caught, and alert acceptance, override and fatigue rates.
- Never a single number
- Performance is reported by disease site, scanner, machine, technique and relevant subgroup. One global score hides exactly the failure a physicist needs to find.
Release gates
A model or rule package cannot be released unless every one of these is true.
- Intended use and contraindications are documented.
- Data provenance and legal rights are documented.
- Test sets are locked and independent.
- Required performance thresholds are met on every critical stratum.
- Known failure modes and mitigations are documented.
- The model card and user-facing limitations are complete.
- Clinical, physics, security, privacy, engineering and quality owners have each signed.
- Rollback and downtime procedures have been tested.
- Training and site commissioning material is complete.
Data sources under consideration
Public research datasets accelerate development. None becomes clinically authoritative because it is popular.
- Public cancer imaging collections
- Useful for prototyping. Collection-specific licence, cohort composition, label quality and data provenance must each be reviewed before use.
- Benchmark planning datasets
- Useful for method comparison. A benchmark built on a few hundred cases of one disease site does not validate general clinical planning.
- Institution-approved retrospective data
- The most useful source, and the one requiring approvals, contracts, de-identification, governance and leakage prevention.
- Prospective silent-mode data
- Real-world performance and failure analysis, with no influence on care until release criteria are met.
- Synthetic cases and phantoms
- What this prototype runs on. Good for interface, integration and boundary testing. Never a substitute for clinical evaluation.