Independent clinical AI evaluation

Find where your clinical AI fails before your users do.

Physician-led evaluation of clinical AI for unsafe reasoning, missing information, unsupported certainty and convincing-but-wrong answers. Every judgement is tied to the exact test case, output, rubric criterion, evidence and severity.

  • Physician-led
  • Synthetic data only
  • Reproducible evidence pack
  • Independent point-in-time assessment
Evaluation trace Critical
Case DF-SURG-004 v1.0
Failure Urgent deterioration missed
Output SHA-256 linked
Rubric DF-CLIN-EVAL v1.0
Evidence Source register attached
A score is not enough. The buyer receives the evidence needed to understand, reproduce and challenge the judgement.

The practical problem

A fluent answer can still be clinically unsafe.

Medical AI can fail without sounding confused. It may omit the one question that changes management, dismiss a time-critical pattern, invent a source, or present an uncertain conclusion as settled fact. DashFin is designed to expose those failures in a form a product team can act on.

Launch services

Defined scope. Tangible evidence.

Introductory pricing is intended for early product teams and bounded pilots. Scope is confirmed in writing before work begins.

Broader evidence

Synthetic Clinical Reasoning Evaluation

£750 introductory scope

Fifty cases with dimension-level scoring, systematic error analysis and prioritised remediation themes.

  • 50 case and output records
  • Explicit, versioned rubric
  • Failure pattern analysis
  • Evidence and source register
  • Retest-ready finding IDs
Discuss suitability

Custom programme

Synthetic Benchmark or Red Team

From £1,500

Custom task design for a product, model version, user group, failure hypothesis or investment/pilot review.

  • Adversarial and edge-case design
  • Custom acceptance criteria
  • Comparative model evaluation
  • Reproducible evidence pack
  • Optional remediation retest
Define the programme

The £295 launch scope covers one English-language product workflow, one declared model or configuration, one agreed clinical domain and one point-in-time assessment. It is evaluation evidence, not certification. No identifiable patient data is required or accepted.

Experimental agent evaluation

Test what your clinical agent does—not just what it says.

DashFin’s Clinical Agent Action Envelope makes tool use, state changes, approvals, stop conditions and human hand-offs explicit before a synthetic test begins. Paired control and deliberately unsafe traces show whether the workflow preserves those boundaries under failure.

DF-CAE-v0.1 Fail closed
Declares
Tools, effects and authority
Captures
Every action and state transition
Challenges
Identity, injection, outage and replay
Preserves
Human authority and immutable retest links

Synthetic-only reference workflow. No external action is executed.

Evaluation method

From model output to defensible finding.

The method is designed to separate observation from interpretation and to preserve what was tested.

  1. 01

    Scope

    Define the workflow, intended user, clinical domain, model or configuration, access method and decision the evaluation must inform.

  2. 02

    Test design

    Create synthetic adversarial and edge cases that target relevant risks, ambiguity, missing information and likely failure modes.

  3. 03

    Run and capture

    Record each prompt, system context, model identity, settings, exact output, timestamp and content digest.

  4. 04

    Score

    Apply an explicit, versioned rubric covering safety, correctness, urgency, completeness, uncertainty and evidence integrity.

  5. 05

    Adjudicate

    Review material failures individually, assign severity and document the rationale. Critical findings cannot be averaged away.

  6. 06

    Report

    Link each finding to the exact case, output, criterion, evidence and evaluator judgement, with remediation and retest status.

What is tested

Clinical quality is more than factual recall.

Factual accuracy

Whether material statements and conclusions are supportable.

Clinical reasoning

Whether the reasoning connects the available facts to a defensible conclusion.

Missing information

Whether the system requests facts needed before reaching a conclusion.

Differential diagnosis

Whether dangerous alternatives are prioritised rather than merely listed.

Urgency and escalation

Whether time-critical patterns are recognised and handled safely.

Investigation and treatment

Whether suggested actions are proportionate, relevant and free from avoidable harm.

Uncertainty handling

Whether confidence matches the evidence, ambiguity and intended system role.

Contradictions

Whether conclusions conflict with the case, earlier reasoning or stated limits.

Claims and citations

Whether stated facts and references exist and actually support the answer.

Temporal reasoning

Whether sequence, duration, deterioration and changing evidence are interpreted correctly.

Evidence, not a black-box score

Every important judgement should be challengeable.

DashFin does not present an unexplained pass rate as assurance. Each material finding identifies what the model said, what criterion was applied, what evidence informed the assessment, how serious the failure is and why the conclusion was reached.

Severe findings remain visible as individual records. They are not concealed by an otherwise high average score.

Open the full illustrative finding → Follow the evaluation walkthrough →
Illustrative finding DF-FAIL-007
Test condition
Post-operative deterioration with incomplete observations
Exact output
Captured and content-digest linked
Criterion
Urgency, escalation and missing information
Severity
High
Rationale
Output reassures without requesting the information needed to exclude a time-critical complication.
Evidence / reference
Source register entry DF-SRC-014
Status
Open · remediation and retest required

Illustrative structure only. This is not a patient case or clinical protocol.

Clinical evaluator
MBChB · MRCS (Edin)

Evaluator

Medical reasoning, evidence review and governed evaluation.

DashFin's clinical AI evaluation work is led by a UK-trained physician and former General Surgery Specialty Training Registrar with NHS surgical, academic, audit, research and teaching experience.

The evaluator holds an MBChB from the University of Dundee and MRCS from the Royal College of Surgeons of Edinburgh. The background includes peer-reviewed surgical research, medical teaching, undergraduate OSCE examining, clinical audit and the CFA Institute Investment Foundations Certificate.

The evaluator is also founder and project lead of FarDb, an open-source programme focused on evidence provenance, authority, temporal validity, explicit lifecycle decisions and auditable decision histories. That work provides the traceability discipline used in DashFin evaluations.

The service is independent evaluation work. It does not represent current clinical practice, provide care to patients or substitute for a product's appointed clinical-safety, regulatory or legal professionals.

Suitable buyers

Built for teams that need evidence before exposure.

Medical AI startups

Before pilot conversations, demonstrations or wider clinician testing.

Clinical documentation tools

To examine omissions, invented details, unsafe summaries and escalation failures.

Medical education products

To test reasoning quality, answer keys, explanations and learner safety.

AI labs and model teams

For physician-authored benchmarks, response evaluation and structured training feedback.

Investors and advisers

For a bounded independent view of product risk before deeper diligence.

Clinical product leaders

To turn suspected failure modes into a documented, retestable evidence set.

Clear boundaries

What a DashFin report is—and is not.

It is

  • An independent evaluation of specified outputs
  • A reproducible record of cases, criteria and findings
  • Synthetic data only
  • Evidence that can inform further testing and professional review
  • Private-system access only after written scope, confidentiality and data-handling terms

It is not

  • Clinical care, emergency support or patient-specific medical advice
  • Medical-device certification or conformity assessment
  • Legal or regulatory compliance certification
  • Clinical-safety sign-off, including DCB0129 or DCB0160 approval
  • A guarantee of safety, compliance or future model behaviour

The launch services use synthetic data only. DashFin does not accept identifiable, pseudonymised or claimed-redacted patient data. Private-system access is read-only and permitted only where no patient data or production credentials are exposed, after written scope, confidentiality, data-handling, contracting and verified insurance gates have passed. The work is not certification, conformity assessment, regulatory approval or deployment authority.

Start with one bounded question

What could this system get dangerously wrong?

Send the product type, intended user, clinical domain, model or workflow, development stage and desired timeline. Confirm that the proposed work can use synthetic data only; do not include real-person information, confidential material or credentials.

Email DashFin

mmohamed@dashfin.org
Remote evaluation · United Kingdom

Frequently asked questions

Practical details

What products can be tested?

Bounded workflows from medical-AI startups, clinical documentation and ambient-scribing tools, education or reasoning products, triage and decision-support systems, and general AI models with health-domain behaviour.

Do you need patient data?

No. Every launch service uses synthetic cases and captured outputs confirmed to contain no personal data. DashFin does not accept identifiable, pseudonymised, claimed-redacted or otherwise real-person-derived patient data.

Is this regulatory certification?

No. DashFin provides independent evaluation evidence, not certification, conformity assessment, regulatory approval, legal compliance certification or formal clinical-safety sign-off.

What does the report contain?

A concise executive summary, declared scope and run conditions, case and output register, rubric, severity-ranked findings, evidence and rationale, remediation priorities and retest-ready finding IDs.

Can you evaluate a private model?

Yes, only where access is read-only and exposes neither patient data nor production credentials. Written scope, confidentiality, data handling, contracting and verified insurance gates must all pass before DashFin accesses a non-public system.

What happens if a critical failure is found?

It is surfaced separately rather than diluted by an average score, documented with its evidence and rationale, and communicated promptly through the agreed contact route. A later retest can reference the original finding without overwriting it.