Aiscéal Evals

Model leaderboard

Expert-authored benchmarks, independently governed. Every score carries a confidence interval and its sample size; sub-scores are always published alongside the headline. Methodology and the full protocol are published with each release.

Clinical
ResultsScore vs price
✓ EU-deployable only
Domain Clinical · Version 0.1.0
Clinical consultation role-play v1 is not yet scored

No model has been run against this benchmark under its published protocol. When one has, the table below will carry every result — including the ones the provider would rather it did not.

Expert-authored
Tasks and signed-weight rubrics are written by credentialed practitioners in the domain, not scraped from exam banks.
Frozen protocol
Temperature, seed, token limits and the judge panel are pinned per version and published verbatim with the results.
Every number carries its uncertainty
A confidence interval and a sample size on every row. Results below the minimum sample are labelled preliminary rather than hidden.
Independently judged
A panel of judges from different providers, and no model is ever graded by a judge from its own provider.
How this leaderboard is governed
No model of our own
Aiscéal builds no models, so there is nothing here we could be measuring in our own favour.
No money from the measured
We take no investment from evaluated providers.
No self-grading
No model is graded by a judge from its own provider.
Nothing withheld
Providers see their results before first publication and may respond. No result is edited, withheld or withdrawn.

All results are produced under a frozen, published protocol with retrieval disabled.