Research

Benchmarks, methodology, and what we're building


Expert-authored evaluation, transparent methodology, and announcements — the public record of how Aiscéal measures and is measured.

AISCÉAL EVALS
Model leaderboard

Expert-authored benchmarks, independently governed. Every score carries a confidence interval and its sample size — no bare rankings.

View the leaderboard →
METHODOLOGY
How we measure

Signed-weight rubrics, adversarial review, named human verification, and confidence intervals — the standard behind every number we publish.

Read the methodology →

Documentation

  • How expert assessments work
    Real work, adversarially reviewed, human-verified — and treated as portfolio evidence, not pass/fail scores.
    Read →
  • How our data quality works
    Blind double-annotation, agreement thresholds, gold-seeded sampling, and human adjudication that feeds back into the standard.
    Read →
  • Provenance and EU AI Act conformity
    Every deliverable carries who did what, under which clearance, on sovereign infrastructure — the paperwork regulated deployments need.
    Read →

News