Kaia Benchmarks

The serving system, measured, on the record.

We publish the current accuracy of the serving model on frozen, versioned evaluation sets, release over release — the same number a GC's diligence team, a payer's compliance office, or a regulator should be able to demand. Vol. 1 is live below; monthly is the cadence we intend. No cherry-picked demos, no one-time lab results.

August 2026

Vol. 1

Legal E-Discovery

95.7%

Privilege-detection accuracy — measured internally on 810 documents, published as-is.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: kaia-native-legal-8b-v2 — the model serving production traffic as of July 31, 2026

Evaluation set: 810 held-out gold-standard records across 94 distinct document bodies — frozen, never used in training

Method: The exact production prompt and production response parser, temperature 0, run against the full set — not a sample

Privilege detection is the axis that matters most in regulated e-discovery: a missed privileged document is the failure mode that ends careers and waives rights. That is why it is the axis we publish first.

The full results — including the axes we're not proud of yet

Transparency means the whole table, not the best row. Per-record results on the same 810-document frozen set:

Axis (per record, n=810)CorrectRate
Privilege(the regulated claim)77595.7%
Sensitivity74291.6%
Relevance54667.4%
All three axes exact51663.7%
Unparseable responses192.3%

What we say about the weaker rows, on the record

Relevance is flat at 67.4%. This is a known, tracked limitation. Relevance scoring is a separate training-corpus workstream now underway; it is not hidden, and we will publish its movement in this series — up or down.

2.3% of responses fail parsing. A parse failure is fail-closed by design: the document routes to human review. The system never silently guesses. We count these against ourselves here.

How to read this number honestly

This is an internal measurement. The evaluation set is frozen and held out from training, but the measurement is run by us. We publish it because buyers deserve the number and the method — and we invite design partners to run their own documents through the system and hold us to it.

The number moves. Every client correction routes through the Intelligence Engine and improves the serving model. As engagements go live, this series will publish the movement — this is the self-improving loop, measured in public.

A note on metric history: our July acceptance bar was privilege recall on the privileged class (94.74% against an 84.7% bar, up from 74.9% for the v1 model). This publication reports per-axis accuracy. Both metrics stand on the record; going forward this series standardizes on per-axis accuracy over the frozen set.

Methodology appendix

Set construction

810 records / 94 distinct bodies, gold-standard labeled, held out from all training runs. Gate-verified (#425).

Serving path

Bedrock Custom Model Import, production registry entry — the measurement exercised the same prompt, parser, and temperature that client traffic receives.

Cold start

The run exercised real cold-start behavior (retry-until-warm, then ~1.4–1.9s per call) — production conditions, not lab conditions.

Per-body consistency

41 of 94 distinct document bodies scored perfectly across all their records; macro-average exact-three-axes rate 62.5%.

Next in this series

  • September 2026:movement on all axes; first read on the relevance training-corpus workstream.
  • Now published below:Accounts Payable and Healthcare Claims on the same frozen-set standard — and, once client engagements are running, measured per-engagement improvement from the correction loop.

Questions from diligence teams: benchmarks@kaiaai.ai. We answer with the raw data.

Accounts Payable

92.75%

Exact extraction accuracy — the complete extracted record and routing decision, matched exactly, measured internally on 1,200 held-out invoices.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: kaia-native-ap-8b — the model serving the accounts-payable lane in production since August 5, 2026

Evaluation set: 1,200 held-out gold-standard records — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set — not a sample. An exact score means everything matched: a partially right invoice counts as wrong.

The residual errors run conservative by design. On this set, sanctions-screening recall was 100% and straight-through-payment precision was 100% — no invoice that should have been held was ever paid. The failure mode we refuse is the expensive one.

Healthcare Claims

99.66%

Fraud-signal recall — 292 of 293 flagged-class claims caught on the held-out set, at 100% precision, measured internally.

Status: MEASURED (internal) — internal measurement on a frozen evaluation set, published as-is

Model: kaia-native-claims-8b — the model serving the claims lane in production

Evaluation set: 860 held-out gold-standard records — frozen, never used in training

Method: The exact production adapter path and production response parser, run against the full set. The acceptance bar was pre-registered before the measurement ran — the number had to clear a written-down falsifier, not a retrospective judgment.

The whole-set number, same run: overall exact accuracy across all decision classes on the same 860 records was 96.05%. Both numbers stand on the record together.

Oil & Gas — published to a different standard

There is no model-accuracy score in this section — deliberately. The reserves lane is a deterministic engineering engine, not a trained classifier. Every figure it produces carries the calculation that produced it, and every result is reproducible from stored evidence — the inputs, the method, and the number, together.

For reserves estimation and disclosure work, the standard is the benchmark: a figure either carries its audit trail or it does not ship. When this series publishes for the Oil & Gas lane, it will publish engineering-verification results — never a model score dressed as one.

in build Electrical design (construction): the same engineering-grade standard, extended to construction electrical design. Produced to date: a design-basis report, load calculations, and a bill of quantities. No benchmark until the lane earns one.

Run your own documents. Hold us to it.

Design partners get access to the serving model and frozen evaluation methodology. We publish the results — yours and ours.