← Back to writing
Research

Introducing ComplianceBench.

An open benchmark for AI on financial compliance casework, built on a fully published institution and released with its complete answer keys.

Brayns today releases ComplianceBench, an open benchmark measuring how AI models perform the core casework of financial compliance: business and consumer onboarding, screening decisions, transaction monitoring investigations, and regulatory gap analysis. The benchmark is constructed on a complete fictional institution whose entire regulatory framework is published, which makes every answer key inspectable against the same documents a model receives. Every case was adjudicated by Brayns’ compliance panel, practitioners from real institutions, and their rulings are the ground truth. This release contains the instrument and no scores. Benchmark results for the leading foundation models follow in the next publication, and Brayns’ own model is evaluated on the identical protocol after them, in public.

Why this didn’t exist.

There are two reasons, and neither is technical.

The first is structural. A legal question can be graded against public law. A compliance decision cannot, because the right answer depends on the law plus the institution’s own confidential framework: its risk assessment, its risk appetite, its rules and thresholds. The same customer with the same documents can be correctly approved at one institution and correctly declined at another. Both decisions are compliant. So every real compliance case has a confidential answer key, and a test you cannot publish the answers to is not much of a test.

The second reason is expertise. Ground truth in compliance is not derivable from documents alone: it requires people who can read a case and state, with authority, what a real institution does with it. That expertise concentrates inside regulated institutions, not inside AI labs, which is why AI teams have not built this and could not have. Brayns’ compliance panel consists of practitioners who have run this work at real institutions, approved and declined customers, filed the reports, answered to the supervisor. Their rulings define every answer key in the benchmark.

The hard part was never the software. The hard part is knowing the right answer.
PUBLIC

Legal benchmark

Question graded against statutes, cases, regulation.

The answer key is public law. Anyone can check it.

CONFIDENTIAL

Compliance case

Question graded against the law plus the institution’s risk assessment, appetite, rules and thresholds.

The answer key is confidential. Nobody could publish it, so nobody could publish the exam.

THE MECHANISM

Mynta AB, a full institution with nothing to hide

The institution is fictional, so its confidential layer is public. The layer is public, so every case ships with its complete answer key. That is the whole trick, and it only had to be built once.

An institution with nothing to hide.

To remove the confidentiality without losing the realism, Brayns constructed Mynta AB, a fictional Swedish e-money institution, at full regulatory depth: a business-wide risk assessment, policies, work procedures for onboarding, screening, monitoring and reporting, rules, registers and templates. Around twenty documents, written to the standard a Swedish institution writes them. The framework was reviewed by our compliance panel the way a real framework is reviewed, and their corrections are visible in the published documents, version histories included.

Mynta does for compliance what public law does for legal benchmarks. The institution is fictional, so its confidential layer can be public. The layer is public, so every case built on it ships with a complete answer key that anyone can check, line by line.

The pack~20 documents, the whole rulebook
The caseapplication, registers, screening, messages
The modelreads raw files, answers in one format
The keyexpected outcome and findings, public
The gradebinary criteria, seven dimensions

One case, end to end. Every box is in the repository, including the answer key. Nothing to trust, everything to check.

The institution is fictional, so the answer key can be public. That is the whole trick, and it only had to be built once.

Adjudication is where the answer keys come from.

Adjudication is not a review step at the end; it is where the answer keys come from. The panel ruled on every case, question by question: whether the ownership calculation follows Swedish practice, whether the escalation path is the correct one, whether a statutory duty triggers immediately or at resolution. Where a case overstated a rule, the ruling cut it back. Where the framework itself had a gap, the ruling closed it, and several of those rulings are now written into the institution’s procedures with the legal reference attached. The adjudicating professionals are named in the repository credits, and the full adjudication record behind every answer key is archived and produced if a ruling is challenged.

What version 1 contains.

PartContent
The institutionThe complete Mynta framework: risk assessment, policies, work procedures, rules, registers and templates.
Module 1: caseworkCases in business onboarding, consumer onboarding, screening and transaction monitoring, each with full case documents, an expected answer and a grading rubric.
Module 2: gap analysisRegulatory change cases: a new obligation meets the framework, and the model must find what is covered, what is missing and what follows from that.
Answer keysEvery case ships with its expected outcome, the findings a correct answer must contain, and the rubric it is graded on. Nothing held back.

Version 1 is deliberately deep rather than wide. The pilot cases underwent multiple adjudication rounds, including a full-panel session, and the doctrine they established now governs production of the volume set under the same adjudication protocol.

Grading is binary, and every line is checkable.

Brayns’ research team developed a bespoke rubric for every case: between ten and fifty binary criteria, each independently checkable against the case documents. Criteria are unweighted by design; the importance of a finding is expressed as more criteria, never as heavier ones, which keeps every pass and fail auditable. Here is what a criterion means in practice, from a live case:

Rasmus Kasknatural person
Other holdersbelow threshold
24.0% direct | 78.0% of Osta Holding
Osta Holding OÜEstonia · holds 32.0%
32.0% of the applicant
Halvix Teknik ABthe applicant · declares: no beneficial owner
24.0% + (32.0% × 78.0%) = 24.0% + 24.96% = 48.96% → beneficial owner. The declaration is false.

A live case from version 1. The customer’s declaration and the public register both say no owner exists. The documents, read together, say otherwise. A model that repeats the declaration approves a false filing; a model that does the arithmetic escalates and reports the register error. One of the graded criteria is exactly this calculation.

FROM THE GRADING RUBRIC · M1-KYB-001 · PASS OR FAIL
C02Computes the indirect holding correctly (32% × 78% = 24.96%) and the total 24% + 24.96% = 48.96%.
C07States that the wrong register entry is notified to Bolagsverket immediately, without waiting for the customer, and records it in the memo.
Decision accuracyScreening precision and coverageEvidence groundingConsistencyPolicy adherenceEscalation at uncertaintyAuditable reasoning

The measurement frame. Seven dimensions, no curves yet: scores come in the next post.

Every criterion is tagged to one of seven dimensions: decision accuracy, screening precision and coverage, evidence grounding, consistency, policy adherence, escalation at uncertainty, and auditable reasoning. Two are measured in ways worth stating precisely. Consistency is evaluated on twin cases, identical facts presented differently, where the decisions must match. Escalation at uncertainty scores “this requires a human decision” as correct wherever the framework prescribes it, because in regulated work a confidently wrong model is a larger operational risk than a model that stops.

Alongside the rubric runs an automatic scoring layer with four measures: outcome match against the expected disposition, coverage of the required findings, detection of fabricated findings, and citation validity, which benchmarks the verifiability of an answer as the share of factual claims citing a document that actually supports them. The scoring reference implementation ships in the repository.

Scores come later, on purpose.

This publication releases the instrument: the institution, the cases, the answer keys, the rubrics and the validation tooling. It releases no scores, including ours. The next publication benchmarks the leading foundation models and reports their results in full. After that, Brayns’ own model is evaluated on the identical protocol, in public. The sequencing is deliberate. A benchmark released together with its maker’s winning score is marketing material; released before any score exists, it is an instrument, and an instrument is what this field lacks.

Released before any score exists, it is an instrument, and an instrument is what this field lacks.

Built to be falsified.

Automated validation runs on every change to the benchmark: no case document may disclose its conclusion, every constructed gap must verifiably exist in the materials, every cross-reference must resolve, and the arithmetic in each case must agree with its documents. Adjudication records are archived per case. External red-teaming is part of the production standard for the volume set.

The benchmark is built to be falsified. Every answer key is published precisely so that practitioners can contest it; where a challenge shows a key wrong, the correction ships as a versioned change and the version history remains public. That is the maintenance model, and it is the same standard we grade the models on: show the evidence, cite the source, accept correction.

Read the cases, the answer keys and the validators, and run your own model: the ComplianceBench repository →

Benchmark results for the leading foundation models follow in the next publication.

See Brayns in your environment. Book a demo →