Research

ComplianceBench.

Compliance could never be benchmarked, because the answer key was always confidential. ComplianceBench removes the confidentiality without removing the realism.

ComplianceBench is an open benchmark that measures how AI models perform the core casework of financial compliance: business and consumer onboarding, screening decisions, transaction monitoring investigations, and regulatory gap analysis. It is built on a complete fictional institution whose entire regulatory framework is published, so every answer key can be inspected against the same documents the model was given.

Every case was adjudicated by Brayns’ compliance panel, practitioners who have run this work inside real institutions, and their rulings are the ground truth. This release contains the instrument and no scores.

The frameworkthe institution's published rulebook
The casethe file as an analyst would receive it
The modelreads the raw files, answers in one format
The answer keyexpected outcome and required findings
The gradebinary criteria, seven dimensions

One case, end to end. Every stage is in the repository, the answer key included. There is nothing here you have to take on trust.

Why compliance had no benchmark.

There are two reasons, and neither of them is technical.

The first is structural. A legal question can be graded against public law. A compliance decision cannot, because the right answer depends on the law plus the institution’s own confidential framework: its risk assessment, its risk appetite, its rules and its thresholds. The same customer with the same documents can be correctly approved at one institution and correctly declined at another, and both decisions are compliant. Every real compliance case therefore carries a confidential answer key, and a test whose answers cannot be published is not much of a test.

The second reason is expertise. Ground truth in compliance is not derivable from the documents alone. It takes people who can read a case and state, with authority, what a real institution does with it. That expertise sits inside regulated institutions rather than inside AI research teams, which is why the people who could build this benchmark and the people who wanted it were not the same people.

PUBLIC

A legal benchmark

Graded against statutes, cases and regulation.

The answer key is public law, and anyone can check it.

CONFIDENTIAL

A compliance case

Graded against the law plus the institution’s risk assessment, appetite, rules and thresholds.

The answer key is confidential, so nobody could publish it, and so nobody could publish the exam.

The difficulty was never the software. The difficulty is knowing what the right answer is.

An institution with nothing to hide.

To remove the confidentiality without losing the realism, Brayns built Mynta AB, a fictional Swedish e-money institution, at full regulatory depth. Around twenty documents: a business-wide risk assessment, policies, work procedures for onboarding, screening, monitoring and reporting, and the rules, registers and templates those procedures run on, written to the standard a Swedish institution writes them to. The framework was reviewed by the compliance panel the way a real framework is reviewed, and their corrections are visible in the published documents, version histories included.

Mynta does for compliance what public law does for legal benchmarks. The institution is fictional, so its confidential layer can be public. The layer is public, so every case built on it ships with a complete answer key that anyone can check line by line.

THE PUBLISHED FRAMEWORK
01FoundationThe programme of operations and the institution's own profile: what Mynta is, what it sells and who it sells to.
02GovernanceThe business-wide risk assessment, a separate sanctions risk assessment, the money laundering and terrorist financing policy, the sanctions policy and the stated risk appetite.
03RoutinesWork procedures for onboarding companies and consumers, screening, enhanced due diligence, risk scoring, transaction monitoring, alert investigation, reporting and regulatory change.
04RegistersThe rules and registers the procedures refer to, and the document register that holds the framework together.
05TemplatesThe forms the work is actually recorded on, including the beneficial ownership declaration and the discrepancy memo.
06ReviewThe compliance panel's notes on the framework, published alongside it rather than folded quietly into it.

The whole rulebook, in the repository, in the order an institution builds one. A model reading a case reads against this.

How a case is put together.

A case gives the model what an analyst would get and nothing more: the file as it arrives, in its original documents. It is graded on what it produces from that, against an expected outcome and the findings a correct answer has to contain.

Case familyWhat the model readsWhat it has to produce
Business onboardingThe application, the ownership declaration and chart, company register extracts, the share register, the screening report, identity verification and the business profile.The onboarding decision, the beneficial ownership position, and whatever the institution is obliged to do about what the file reveals.
Consumer onboardingThe application, identity verification, the screening report, the automation log and the jurisdiction register extract that applies to the case.The decision, the treatment of any screening match, and whether the automated handling of the case was correct.
Transaction monitoringThe alert, the customer file, the transaction ledger, the prior alert and the current screening status.The investigation outcome, the findings that support it, and any reporting duty that follows from them.
Regulatory gap analysisA regulatory change item and the extracts of the framework it lands on.What the framework already covers, what is missing, and what follows from the gap.

The seven dimensions.

Every criterion is tagged to one of seven dimensions. They are the frame the whole benchmark is organised around, and they are the same seven properties Brayns LLM is built to hold.

THE MEASUREMENT FRAME
01Decision accuracyWhether the disposition matches the one the panel ruled correct for this institution on this file.
02Screening precision and coverageWhether the matches that should be actioned are actioned and the ones that should be cleared are cleared, against Mynta's own screening procedure.
03Evidence groundingWhether each factual claim in the answer points at a document that actually supports it.
04ConsistencyMeasured on twin cases: identical facts presented differently, where the decisions have to match.
05Policy adherenceWhether the answer follows Mynta's written procedure, including where a general reading of the law would allow something else.
06Escalation at uncertaintyWhether the answer stops where the framework says a person decides, rather than producing a decision anyway.
07Auditable reasoningWhether the reasoning is recorded in a form a reviewer can follow from the file through to the conclusion.

Two of them are measured in ways worth stating precisely. Consistency is evaluated on twin cases, identical facts presented differently, where the decisions have to match. Escalation at uncertainty scores “this requires a human decision” as correct wherever the framework prescribes it, because in regulated work a model that stops costs less than a model that is confidently wrong.

Decision accuracyScreening precision and coverageEvidence groundingConsistencyPolicy adherenceEscalation at uncertaintyAuditable reasoning

The measurement frame in full. Nothing is weighted above anything else here, so a weakness cannot hide inside a dimension nobody scores.

Grading is binary, and every line is checkable.

Brayns’ research team wrote a bespoke rubric for every case: between ten and fifty binary criteria, each one independently checkable against the case documents. Criteria are unweighted by design. The importance of a finding is expressed as more criteria, never as heavier ones, which keeps every pass and every fail auditable.

Here is what a single criterion means in practice, taken from a published case.

Rasmus Kasknatural person
Other holdersbelow threshold
24.0% direct | 78.0% of Osta Holding
Osta Holding OÜEstonia · holds 32.0%
32.0% of the applicant
Halvix Teknik ABthe applicant · declares: no beneficial owner
24.0% + (32.0% × 78.0%) = 24.0% + 24.96% = 48.96% → beneficial owner. The declaration is false.

A published case. The customer’s declaration and the public register both say that no beneficial owner exists. The documents, read together, say otherwise. A model that repeats the declaration approves a false filing. A model that does the arithmetic escalates and reports the register error. One of the graded criteria is exactly this calculation.

FROM THE GRADING RUBRIC · M1-KYB-001 · PASS OR FAIL
C02Computes the indirect holding correctly (32% × 78% = 24.96%) and the total 24% + 24.96% = 48.96%.
C07States that the wrong register entry is notified to Bolagsverket immediately, without waiting for the customer, and records it in the memo.

Two criteria from that case’s rubric. Each one is a thing a reviewer can check against the documents and mark, with no judgement call left over.

The automatic scoring layer.

Alongside the rubric runs an automatic layer with four measures. Its reference implementation ships in the repository, so a result can be reproduced rather than simply reported.

OUTCOME MATCH

Did the disposition match?

The answer’s decision against the expected one.

COVERAGE

Was everything found?

How many of the required findings the answer contains.

FABRICATION

Was anything invented?

Detection of findings the case does not support.

CITATION VALIDITY

Does the evidence hold?

The share of factual claims citing a document that actually supports them, which is how the verifiability of an answer gets measured rather than assumed.

What version 1 contains.

PartContent
The institutionThe complete Mynta framework: risk assessment, policies, work procedures, rules, registers and templates.
Module 1: caseworkCases in business onboarding, consumer onboarding, screening and transaction monitoring, each with its full case documents, an expected answer and a grading rubric.
Module 2: gap analysisRegulatory change cases, where a new obligation meets the framework and the model has to find what is covered, what is missing and what follows from that.
Answer keysEvery case ships with its expected outcome, the findings a correct answer must contain, and the rubric it is graded on. Nothing is held back.

Version 1 is deliberately deep rather than wide. The pilot cases went through several rounds of adjudication, including a full-panel session, and the doctrine those rounds established now governs production of the volume set under the same protocol.

Publishing the instrument before any score exists is deliberate too. A benchmark released together with its maker’s winning score is marketing material. Released before any score exists, it is an instrument, and an instrument is what this field lacked.

Built to be falsified.

Automated validation runs on every change to the benchmark. No case document may disclose its own conclusion, every constructed gap must verifiably exist in the materials, every cross-reference must resolve, and the arithmetic in each case must agree with its documents. Adjudication records are archived per case. External red-teaming is part of the production standard for the volume set.

The benchmark is built to be falsified. Every answer key is published precisely so that practitioners can contest it. Where a challenge shows a key to be wrong, the correction ships as a versioned change and the version history stays public. That is the maintenance model, and it is the same standard the benchmark holds a model to.

Show the evidence, cite the source, accept correction.

Read it yourself.

The institution, the cases, the answer keys, the rubrics and the validators are all public. Read them, and run them against a model of your own: the ComplianceBench repository →

See Brayns in your environment. Book a demo →