ComplianceBench.
Compliance could never be benchmarked, because the answer key was always confidential. ComplianceBench removes the confidentiality without removing the realism.
ComplianceBench is an open benchmark that measures how AI models perform the core casework of financial compliance: business and consumer onboarding, screening decisions, transaction monitoring investigations, and regulatory gap analysis. It is built on a complete fictional institution whose entire regulatory framework is published, so every answer key can be inspected against the same documents the model was given.
Every case was adjudicated by Brayns’ compliance panel, practitioners who have run this work inside real institutions, and their rulings are the ground truth. This release contains the instrument and no scores.
One case, end to end. Every stage is in the repository, the answer key included. There is nothing here you have to take on trust.
Why compliance had no benchmark.
There are two reasons, and neither of them is technical.
The first is structural. A legal question can be graded against public law. A compliance decision cannot, because the right answer depends on the law plus the institution’s own confidential framework: its risk assessment, its risk appetite, its rules and its thresholds. The same customer with the same documents can be correctly approved at one institution and correctly declined at another, and both decisions are compliant. Every real compliance case therefore carries a confidential answer key, and a test whose answers cannot be published is not much of a test.
The second reason is expertise. Ground truth in compliance is not derivable from the documents alone. It takes people who can read a case and state, with authority, what a real institution does with it. That expertise sits inside regulated institutions rather than inside AI research teams, which is why the people who could build this benchmark and the people who wanted it were not the same people.
A legal benchmark
Graded against statutes, cases and regulation.
The answer key is public law, and anyone can check it.
A compliance case
Graded against the law plus the institution’s risk assessment, appetite, rules and thresholds.
The answer key is confidential, so nobody could publish it, and so nobody could publish the exam.
The difficulty was never the software. The difficulty is knowing what the right answer is.
An institution with nothing to hide.
To remove the confidentiality without losing the realism, Brayns built Mynta AB, a fictional Swedish e-money institution, at full regulatory depth. Around twenty documents: a business-wide risk assessment, policies, work procedures for onboarding, screening, monitoring and reporting, and the rules, registers and templates those procedures run on, written to the standard a Swedish institution writes them to. The framework was reviewed by the compliance panel the way a real framework is reviewed, and their corrections are visible in the published documents, version histories included.
Mynta does for compliance what public law does for legal benchmarks. The institution is fictional, so its confidential layer can be public. The layer is public, so every case built on it ships with a complete answer key that anyone can check line by line.
The whole rulebook, in the repository, in the order an institution builds one. A model reading a case reads against this.
How a case is put together.
A case gives the model what an analyst would get and nothing more: the file as it arrives, in its original documents. It is graded on what it produces from that, against an expected outcome and the findings a correct answer has to contain.
| Case family | What the model reads | What it has to produce |
|---|---|---|
| Business onboarding | The application, the ownership declaration and chart, company register extracts, the share register, the screening report, identity verification and the business profile. | The onboarding decision, the beneficial ownership position, and whatever the institution is obliged to do about what the file reveals. |
| Consumer onboarding | The application, identity verification, the screening report, the automation log and the jurisdiction register extract that applies to the case. | The decision, the treatment of any screening match, and whether the automated handling of the case was correct. |
| Transaction monitoring | The alert, the customer file, the transaction ledger, the prior alert and the current screening status. | The investigation outcome, the findings that support it, and any reporting duty that follows from them. |
| Regulatory gap analysis | A regulatory change item and the extracts of the framework it lands on. | What the framework already covers, what is missing, and what follows from the gap. |
The seven dimensions.
Every criterion is tagged to one of seven dimensions. They are the frame the whole benchmark is organised around, and they are the same seven properties Brayns LLM is built to hold.
Two of them are measured in ways worth stating precisely. Consistency is evaluated on twin cases, identical facts presented differently, where the decisions have to match. Escalation at uncertainty scores “this requires a human decision” as correct wherever the framework prescribes it, because in regulated work a model that stops costs less than a model that is confidently wrong.
The measurement frame in full. Nothing is weighted above anything else here, so a weakness cannot hide inside a dimension nobody scores.
Grading is binary, and every line is checkable.
Brayns’ research team wrote a bespoke rubric for every case: between ten and fifty binary criteria, each one independently checkable against the case documents. Criteria are unweighted by design. The importance of a finding is expressed as more criteria, never as heavier ones, which keeps every pass and every fail auditable.
Here is what a single criterion means in practice, taken from a published case.
A published case. The customer’s declaration and the public register both say that no beneficial owner exists. The documents, read together, say otherwise. A model that repeats the declaration approves a false filing. A model that does the arithmetic escalates and reports the register error. One of the graded criteria is exactly this calculation.
Two criteria from that case’s rubric. Each one is a thing a reviewer can check against the documents and mark, with no judgement call left over.
The automatic scoring layer.
Alongside the rubric runs an automatic layer with four measures. Its reference implementation ships in the repository, so a result can be reproduced rather than simply reported.
Did the disposition match?
The answer’s decision against the expected one.
Was everything found?
How many of the required findings the answer contains.
Was anything invented?
Detection of findings the case does not support.
Does the evidence hold?
The share of factual claims citing a document that actually supports them, which is how the verifiability of an answer gets measured rather than assumed.
What version 1 contains.
| Part | Content |
|---|---|
| The institution | The complete Mynta framework: risk assessment, policies, work procedures, rules, registers and templates. |
| Module 1: casework | Cases in business onboarding, consumer onboarding, screening and transaction monitoring, each with its full case documents, an expected answer and a grading rubric. |
| Module 2: gap analysis | Regulatory change cases, where a new obligation meets the framework and the model has to find what is covered, what is missing and what follows from that. |
| Answer keys | Every case ships with its expected outcome, the findings a correct answer must contain, and the rubric it is graded on. Nothing is held back. |
Version 1 is deliberately deep rather than wide. The pilot cases went through several rounds of adjudication, including a full-panel session, and the doctrine those rounds established now governs production of the volume set under the same protocol.
Publishing the instrument before any score exists is deliberate too. A benchmark released together with its maker’s winning score is marketing material. Released before any score exists, it is an instrument, and an instrument is what this field lacked.
Built to be falsified.
Automated validation runs on every change to the benchmark. No case document may disclose its own conclusion, every constructed gap must verifiably exist in the materials, every cross-reference must resolve, and the arithmetic in each case must agree with its documents. Adjudication records are archived per case. External red-teaming is part of the production standard for the volume set.
The benchmark is built to be falsified. Every answer key is published precisely so that practitioners can contest it. Where a challenge shows a key to be wrong, the correction ships as a versioned change and the version history stays public. That is the maintenance model, and it is the same standard the benchmark holds a model to.
Show the evidence, cite the source, accept correction.
Read it yourself.
The institution, the cases, the answer keys, the rubrics and the validators are all public. Read them, and run them against a model of your own: the ComplianceBench repository →
See Brayns in your environment. Book a demo →