Research

A compliance model cannot be built by one discipline.

The hardest part of building a model for compliance is not the model. It is knowing what a correct decision is, and that knowledge sits with people who have spent their careers making these decisions rather than in any public text.

What makes a compliance decision correct.

Every method begins by answering one question: what is the system being built toward. In most fields that answer is written down somewhere. A medical question has a literature. A legal question has statutes and decided cases. Anyone who disagrees with the answer can go and read the same source.

Compliance has no such source. The right answer to a case depends on the law plus one institution’s own framework: its risk assessment, its risk appetite, its rules and its thresholds. Two institutions can take the same customer, with the same documents, and correctly reach opposite decisions. What makes a decision correct is therefore not a fact anyone can look up. It is a judgement about what a particular institution does, made the way a practitioner makes it.

That judgement is not derivable from the documents alone. It takes people who can read a case and state, with authority, what a real institution does with it, and that expertise sits inside regulated institutions rather than inside AI research teams. This is the first constraint on any method for building a compliance model, and it is a constraint about people before it is a constraint about engineering.

What makes a decision correct is not a fact anyone can look up. It is a judgement about what a particular institution does.

Two kinds of knowledge, and neither is enough.

Two kinds of expertise bear on this problem. Each one fails on its own, and the two failures are worth stating plainly, because the method is a response to both.

The first failure is building well towards the wrong target. Model building is a real and scarce skill, and it optimises towards whatever objective it is handed. Where the true objective is expensive to evaluate, a cheaper proxy tends to stand in for it, and the proxy quietly becomes the thing the system is good at. Amodei and colleagues (2016) set this out as the problem of scalable oversight: how to hold a system to an objective you cannot afford to check directly. In compliance the true objective is a practitioner’s judgement on a specific file against a specific framework, which is precisely the kind of objective that is expensive to check. A team without access to that judgement does not build nothing. It builds something that resembles compliance work and is graded against something else.

The second failure is judgement that stays tacit. A practitioner who has run onboarding for most of a career can read a file and reach the right answer, and still have no way to make a system reach it every time without drifting. Knowing the answer and being able to specify it are different skills. Expertise that is never written down as something checkable cannot be built into anything: it stays case by case, resident in one person, and it leaves when they do.

MODEL BUILDING ALONE

Built well, towards the wrong target

The objective becomes whatever the team can measure.

Where the real objective is expensive to check, a proxy stands in for it, and the system becomes good at the proxy.

DOMAIN EXPERTISE ALONE

Right answers that never become a system

The judgement is sound and stays tacit.

It is never expressed as something checkable, so nothing can be built on it and it does not survive the person holding it.

Neither discipline covers for the other. Knowing what correct means and being able to turn that into a system are held by different people, and a compliance model needs both at full strength.

The method is the team.

The method follows from that. Brayns builds the model with both disciplines inside the work rather than one consulting the other: AI researchers from frontier AI labs, and compliance practitioners whose careers are measured in decades across banks, payment institutions and supervision. Neither group is an advisory board to the other.

The split is not by seniority, and it is not by phase. It is by what each discipline can actually settle. Practitioners settle what correct means. Researchers settle how a system is held to it. That is the sense in which this is a compliance model rather than a general one pointed at compliance: the people who decide what correct means are compliance people, and they decide it inside the build rather than around it.

On ComplianceBench, where this arrangement is visible in public, adjudication is not a review step at the end. It is where the answer keys come from. The panel ruled on every case, question by question. Where a case overstated a rule, the ruling cut it back. Where the framework itself had a gap, the ruling closed it, and several of those rulings are now written into the institution’s procedures with the legal reference attached. Brayns’ research team wrote the rubric for each case, which is the other half of the same operation: turning a ruling into criteria that can be checked one by one, against the documents, by anyone who cares to.

The panel is not a separate body from the product either. The practitioners who ruled on those cases are the same people whose judgement runs in the Brayns product. That is the point of the arrangement rather than a detail of it.

WHERE THE WORK CROSSES
Compliance practitionersSettle what correct means on a given file, for a given institution, and rule on it.
AI researchersSettle how a system is held to that ruling, case after case, without drifting.
The research team turns a practitioner's ruling into criteria that can be checked one at a time against the documents.
A ruling that cannot be expressed as something checkable goes back to the practitioner who made it, because a standard nobody can verify is not a standard.
A criterion that no institution would actually follow is overruled by the practitioner, however cleanly it formalises.

The return paths are the part that matters. Without them this is a handoff, and a handoff only carries the knowledge that survived being written down once.

The checking runs in both directions, and that is what makes this a method rather than a division of labour. Each discipline holds the other to its own test. The researcher asks whether a standard can be verified. The practitioner asks whether it is what a real institution does. A claim that fails either test is not part of the standard the model is built to.

A standard nobody can verify is not a standard. A criterion no institution would follow is not compliance.

Law has taken a comparable route. Guha and colleagues (2023) assembled a legal reasoning benchmark with the profession rather than around it. The pattern it suggests is the same one: where correctness is a professional judgement, there is a case for putting the professionals inside the construction rather than surveying them after it.

Held to the standard the domain sets.

Both disciplines need a language they can each hold to. That is what the seven dimensions are for. They are not a summary of the model’s qualities. They are the list of questions a compliance decision has to survive, and each one is stated in a form a practitioner recognises and a researcher can build against.

A practitioner reads them as the questions a supervisor asks about a decision already made. A researcher reads them as seven separate things that can be checked, one at a time, against a file. The same seven carry both readings, which is why they work as a shared standard rather than a shared slogan.

DimensionWhat a practitioner recognisesWhat a researcher can build against
Decision accuracyThe decision you would have to defend if it turned out to be wrong.A stated expected outcome for a file, which an answer either matches or does not.
Screening precision and coverageWhether what mattered was caught and the rest was cleared, against your own thresholds.A defined set of matches to action and to clear, checkable case by case.
Evidence groundingBeing able to point at the document behind every line of a memo.A claim-by-claim test of whether the cited document actually supports the claim.
ConsistencyTwo customers with the same facts being treated the same way.Twin cases whose decisions have to agree, which makes drift visible instead of anecdotal.
Policy adherenceFollowing the procedure as written, including where you would personally have done otherwise.A written procedure to compare an answer against, rather than a general reading of the law.
Escalation at uncertaintyKnowing when the file is not yours to decide.A defined condition where stopping is the correct output rather than a failure to answer.
Auditable reasoningA file another person can pick up years later and follow.A recorded line of reasoning that can be read from the documents through to the conclusion.

The same seven, read twice. A dimension only one discipline can state precisely is not yet a shared standard, and it would not survive the exchange above.

ComplianceBench is the public form of that standard. It publishes the institution, the cases, the answer keys and the rubrics, so the definition of correct this method is built around can be read and contested by people who do not work here. Brayns LLM is built around those seven properties and held to the standard the benchmark sets out.

Why it has to be built this way.

A general model is, on the published evidence about where this expertise sits, built largely by one of these disciplines without the other. That is not a criticism of the people who build them. It is a description of where the two kinds of knowledge sit. The result is a system optimised towards a target that was defined without the people who know what correct means, and no amount of capability closes that gap, because the gap is in the objective rather than in the model.

A compliance model has to arrive somewhere else. It can only get there if the people who know what correct means are inside the build and holding the standard, and the people who know how to build are holding that standard to something that can be checked. That is the method. The rest of it is consequence.

See Brayns in your environment. Book a demo →