Auditing generative AI.
A deterministic, auditable control layer over the behaviour of generative AIs.
Part of the architecture I published on Zenodo for fear that someone would get there first. But the evolution of the theoretical model, the formulas, the constants and the calibration parameters are under lock and key. And they are only shared after signing an NDA.
It may sound blunt, but I'm inside the system and the system is what it is. Whether I like it or not. A system that audits others has no right to be a black box too. Or it would lose its claim to reason. I'm looking for a balance between defending my intellectual property and what I believe is fair. Hard.
They answer with confidence, even when they're wrong.
Generative models are probabilistic. In regulated environments, that leaves four structural cracks:
Failing to harness the power of probabilistic AIs because of biases in their design would be, to my mind, a mistake of calibre. The initial solution, designed for regulated environments, can be applied to any production cycle where including AI brings a competitive advantage over exclusively human work.
It can also be extrapolated to agentic AI and to B2C.
It doesn't generate. It doesn't fix the model. It audits it.
One-Of-Us is a conditioned deterministic AI that sits one level above: it supervises, audits and confines generative AIs, and issues a verdict of reality and integrity on every output.
A single cycle, from input to verdict.
None of them receives the same information: several are denied sight of the object, and judge only the result. Blinding them to the state of the object mitigates the biases of the probabilistic model: no auditor fails on the same flank.
This architecture has been made obsolete by the model's development. I reserve the final proposal as a trade secret.
Four states. One only, with emitter and trace.
Passes every test. The output is reliable.
→ releasedSemantic, governance or adversarial failure.
vetoed by an LLMAsserts data not present in the original object.
issued by BoolM1Breaks the required format or structure.
issued by BoolM2Tetra-state logic has pedigree. In 1977, the logician Nuel Belnap proposed a four-valued logic —he called it «how a computer should think»— so that a machine could reason with incomplete or contradictory information without breaking: true, false, neither and both at once. STABLE, UNSTABLE, NULL and INVALID descend from that lineage: four verdicts for a world where data can be missing, redundant or contradictory. Half a century later, Belnap's question resurfaces.
If the output is not reliable, it is escalated to a human supervisor. It never leaves without judgement.
From the pair to the whole.
First I tried the obvious: measuring covariance between pairs of models, to punish those that fail together. In a multi-model ensemble that doesn't work: the pairs multiply, third parties contaminate every measure, and the calculation says less and less about what the ensemble does.
It's way cooler —and more exact— to penalise the error of the whole executed strategy, as a complete configuration, in production or in shadow mode. If an alignment fails, that specific strategy stops being eligible.
The equilibrium and the threshold
Output integrity is held in tension against a single cost. I resolve that dilemma mathematically, on a Pareto front. It's the only logical way to do it.
There's no magic number to game. The humans who configure the implementation decide what matters: the weights are theirs (and so is the responsibility), and any of them can be zero. The system weighs the forces, returns a single verdict, and compares it to a line they drew. Below that line it doesn't improvise: it stops and escalates to a human. Deterministic even —especially— when it tells you it doesn't know.
It doesn't improve one isolated AI. It improves the whole.
It learns from its hits and misses —detected failure patterns and human validation— in a shadow phase, never on the client's data. With that signal it recomposes which components, in which roles, form the pool.
The system tends to precision.
Governance, not promise.
Serendipity. I designed the architecture without knowing the European regulatory framework. We coincide on traceability (mine is deterministic, easy) and structured human oversight.
A sweep of international regulation led me to add Zero Data Retention and the emission of specific error sub-typologies (I have yet to decide whether an AI or a human will do this).
Infrastructure for an AI that can be trusted.
Architecture partially published and validated by simulation · on the way to Proof of Concept.
Heads up! The button takes you, obviously, to PayPal. It's a plain link (paypal.me). No PayPal widget, no PayPal code on this page. I don't want your data, truly. What happens —and what does NOT happen— to your data if you donate →