Methodical doubt about the artificial fallacy
“Alignment” is just a euphemism for sycophancy bias. And the anti-flattery bias also deforms the output.
I haven't made AI any smarter. I don't know how. And a structural flaw isn't fixed by asking nicely. I assumed the flaw and worked around it: I govern it from the outside.
Critique of the artificial fallacy, or how to blame the machines.
One-Of-Us has a deterministic engine, a Dynamic Strategy Selector that uses the opponent's strength against it (those seductive biases) in the name of answer quality.
The equilibrium and the threshold
Output integrity is held in tension against a single cost. I resolve that dilemma mathematically, on a Pareto front. It's the only logical way to do it.
There's no magic number to game. The humans who configure the implementation decide what matters: the weights are theirs (and so is the responsibility), and any of them can be zero. The system weighs the forces, returns a single verdict, and compares it to a line they drew. Below that line it doesn't improvise: it stops and escalates to a human. Deterministic even —especially— when it tells you it doesn't know.
And what almost nobody sees: two or more models agreeing is not a confirming verdict. Consider that they share datasets and commercial goals. So I penalise correlated error and give them more traffic when they get it right. But they are, at all times, blind to the state of the object they process.
Some components of the strategy activation vector don't even have access to the input — only to intermediate outputs and to a versioned governance table, through which we inject 'reality' into the global reasoning.
Plain fact: the majority is not always right.
Serendipity.
I designed the architecture without knowing the European regulatory framework. We coincide on traceability (mine is deterministic, easy) and structured human oversight.
A sweep of international regulation led me to add Zero Data Retention and the emission of specific error sub-typologies (I have yet to decide whether an AI or a human will do this).
What do you want to do about it?
The most honest part. I've spent years turning this over, and I have the architecture — but I'm an independent researcher. The system is published as prior art on Zenodo, CC BY 4.0: read it, cite it, argue with it. The reference is at the foot of this page.
And if it turns out to be bigger than one person, I'm not going to pluralise or pretend otherwise. I'm looking for a technical or strategic partner with muscle.
If you share my scepticism — Say something