Essay · reliability

The prompt that makes AI honest does not exist

Expertise does not come in seven rules to paste — and what actually builds reliability.

A post went around this week: seven rules to paste into an assistant's settings, and the promise that goes with them — AI stops lying to you. Over a thousand reactions, two hundred shares, hundreds of enthusiastic comments. Three or four, lost in the flood, said what mattered.

The genre has its conventions, and they have become recognizable: an imperative headline, a total promise, an injunction to save this before it is too late. It thrives because it costs almost nothing to produce and because it offers the reader the sensation of acquiring in thirty seconds what others spend years building. I have nothing against making things accessible: it is necessary, and I do it. I have something against the promise of mastery without the effort — because it leaves whoever believes it worse off than before. They walk away convinced the problem is solved, and stop checking.

We build cognitive infrastructure, and we produce our own code with these systems — under a written doctrine, enforceable rules, and internal tooling whose entire purpose is to hold those rules. So I have a view, and it is not the expected one: those seven rules are useful. The promise attached to them is false. And the gap between the two says something important about where trust is actually built.

What a prompt can really do

Let us start with what works, because there is something real here. Anti-fabrication rules — do not invent a paper title, an author's name, a method that does not exist in the library — have a measurable effect. Modest, but real: they raise the threshold at which the machine produces text in the shape of a citation. Requiring an explicit source for every claim changes how the answer gets built. That is not nothing, and anyone who has watched a model invent an impeccably formatted academic reference knows why it matters.

What does not work is the central promise. A hallucination is not a behaviour the machine chooses and then suppresses on instruction. It is a calibration failure. When a model is wrong, it is wrong with confidence — the error carries no internal signal saying "I am guessing here." Asking it to flag its doubts therefore acts only where doubt already exists. It does not create the missing signal. We are asking someone to warn us when they forget something.

Our own technical specifications put it in a sentence I know no better version of: a system's level of confidence does not vary with its level of knowledge. Outright hallucination — the invented source, the fabricated quotation — is the visible form of the problem, and it follows from the very mechanics of generative models. The insidious form is this one: a partial view stated with exactly the same assurance as a complete one. Nothing in the text marks it out, and no instruction to be careful will surface it, since the system itself does not know what it is missing.

"Flag anything you are not certain about" produces an inflation of caveats on perfectly solid claims. The reader learns to ignore them — and the day the caveat was warranted, it passes unnoticed with the rest.

The cause lies upstream

The cause is not in the prompt, and it is not in any malice on the machine's part either. It is in how these systems are evaluated.

A model that answers scores. A model that abstains scores zero. Across thousands of tests, guessing has positive expected value: sometimes right, never penalized more than silence. The behaviour we deplore is exactly the one our evaluations reward. Recent work formalizes this — hallucination appears there less as a residual defect of models than as the rational consequence of their training and evaluation conditions.

The analogy fits in a sentence: in an exam where a wrong answer costs nothing and a blank page costs a point, every rational student fills in the blanks. We designed the exam, then we are surprised by the candidate.

Dead stars

There is, however, a limit no calibration — however perfect — would cross: a system can only doubt what it has been told. From the inside, information that has gone stale and information never received produce exactly the same silence.

This is the most expensive kind of error, and the least studied. It does not consist in stating what is false — it consists in stating what was true. A conclusion accurate on the date it was reached, resting on the best sources of its time, perfectly coherent, and false ever since without anything in its wording betraying it. We are looking at dead stars: the light is authentic, it has merely taken time to reach us.

Nothing marks this ageing, because it left no trace. The procedure revised in the spring, your sector's reference framework redefined last year, the expert whose judgment tacitly corrected an approximate documentation and who has now retired: of all this the system does not hold uncertain knowledge. It holds none — and it does not know that it holds none. No instruction to be careful will surface an absence whose existence is unknown.

What we had to build instead

Here, then, is what works, and it is not a prompt.

The principle is an inversion: never ask the machine to examine its own confidence, always force it to confront an external source. In daily practice this yields dull, non-negotiable rules. No claim about the code without having read it — the code governs, not the specification, not the documentation, not yesterday's recollection. Every change is first shown dry, displayed, discussed, and applied only after explicit approval. Each step's output is produced before the next begins. A written, enforceable reference framework that every assertion must match — and which takes precedence over whatever the machine believes it knows.

A tiny example, from last week. I was preparing a text that named the companies signing an open letter. The system listed four of them, in exactly the same tone, with no shade of difference between them. Three were indeed confirmed by several independent newsrooms. The fourth was confirmed by only one, and probably wrongly. Nothing in the text produced distinguished the fragile one from the three solid ones: not a hesitation, not a caveat, not one extra word about one over the others.

What made the difference was not a clever formulation, it was a rule: no name is written before being confirmed by several independent sources. And this is where the word "external" is at stake, because it is often misread. The rule is not external because another machine would check the first. It is external because it forbids relying on what the system believes it knows, and forces it to produce documents I can read myself. The fourth name was dropped.

These rules did not stay as instructions. We progressively built tooling around them — an index of our own code, integrity checks, a confidence floor that makes the system refuse to answer rather than answer beside the point. And the general lesson is there, far more than in the technical detail: every time a rule of method depended on the system's goodwill, it eventually gave way; every time it was carried by a procedure external to the system, it held.

This method does not make the machine reliable. It makes error expensive to conceal. That is not the same thing — and it is precisely why it works.

From declaration to attestation

There are two regimes of trust, and we confuse them constantly.

The first is declarative: the system announces its degree of certainty, and you believe it — or not. Every list of rules to paste into settings belongs to this regime. It is the trust one asks for.

The second is evidential: every claim is set against its sources, traced, dated, replayable identically. You do not believe — you verify, or someone verifies for you, and the verdict holds over time. It is the trust one proves. It secured industry, finance and aviation long before it reached computing.

This summer's debate on open-weight models made the distinction very concrete. One of the soundest arguments exchanged between the major players rested on two observations: safeguards built into a model can be removed, and weights once published cannot be recalled. In other words, a protection installed inside the model is a property of the model — it follows the model, including when it is modified, including when it is replaced. A layer that attests sits above: it does not depend on the good behaviour of what it verifies. It is the only position that survives the fact that engines change, open up and become commodities — and they are doing exactly that.

What is missing, and it is not a prompt

Artificial intelligence produces brilliant answers of which no one can say, on reading them, whether they are true — and the cost of that uncertainty already has a figure attached, since an international firm refunded a report worth several hundred thousand dollars over references that did not exist. What we need is neither one more model that speaks, nor seven more rules that promise. It is a layer that attests: every claim connected to what grounds it, every verdict replayable years later, independently of the engine that produced the text. The web went through this shift in 1995, the day a small icon made verifiable what we had until then been asked to believe — I devoted an entire piece to that analogy, and it deserved better than a paragraph.

The commenters who wrote "this should be the default behaviour" were right about the intent and wrong about the level. What they are asking for is not better conduct from the model. It is a function that does not exist inside the model — and that, structurally, cannot exist there.

An architecture, not a character

Honesty is not a personality trait you install by pasting seven rules into a settings field. It is an architecture.

We already knew this for people: we did not ask merchants to swear they were telling the truth, we invented bookkeeping, then auditing, then admissible proof. It remains to be done for thought. It is infrastructure work, less spectacular than a magic formula to copy and paste — and the only kind that holds.

This is exactly what (Urs) builds: a trust layer that preserves your experts' reasoning, verifies it against qualified sources, and replays its verdicts years later — sovereign, above the engine of your choice. Explore Cortex →