The invented risk — why a fluent AI status report is more dangerous than a boring one
Thomas Thejn4 min read
An AI status narrative is dangerous when it is fluent and wrong — a risk nobody logged, a decision nobody took, stated with confidence in a steering pack. TransformRadar's AI may only reframe the project record: every prompt carries the data that licenses each claim, outputs are schema-validated, and a person edits every draft before it goes anywhere.
A steering meeting. The pack was drafted by an AI feature, and it reads beautifully. Under "key risks" there is an elegant paragraph about integration testing capacity in the third quarter. Confident. Well structured. Correctly formatted.
The sponsor looks up. "Who owns that one?"
Silence. Nobody logged that risk. Nobody raised it. The governance log has three entries, one of which is "TBD". The model, asked for a risks section and given very little to work with, produced the kind of risk that programmes like this one usually have. It was not lying, exactly. It was doing what a well-brought-up language model does when the fridge is empty: describing a nice meal.
That is the invented risk, and it is the single most important failure mode in AI-assisted governance. Not because it is common, but because you cannot see it from the inside. The paragraph is grammatical, plausible and identical in tone to everything around it. The only way to catch it is to know the record better than the model does, which is precisely the work the AI was supposed to save you.
Fluency is not evidence
Language models are trained to write well. They are astonishingly good at it. That is also the problem: a model asked for a narrative will produce one whether or not the inputs support it, and the thinner the record, the more it fills in from its general sense of how IT and digital transformation programmes tend to go. The result is most confident exactly where the evidence is weakest, which is the opposite of what a CIO wants from a portfolio pack.
A status report is worth reading because someone accountable stood behind each line. A narrative that may or may not correspond to the record is a liability wearing a nice suit.
The rule: reframe, never invent
TransformRadar's AI works under one rule, and it shapes every feature: it reframes the project's own data, and it does not invent. Four things follow.
- The prompt carries the evidence. Every prompt includes the structured
input that licenses every claim the model is allowed to make — the slipped tasks, the open risks by id, the movement in the health score — and instructs the model to mention nothing outside it.
- The output is validated. Every response is checked against a schema
before anyone sees it. Raw model text is never rendered. A short list of banned phrases catches the hedging and filler that signal a model reaching beyond its inputs.
- A person edits every draft. The weekly status is filled by the AI and
submitted by the project lead, field by field. The steering pack is drafted and reviewed by the chair. Nothing is circulated on its own.
- An empty week produces an empty draft. The evaluation harness checks,
against fixtures, that a project where nothing happened yields no invented entities. Silence is the correct output for silence.
The clever check we had to remove
It is worth being honest about how this evolved, because the story is funnier than the rule. An early version added a runtime heuristic on top: every capitalised word in a draft had to appear somewhere in the input, on the theory that invented names and places would be capitalised. Rigorous. Elegant. Guaranteed to catch a made-up supplier.
It rejected a draft because a sentence began with "Finalised".
Then it did it again. The weekly status draft went down in production twice, defeated by the English convention of starting sentences with a capital letter. We removed it. Grounding is now enforced by the prompt, by evaluation against fixtures, and by the human gate, plus two objective runtime checks: the schema and the banned-phrase list.
The lesson was not "checks are bad". It was that a clever check which fails on normal language is worse than no check, because it teaches people to distrust the safeguard instead of the model. The safeguard has to be as reliable as the thing it guards.
Why this matters more, not less, as agents arrive
As AI moves from drafting to acting — chasing task owners, tracing a health drop to its causes, coordinating with other systems — the grounding rule becomes the whole game. An agent that may only cite the record has to find the actual causes of a problem, and when it cannot, it says so. That is a feature. A Programme Health drop with no traceable cause in the record is itself a governance finding: something is happening that nobody has written down, and someone should go and ask.
The previous post listed five questions to ask a vendor. This was the fourth one, at length. If you want to see the rule in action, start with one project, leave the governance log empty, and ask for a status draft. The honest answer is a short one, and it will not mention Q3.
- AI
- governance
- status reporting
- steering