Key takeaways in 3 minutes
A confidence score tells you something about a model output. It does not tell a decision-maker whether the evidence is sufficient, current or relevant.
A useful AI decision brief separates:
- the decision and consequence;
- the recommendation;
- evidence and provenance;
- uncertainty, gaps and conflicts;
- alternatives and counter-evidence;
- the action being requested.
Design explanations for verification, not persuasion. If the person cannot find a practical reason to disagree, the interface is presenting an answer rather than supporting a decision.
94% confident.
It looks reassuring in the header.
It is crisp, numerical and just scientific enough to suggest that somebody, somewhere, has done the difficult thinking.
But confidence in what?
Perhaps the AI is highly confident that a supplier will miss a delivery. Perhaps the underlying shipment data is six hours old. Perhaps it has not seen the temporary production shutdown recorded in an email. Perhaps three other suppliers are worse. Perhaps the recommended intervention costs more than the late delivery.
The percentage may describe the model output perfectly.
It still does not tell a planner whether to act.
A confidence score tells you something about the model. A decision brief tells you whether a person has enough to act.
The Answer And The Decision Are Different Objects
An AI answer is usually a prediction, classification, summary or recommendation.
A decision brief is an interface between that answer and accountable human judgement.
It has to answer a wider set of questions:
- What decision are we actually making?
- What supports this proposal?
- Where did that evidence come from?
- What is missing, stale or disputed?
- What alternatives were considered?
- What happens if we do nothing?
- What happens if the recommendation is wrong?
This is why adding a paragraph labelled “Reasoning” does not finish the design.
A fluent explanation can make a weak answer feel more coherent. Studies of human–AI decision-making have found that explanations can increase reliance and that useful friction or source inspection may sometimes reduce over-reliance more effectively. The right question is not “Did we explain the model?” It is “Did we help the person verify the proposal?” See Buçinca, Malaya and Gajos and Vasconcelos et al..
Design The Brief In Six Layers
Anatomy Of An AI Decision Brief

1. The Decision
Lead with the decision, not the model activity.
“Approve a two-day supplier hold” is clearer than “Risk anomaly detected”. It tells the person what is being asked of them and establishes the scope.
Include the affected supplier, order, customer or data object. Include the decision deadline and what happens if nobody responds.
2. The Consequence
Show what changes if the proposal is accepted.
Cost, service, time, access, customer commitment, regulatory exposure or operational disruption may matter more than the model's confidence. Put the consequence close to the decision rather than burying it in an expandable panel called Details.
3. The Evidence
Connect each important claim to inspectable evidence.
Distinguish direct records from inferred content. Show source, owner and freshness where they affect judgement. If a policy rule triggered the proposal, name the rule and the version applied.
A citation badge is not provenance. The reviewer needs to see which claim the source supports and whether it was actually consulted.
4. The Gaps And Conflicts
Absence is part of the brief.
Show stale data, missing fields, unresolved discrepancies and inputs outside the model's normal operating range. If the invoice and purchase order match but the bank-detail change has not been verified, that missing verification belongs in the primary reading path.
Uncertainty should change the flow, not decorate it. Weak evidence may require a hold. Conflicting signals may require escalation. High consequence may require verification regardless of confidence.
5. The Alternatives
One polished recommendation creates an anchor.
Show the credible alternatives and their trade-offs: approve the hold, reduce its scope, request evidence, use a different supplier or take no action. The person should be able to compare outcomes rather than reverse-engineer alternatives from the machine's preferred answer.
6. The Ask
End with the exact decision or action required.
The reviewer should know whether they are approving a proposal, authorising an external action, acknowledging a risk or choosing between scenarios. Those are different commitments and should not share one vague Continue button.
Progressive Disclosure Without Convenient Amnesia
Not every reviewer needs a wall of provenance and model diagnostics.
Progressive disclosure is the right pattern. The mistake is using it to hide the awkward evidence.
The first layer should contain what any responsible decision-maker must know. Deeper layers can hold source documents, rule details, system versions and trace data for expert inspection.
The ICO's explaining-AI guidance similarly emphasises contextual, meaningful explanation and explicitly recognises the value of UX design in presenting it appropriately.
The product designer's job is not to display everything.
It is to decide what cannot responsibly remain hidden, and/or discoverable.
Example: The Supplier Risk Recommendation
Bad brief:
`text
SUPPLIER RISK: HIGH
Recommendation: switch supplier
Confidence: 94%
[Approve]
`
Decision brief:
`text
Decision: move Order 1847 from Northstar to Vale Components
Deadline: before the 16:00 production release
Consequence: +£18,400 cost; protects 82% of expected service
Evidence:
- Northstar delivery estimate: 9–12 days (updated 10:42)
- Required arrival: 6 days
- Vale confirmed capacity for 60% of volume (email verified 11:06)
Unknowns:
- Vale capacity for remaining 40% is not confirmed
- Northstar recovery plan requested, not received
Alternatives:
- Split the order
- Hold for Northstar response until 14:00
- Accept the service risk
Requested action: choose a recovery path
`
The second version is longer.
That is because the decision is larger than the answer.
The Practical Move: Run The Disagreement Test
Open one AI recommendation in your product and ask a domain expert:
What could you inspect here that might reasonably make you disagree?
If the answer is “not much”, do not celebrate consistency.
Inspect the brief for:
- source-to-claim links;
- freshness and coverage;
- missing and conflicting evidence;
- alternatives;
- consequence;
- scope and deadline;
- a legitimate route to ask for more information.
The Accountability Lab Connection
An accountable decision record starts before the click.
It needs a snapshot of the brief as it was presented: the evidence available, what remained unknown, the applicable rule and the alternatives the human could see.
That is one of the ideas I am exploring through Accountability Lab: how the interface can produce better accountability evidence as a natural by-product of helping someone make a better decision.
No retrospective archaeology required.



