T.R.U.S.T.
Trust should be earned at the moment of reliance.
A one-line summary needs a light scaffold. A hiring, financial, medical, legal, or customer-impacting recommendation needs evidence, assumptions, uncertainty, alternatives, accountable review, and a clear route to override.
Design principle
Design so people can say: I know what this system did, what it based that on, where it may be wrong, and what I can do about it.
T.R.U.S.T.
Five pillars of calibrated trust. Each section includes what to build and patterns that put it into the product.
Truthful expectations
Realistic mental model
- Question
- What can this AI do, how well, and where does it fail?
- Principle
- Trust begins before the first output. Set an accurate mental model of role, capabilities, data boundaries, and non-capabilities. Aim for informed use, not disclaimer overload.
- Breaks as
- “AI-powered” with no benefit or limit, human-like certainty the system does not have, or fine-print limits only after someone acts on bad output.
- Build
- Clear AI role label: draft, recommendation, or assisted result
- Capability framing in task language, not model marketing
- Boundary statement for what the system does not do
- First-use orientation at the moment the feature is used
- Sandbox mode for low-stakes trials before real work
- Model or mode disclosure when speed, quality, privacy, or cost change
Reasons & provenance
Inspectable basis for reliance
- Question
- What evidence, inputs, and criteria informed this output?
- Principle
- People should answer “Why did the AI say this?” with task-useful evidence. Provenance is often more actionable than a bare confidence score.
- Breaks as
- A source dump at the bottom with no claim mapping, or explainability theater that does not help someone challenge the basis of a claim.
- Build
- Claim-level citations tied to specific factual statements
- Evidence excerpts: the paragraph, record, or row behind the claim
- “Why this?” disclosure for recommendations and classifications
- Input summary of active documents, range, filters, and instructions
- Assumption cards for high-impact premises
- Source quality labels: freshness, coverage, conflicts, missing evidence
| Layer | User need | UI pattern |
|---|---|---|
| Glance | Can I broadly rely on this? | AI label, source count, freshness, concise limitation |
| Inspect | What supports this claim? | Inline citation, evidence excerpt, assumption |
| Investigate | How did the system reach this? | Source panel, criteria, input context, method |
| Audit | Can I reconstruct and defend this decision? | Exportable evidence record, data lineage, version history |
Uncertainty & limits
Appropriate caution
- Question
- What is unknown, variable, incomplete, or contested?
- Principle
- Make uncertainty specific and actionable. Communicate when results are strong and when users should apply more judgment.
- Breaks as
- An ambiguous 78% score with no meaning, or hiding incompleteness until the user has already relied on the answer.
- Build
- Plain-language uncertainty labels for incomplete, estimated, or conflicting evidence
- Assumption confirmation before high-impact continuation
- Alternative interpretations when the request is ambiguous
- Evidence gap callouts: what would increase confidence
- Verification prompts before external or consequential use
- Ranges or scenarios instead of a falsely exact single forecast
- Confidence only with context: what it measures and what to do next
Safeguards & user control
Preserved agency
- Question
- Can the user inspect, correct, reject, or override the AI?
- Principle
- Every AI output should offer a path to inspect, edit, reject, correct, or escalate. Trust is earned when people can influence the system, not only accept it.
- Breaks as
- Accept-only flows, buried human handoff, or no undo when AI changes durable content or records.
- Build
- Edit-before-use for outputs treated as drafts
- Alternative non-AI path when users need it
- Fine-grained feedback: inaccurate, stale, irrelevant, unsafe, biased, misunderstood
- Appeal or review for access, eligibility, money, health, safety, or employment
- Data controls: scope, removal, personalization reset, retention clarity
- Undo and version history for durable changes
- Visible human handoff, not a buried support link
| Control | What it enables | Example |
|---|---|---|
| Inspect | See evidence, context, and limits | Open citations and source excerpts |
| Adjust | Change inputs or preferences | Change audience, source scope, or constraints |
| Correct | Fix an inaccurate result | Edit an extracted field or flag a false claim |
| Reject | Decline without penalty | “Not relevant” or “Use a different approach” |
| Override | Replace AI judgment | Choose another candidate, outcome, or route |
| Escalate | Bring in an accountable owner | Send to compliance, support, or reviewer |
| Revoke | Withdraw data or permission | Disable a connector or reset personalization |
Other related patterns
- Human in the loopKeep approve and override next to the output
- Human in the loopGate high-impact actions before they land
- Memory Scope ToggleControl what persists before it sticks
- Data Ownership & ControlSee and revoke data used for personalization
- Granular ConsentWithdraw permissions as easily as granting them
Track record & repair
Durable, earned trust
- Question
- How does the product demonstrate reliability over time and recover from mistakes?
- Principle
- Trust accumulates through repeated, predictable behavior and is lost quickly when failure is opaque. Acknowledge errors, preserve work, correct outcomes, and learn without blaming the user.
- Breaks as
- Silent degradation, no correction receipt after feedback, or incidents with no explanation of what was affected.
- Build
- Output history to revisit, compare, and trace prior AI work
- Version and freshness status when knowledge or models change
- Reliability status for outages, degradation, or unavailable sources
- Correction receipts when feedback changes saved work
- Incident explanation: what happened, what was affected, what was fixed, what to do next
- Quality review loops for recurring domain errors
Scaffold by situation
Match the minimum scaffold to the reliance moment. Stronger stakes need stronger structure.
| Situation | Minimum scaffold |
|---|---|
| Low-stakes creative draft | AI label, editability, simple feedback |
| Factual answer or research summary | Claim-level citations, source freshness, limitation note |
| Recommendation affecting a decision | Reasons, criteria, assumptions, alternatives, rejection path |
| Prediction or forecast | Uncertainty range, input basis, scenario comparison |
| High-impact decision support | Evidence record, human review, override, escalation, auditability |
| Autonomous or external action | Add Agentic UX: authority, approval, monitoring, recovery |
Design workflow
Use T.R.U.S.T. as a design-review and product-definition process.
- 1
Identify the reliance moment
Ask what decision or action a user might take because of this AI output. Document consequence if wrong, user expertise, available evidence, reversibility, and whether an alternative workflow exists.
- 2
Assign the right scaffold
Match stake and output type to the minimum scaffold. Light for drafts; stronger for recommendations, forecasts, and high-impact decisions.
- 3
Test for overtrust and undertrust
Include cases where the AI is correct and well evidenced, correct but weakly evidenced, plausible but wrong, ambiguous, out of scope, or under-informed. Ask what users would do next.
- 4
Measure calibration
Track mental model accuracy, evidence use, appropriate reliance, overtrust, undertrust, control success, repair time, and long-term trust by task type.
Cheatsheet
Eight checks before shipping, aimed at calibrated reliance.
- Can users accurately describe what this AI can and cannot do?
- Can they see the evidence or inputs behind claims that matter?
- Is uncertainty specific and actionable when stakes rise?
- Can they inspect, edit, reject, override, or escalate without dead ends?
- Does the product show reliability over time and repair opaque failures?
- Does scaffold strength match the reliance moment, not a one-size disclaimer?
- Do people verify when evidence is weak, and rely when it is strong?
- After an error, can they recover work and understand what changed?
