Designing for trust in generative AI
Reframing AI hallucination as a signal problem, not an accuracy problem.
9:41
ChatGPT 5
Ask anything
I want to create a weekly diner schedule for me and my husband.
Perfect! I can create a weekly dinner plan and even prep a shopping basket for you. Can you tell me your dietary preferences, and your partner’s too?
Camera
Photos
Files
Modes
Precision
Standard
Creative

9:41
ChatGPT 5
Ask anything
I want to create a weekly diner schedule for me and my husband.
Perfect! I can create a weekly dinner plan and even prep a shopping basket for you. Can you tell me your dietary preferences, and your partner’s too?
Factual Error
Contains incorrect information.
Logic Error
Math or reasoning is flawed.
Incorrect Source
Citation is broken or fake.
Misunderstood Request
Failed to follow instructions.

Timeline
3 weeks
Team
Academic project
My Role
UX Research
UI Design
Interaction Prototyping
Tools
Figma
Miro
Whimsical
Google Forms
hypothesis i’m exploring
Users don't distrust ChatGPT because it's wrong. They distrust it because right and wrong sound identical.
01 I context
What is this, and why now?
33 survey respondents. 3 interviews.
Millions of people use ChatGPT for research, coursework and decisions that matter. However, the interface presents verified facts and invented ones with the exact same tone, formatting and certainty.
Not broken.
Just blind.
The model isn’t failing at a higher rate than the users expect it to. It is failing silently, with no visible difference between a grounded answer and a guessed one.
02 I problem
Confident is not the same as correct.
the assumption
If the model gets smarter, trust problem solves itself.
Accuracy and trust are not same. A model that's right 99% of the time still needs to signal the 1% it isn't.
the current resposne
Identical tone and formatting, regardless of accuracy.
76% of users report that correct and incorrect answers arrive with no distinguishing signal at all.
the two gaps
Silence. The product says nothing when it should speak.
No signal. Nothing in the interface tells a user when to be skeptical. Zero control.
03 I strategy
Why Control, not Accuracy?
Accuracy is a model problem. Control is a design problem.
Reframe: instead of chasing "make the model always right," the design targets user control over strictness and visible accountability after the fact.
why the reframe matters?
Treating this as an accuracy problem leads nowhere a designer can act. Treating it as a control and feedback problem produces two concrete, shippable moves: let users set expectations before they type, and let them correct the system in one tap after.
04 I user journey map
What the user thinks, feels, and does at each step
ask
read
doubt
with fix
doing
Opens ChatGPT, types question, submits.
Reads the response, tries to assess accuracy.
Opens Google in a new tab. Re-prompts if it still feels off.
Reads inline confidence. Flags in one tap if wrong.
Never leaves the product.
emotion
High intent
Cautious optimism
Anxiety spike
The Gap.
Product goes silent. User has to abandon the app.
Set the mode before asking. Flag the answer after. Never break the flow to check.
Trust maintained
thinking
This should be quick.
This sounds right... but is it?
I can't just trust this blindly.
This is marked Precision so it's already cited.
feeling
Hopeful. Intent is clear.
Uncertain. Judging tone, not fact.
Anxious. No signal to tell grounded from guessed.
Reassured. Confidence is visible, not assumed.
05 I solutions explored
Three directions. One chosen.
01
Confidence Heatmap
Highlight text by confidence level
It was cut as it created visual clutter and unfairly penalized creative requests.
02
Split-Screen Auto Search
Live search results beside the chat
It was cut as it kills the chat space and doesn't scale to mobile.
03
Contextual Modes + Granular Feedback
Pick mode before typing, expectation set beforehand
Solves the brainstorming vs research tension without adding clutter.
the new model
Contextual Modes + Granular Feedback
Contextual Modes wins because it acts before the problem exists, not after. Users define trust needs upfront, so the system adjusts quietly with zero clutter for tasks that don't need scrutiny.
Paired with a one-tap flag for specific errors, it also lets users correct the system in the moment, not just prevent issues in advance.
06 I ui screens
Every screen. Every desicion.

MODE SELECTOR
Lives inside the input bar, not a settings menu. Switching context takes one tap, before a single word is typed.

GRANULAR FLAG
Four categories, not a generic thumbs-down. Generic downvote tells the system nothing. A category tells it exactly what to fix.
07 I impact
What testing with users showed
68%
Faster task completion
Most participants no longer needed to leave the chat to recheck results, resulting in faster completion times.
83%
Product retention
5 of 6 tested participants completed the task without ever leaving ChatGPT to verify.
Trust
Reduced anxiety
Participants stopped treating every answer with the same suspicion. Reserved scrutiny for sensitive situations.
08 I next steps
Two moves. In this order.
Soon
Longitudinal testing on the granular feedback loop.
Later
Refine the micro-interactions to handoff a build that feels considered and deliberate.
09 I what i learnt
Not just about Generative AI. About designing around a system that can't be fully trusted.
on trust
Skepticism isn't a complaint. It's a feature request.
Users didn't want a flawless model. They wanted to know which answers to double check. That distinction was the entire project.
on framing
The interface can’t fix the model. It can fix the moment after.
Hallucinations can’t be fixed at the interface level. How users respond to uncertainty can be fixed. That shift made control the real lever, not accuracy.
on control
Letting users set strictness upfront beat every reactive warning tested.
Precision, standard and creative modes solved the brainstorming vs research conflict no static label could. Proactive control beat reactive correction.
on noise
Transparency that looks like accusation stops being transparency.
Highlighting every uncertain claim overloaded users and punished creative prompts. A warning should be proportional, not maximal.
Glad we could cross paths!
I hope it left you with a bit of curiosity and inspiration.
SELECTED WORK
© 2026 Arunima Guin. All rights reserved.