Designing for trust in generative AI

Reframing AI hallucination as a signal problem, not an accuracy problem.

9:41

ChatGPT 5

Ask anything

I want to create a weekly diner schedule for me and my husband.

Perfect! I can create a weekly dinner plan and even prep a shopping basket for you. Can you tell me your dietary preferences, and your partner’s too?

Camera

Photos

Files

Modes

Precision

Standard

Creative

9:41

ChatGPT 5

Ask anything

I want to create a weekly diner schedule for me and my husband.

Perfect! I can create a weekly dinner plan and even prep a shopping basket for you. Can you tell me your dietary preferences, and your partner’s too?

Factual Error

Contains incorrect information.

Logic Error

Math or reasoning is flawed.

Incorrect Source

Citation is broken or fake.

Misunderstood Request

Failed to follow instructions.

Timeline

3 weeks

Team

Academic project

My Role

UX Research

UI Design

Interaction Prototyping

Tools

Figma

Miro

Whimsical

Google Forms

hypothesis i’m exploring

Users don't distrust ChatGPT because it's wrong. They distrust it because right and wrong sound identical.

01 I context

What is this, and why now?

33 survey respondents. 3 interviews.

Millions of people use ChatGPT for research, coursework and decisions that matter. However, the interface presents verified facts and invented ones with the exact same tone, formatting and certainty.

Not broken.

Just blind.

The model isn’t failing at a higher rate than the users expect it to. It is failing silently, with no visible difference between a grounded answer and a guessed one.

02 I problem

Confident is not the same as correct.

the assumption

If the model gets smarter, trust problem solves itself.

Accuracy and trust are not same. A model that's right 99% of the time still needs to signal the 1% it isn't.

the current resposne

Identical tone and formatting, regardless of accuracy.

76% of users report that correct and incorrect answers arrive with no distinguishing signal at all.

the two gaps

Silence. The product says nothing when it should speak.

No signal. Nothing in the interface tells a user when to be skeptical. Zero control.

03 I strategy

Why Control, not Accuracy?

Accuracy is a model problem. Control is a design problem.

Reframe: instead of chasing "make the model always right," the design targets user control over strictness and visible accountability after the fact.

why the reframe matters?

Treating this as an accuracy problem leads nowhere a designer can act. Treating it as a control and feedback problem produces two concrete, shippable moves: let users set expectations before they type, and let them correct the system in one tap after.

04 I user journey map

What the user thinks, feels, and does at each step

ask

read

doubt

with fix

doing

Opens ChatGPT, types question, submits.

Reads the response, tries to assess accuracy.

Opens Google in a new tab. Re-prompts if it still feels off.

Reads inline confidence. Flags in one tap if wrong.

Never leaves the product.

emotion

High intent

Cautious optimism

Anxiety spike

The Gap.

Product goes silent. User has to abandon the app.

Set the mode before asking. Flag the answer after. Never break the flow to check.

Trust maintained

thinking

This should be quick.

This sounds right... but is it?

I can't just trust this blindly.

This is marked Precision so it's already cited.

feeling

Hopeful. Intent is clear.

Uncertain. Judging tone, not fact.

Anxious. No signal to tell grounded from guessed.

Reassured. Confidence is visible, not assumed.

05 I solutions explored

Three directions. One chosen.

01

Confidence Heatmap

Highlight text by confidence level

It was cut as it created visual clutter and unfairly penalized creative requests.

02

Split-Screen Auto Search

Live search results beside the chat

It was cut as it kills the chat space and doesn't scale to mobile.

03

Contextual Modes + Granular Feedback

Pick mode before typing, expectation set beforehand

Solves the brainstorming vs research tension without adding clutter.

the new model

Contextual Modes + Granular Feedback

Contextual Modes wins because it acts before the problem exists, not after. Users define trust needs upfront, so the system adjusts quietly with zero clutter for tasks that don't need scrutiny.


Paired with a one-tap flag for specific errors, it also lets users correct the system in the moment, not just prevent issues in advance.

06 I ui screens

Every screen. Every desicion.

MODE SELECTOR
Lives inside the input bar, not a settings menu. Switching context takes one tap, before a single word is typed.

GRANULAR FLAG
Four categories, not a generic thumbs-down. Generic downvote tells the system nothing. A category tells it exactly what to fix.

07 I impact

What testing with users showed

68%

Faster task completion

Most participants no longer needed to leave the chat to recheck results, resulting in faster completion times.

83%

Product retention

5 of 6 tested participants completed the task without ever leaving ChatGPT to verify.

Trust

Reduced anxiety

Participants stopped treating every answer with the same suspicion. Reserved scrutiny for sensitive situations.

08 I next steps

Two moves. In this order.

Soon

Longitudinal testing on the granular feedback loop.

Later

Refine the micro-interactions to handoff a build that feels considered and deliberate.

09 I what i learnt

Not just about Generative AI. About designing around a system that can't be fully trusted.

on trust

Skepticism isn't a complaint. It's a feature request.

Users didn't want a flawless model. They wanted to know which answers to double check. That distinction was the entire project.

on framing

The interface can’t fix the model. It can fix the moment after.

Hallucinations can’t be fixed at the interface level. How users respond to uncertainty can be fixed. That shift made control the real lever, not accuracy.

on control

Letting users set strictness upfront beat every reactive warning tested.

Precision, standard and creative modes solved the brainstorming vs research conflict no static label could. Proactive control beat reactive correction.

on noise

Transparency that looks like accusation stops being transparency.

Highlighting every uncertain claim overloaded users and punished creative prompts. A warning should be proportional, not maximal.

Glad we could cross paths!

I hope it left you with a bit of curiosity and inspiration.

© 2026 Arunima Guin. All rights reserved.

Create a free website with Framer, the website builder loved by startups, designers and agencies.