Scroll to explore

Scroll to explore

Button text

Scroll to explore

Scroll to explore

Course

Course

Human-AI Interaction, UT Austin

Human-AI Interaction, UT Austin

Duration

Duration

11 Weeks

11 Weeks

My Role

My Role

UX Research & Interface Design

Team

Team

3 students

Date

Date

Feb 2026 - Apr 2026

Status

Status

Prototype & study

AI Interaction

Design Research

Prototype

Everyone is building for speed. No one is building for thinking.

Everyone is building for speed. No one is building for thinking.

The problem

The problem

So I built an AI ideation tool that gets in your way on purpose, in two different ways, then ran a study to find out what that does to a person’s thinking.

Generative Labour

Antagonistic Partner

Standard

One variable

3

3

modes, one toggle

What I did

What I did

Designed and built a text ideation tool with two kinds of friction inside it, then ran a study to see whether friction changes how people think and what they feel they own.

Every person did every mode

5 participants

5 participants

5 participants

The outcome

The outcome

Both friction modes moved thinking from mostly fast to an even split, and people felt more ownership of the ideas they left with. Satisfaction did not drop. In the plain chatbot, four of five people quit before the tool did.

Fast thinking fell from 79% to 42%

Ownership rose in both friction modes

Satisfaction and confidence did not move

22 of 35 baseline turns went unused

The problem

Ten tools, one shared assumption

We audited what people actually reach for when they make things. Every one of them sells the same promise: fewer steps between wanting something and having it.

When that promise is right

Book the flight. Resize the image. Write the email you have already composed in your head. When you already know what you want, the faster the tool gets out of the way, the better it is.

See the study

Scroll to explore

Button text

Interfaces: Opal, Figma Make, Lovable, Magicpath

Campaigns: Pomelli, Flow, Canva, Jasper. Moodboards: Mixboard, Firefly.

When it is exactly wrong

An idea is not like that. You do not know what you want yet. Working it out is the task, and the thinking is what does the working out. Remove the effort and you have not saved time, you have removed the only part that was getting you anywhere.

Oct 2023

Discover

Old app was a booking page. Research with doctors, staff and founders; 5 stakeholders mapped.

Nov 2023

Define

Three bets: rebuild, two generations, operable. Business and user goals split.

Jan 2024

Develop

First design system, then both apps end to end. Accessibility to 9/9 WCAG AA.

Feb 2024

Deliver

Shipped iOS & Android. Design QA became a real step. Now live in US clinics.

Why it happens

A polished answer never trips the wire

Dual-process theory splits cognition in two. System 1 is fast, associative, heuristic and effortless. System 2 is slow, procedural, metacognitive and effortful. Creative work needs both, because System 1 generates and System 2 evaluates.

The System 1 to System 2 transition is not voluntary. System 2 engages when System 1 hits an error it cannot resolve, so confusion is the trigger. AI-assisted ideation is documented as System 1 dominant, and this is why: fluent output produces no error to catch.

“The fluency trap: output that is syntactically perfect and conceptually shallow. Nothing looks wrong, so nothing triggers System 2.”

The three modes, mid-session. Standard on the left, Antagonistic Partner in the middle, Generative Labour on the right.

The baseline: a normal AI chatbot

What that costs

What that costs

The interface is doing the skipping, which makes the interface the place to fix it. Three things go wrong at once when a tool hands you a finished answer.

01
Nothing to push against

People take what the model hands them and move on. The generative half of the work quietly moves to the machine.

02
Agreement arrives too early

Person and model converge within a turn or two. There is no disagreement left to explore.

03
You do not own what you did not make

An idea that arrives finished is easy to abandon. Effort is what makes it yours.

Research

Seven kinds of friction exist. We could build two.

Deliberate friction has a literature. Under desirable difficulties in educational psychology, and positive friction and frictional AI in interaction design, we found seven families of it.

7

7

families of cognitive friction in the literature.

The menu.

2

2

mechanisms we could build and test properly. The scope.

Read the filter

Two questions cut it down. Can three students build it in one semester, and does it change what the person has to do rather than just how the screen looks.

The five we passed on

  • Perceptual disfluency

    education · cognitive psychology

    What it is

    Make the material itself harder to process.

    Looks like

    Degraded fonts, hard-to-read typesetting, disfluent layout

    Evidence

    Diemand-Yauman et al. 2011; Alter et al. 2007

    Why we passed

    Changes how the screen looks, not what the person has to do

  • Temporal barriers

    UX · safety-critical systems

    What it is

    Put time between intent and output.

    Looks like

    Microboundaries, artificial processing delays, task lockouts

    Evidence

    Mejtoft et al. 2019; Cox et al. 2016

    Why we passed

    Makes people wait, which is not the same as making them think

  • Cognitive forcing functions

    AI-assisted decision-making

    What it is

    Withhold the AI’s answer until the person commits to their own.

    Looks like

    Forced choice, unassisted first steps, on-demand suggestions

    Evidence

    Buçinca, Malaya and Gajos 2021

    Why we passed

    Strong evidence, but built for decisions with a right answer

  • Productive struggle

    pedagogy · tutoring systems

    What it is

    Give hints and guiding questions instead of solutions.

    Looks like

    Scaffolding, AI coaching, staged reveals

    Evidence

    Wang and Srivastava 2025; Jun and Wang 2025

    Why we passed

    Aimed at skill development rather than at making something

  • External scrutiny

    social media · AI development

    What it is

    Make someone or something else audit the reasoning.

    Looks like

    Reflective nudges, read-before-you-post, explainability requirements

    Evidence

    Chen and Schmidt 2024; Buçinca et al. 2021

    Why we passed

    Needs a second party, which a solo ideation session does not have

Buildable in one semester

Three students, eleven weeks, a model running on a laptop. Anything needing a second human or a trained classifier was out.

Interactional, not decorative

It had to change what the person does inside a turn. Perceptual disfluency and temporal barriers fail this test.

Untested on creative work

Generative labour has never been applied to ideation, and antagonistic style has never been evaluated for it. That gap is the project.

The two we built:

Mechanism 1

Generative labour

Force effort and input from the person, on the theory that labour creates ownership. The IKEA effect, applied to ideas.

Effortful assembly

Anchors and connections

Norton, Mochon and Ariely 2012

Never tested on ideation before

Mechanism 2

Agonistic design

Make the system disagree. Engineered discomfort, an unfriendly counterpart, a position you have to defend rather than accept.

Unfriendly by design

Ritualistic subversion

Collins et al. 2025

Never evaluated for creative work

The response

So we built the same tool three times, and changed one thing

A text ideation tool for marketing briefs, with three modes behind a single toggle. The model, the task and the length of the session are identical in all three. The only thing that changes is how much work the interface makes you do, which is the only way to find out whether the work is what matters. Standard is a normal AI chatbot: it hands you finished ideas and you pick one. Generative Labour puts rough fragments on a canvas and stops, and will not advance until you have edited them yourself. Antagonistic Partner is helpful for two turns, then turns critic and pushes on your weakest answer.

Three modes, one toggle

Family App

Family books care. Member, service and time slot

Operations

FHM assigned. A real Health Manager with a face, not a ticket

At the Door

Visit verified by a one time code. Proof for the patient, proof for the worker.

Care Team App

Vitals captured. BP, SpO2, heart rate recorded on the spot

Records

Family and Clinic. Structured records flow to the app and the medical record

Reminders, medication schedules and alerts keep the loop running between visits

Service Loop

One booking, two apps, a verified loop of care

Holding it still

To learn anything, only one thing was allowed to differ

The tool was never the point. It is a measuring instrument, and an instrument only tells you anything if everything except the one thing you are testing stays the same. Same model on a laptop, same brief format, same turn cap, same interface shell. What differs: Standard shows chat only and asks nothing of you; Generative Labour adds a node canvas and will not move until you build on it; Antagonistic Partner stays chat but makes you defend the idea you picked. Participants were never told which mode they were in, or that it changed between sessions.

Teal acts,

Cream informs,

Red interrupts

Card 1
Card 2
Card 3
Card 4
Card 5
Card 6
Card 7
Card 8
Card 9
Card 10
  • Poppins

  • Poppins

  • Poppins

  • Inter

  • Inter

  • Inter

  • Inter

Two voices of type

Words people already know

Button

Styles

Radius 16px

Action button

Action button

Action button

Action button

disabled

Default

disabled

Default

Icons

?

What was held constant

UX Research

Wireframing

UI Design

Mode one

Generative Labour

The model opens the space. You make the meaning.

Chat on the left, a node canvas on the right. Every turn the model puts something abstract on the canvas, and the session will not move until you have worked on it. Edit the nodes, add detail, draw the connections yourself. Six turns take you from a one line brief to a campaign with a mechanism: Clarify, Vibe, Medium, Seeds, Mechanism, Finalise.

Turn 4: seed ideas on the canvas

Turn 5: three empty fields

One turn does the real work

Turn five is the only turn where the model writes nothing. Three empty fields appear: what will make the audience act, what will they do, and what happens as a result. Every turn before it was choosing. This one is writing, and it is where people stalled.

The first canvas was open. You connected any node to any other and wrote a sentence explaining the link. The hard part was making that effort feel meaningful rather than tedious, and an open canvas could not promise that. Six fixed turns give the work a direction, so every action builds on the last.

The catch: the session is rigid. A confident user walks the same six steps as everyone else.

The prompt is the interface

The canvas is the part you see. The part that decides how a turn behaves is a prompt. Each of the six turns carries its own instruction set, and every reply comes back as strict JSON against a fixed contract, so the interface can place nodes without guessing at them.

Canvas state, every message

The canvas is fed back to the model on every single message. It reads what is already there, returns only what it is adding, and never re-lists the rest.

The model may not touch what you made

Every node a user creates carries an id the model is forbidden to modify, enforced twice: once in the prompt, once in the code that handles the reply. A study about psychological ownership is worthless if the model can quietly edit the thing the person made.

The catch: the model sometimes refers to user work clumsily, because it cannot tidy it. Correct trade for this project.

Six turns of accumulated work survive intact, which is the entire promise of the mode.

Friction needs an exit

We built resistance, then had to build relief. Too many escape routes and the friction stops working. Too few and people abandon the session. Four exits, each scoped to one problem: Undo on the canvas, Regenerate on turns one to four only, Go Back to restore the turn four snapshot, and a Stuck button that exists on turn five and nowhere else.

One turn was hard enough to need a hint button. It got exactly one, and no other turn has it.

Where the exits sit

Regenerate stops before the commitment turn, so you can ask for different options while ideas are still cheap but not after you have chosen. Hints exist only on the turn people actually stalled on.

Undo · Regenerate (turns 1-4) · Go Back (to turn 4) · Stuck (turn 5 only)

The cost of relief

A person stuck for a different reason has no way out but finishing.

Every exit is a hole in the mechanism, so each one had to earn its place.

Mode two

It is friendly for two turns, then it turns

For the first two turns it is a neutral evaluator. It clarifies your brief, hands you three or four named campaign ideas, and asks which one excites you. Warm, useful, ordinary. The moment you choose, the interface swaps the prompt underneath you. From turn three it is a critic: it acknowledges your idea without praising it, asks questions built to find the weak joint, and never fully accepts an answer. Each turn it pushes harder.

The decision

Let people commit before you attack

Hostility at turn one has nothing to bite. You have not committed to anything yet, so pressure reads as noise and people disengage. The persona switch waits until after the choice, so every challenge lands on something the person has already claimed as theirs.

What we cut: the shape-shifter

The plan had the model role-playing a different stakeholder each turn, a marketer then a customer then a sceptic. That was cut for one critic who escalates, because a rotating cast gives you three shallow objections instead of one that gets deeper.

The catch: those first two friendly turns are a small ambush. One pilot participant said it made them feel they were defending someone else’s idea, and that showed up in the ownership scores months later.

Method

Five people, fifteen briefs, nothing repeated

Within-subjects, counterbalanced. Every participant did all three modes, in a different order, on a different brief each time. Fifteen consumer marketing briefs matched for difficulty, so nobody saw the same problem twice and no mode ever drew the easy question. Sessions were screen and audio recorded, with a timestamp logged at every turn.

01

Counterbalanced Latin square

Five orders across five people, so the effect of going first or last cancels out. The catch: each order runs exactly once, with no redundancy if a session fails.

02

All three capped at six turns

Generative Labour has a fixed length. If the other modes could run longer, session length and mode would be tangled together and neither could be read. The catch: the baseline was cut off at a length it would never have chosen.

03

Retrospective think-aloud, not a test

We dropped the planned Cognitive Reflection Test and replayed each recorded session turn by turn instead, asking people to narrate what they were thinking at that moment. The catch: coding transcripts takes far longer, and needs two people to be credible.

The study shrank, on purpose

The first plan had fifteen participants each doing one mode, between subjects. We rebuilt it as five doing all three. Fewer people, but every person becomes their own control, which is worth more than the extra ten at this sample size.

What we measured

Two research questions, four self-report measures

One: does generative labour or an antagonistic conversation style trigger System 1 to System 2 transitions during ideation. Two: do they affect psychological ownership of the resulting idea. After each task, participants rated psychological ownership, perceived AI contribution, satisfaction with output and confidence in the idea, each on a 1 to 7 scale. Everyone also completed the Metacognitive Awareness Inventory as a baseline.

Nobody was told which mode they were in, or that the mode changed between sessions.

Fifteen briefs, no repeats

Every brief was a consumer marketing problem with no specialist knowledge required: a clothing brand, a music festival, a coffee shop, a smartphone launch, a tutoring platform. Matching them for difficulty is what stops an easy question flattering one mode.

Analysis

How “I just did it” becomes data

We wrote the codebook before anyone looked at a transcript. Three top-level codes and twelve beneath them, built from dual-process theory rather than from whatever we happened to notice. System 1 covers rapid generation, association, gut judgement, heuristics and automaticity. System 2 covers step-by-step reasoning, effortful deliberation, evaluation, rule-based reasoning and handling novelty. A third code, cross-cutting, catches the moment a person starts on instinct and then catches themselves.

κ 0.70 on System 1 versus System 2. Substantial agreement.

κ 0.62 on the twelve sub-codes. Weaker, and expected.

Two coders, independently

Inter-rater reliability

Reliability was calculated as Cohen’s κ. Coders agreed strongly on whether System 2 was engaged and less strongly on which sub-process it was, which is the expected pattern for theory-driven qualitative coding. It is easier to tell that someone is deliberating than to name what kind of deliberating they are doing. Cross-cutting moments appeared only in the friction modes, never in the baseline.

Reliability was calculated as Cohen’s κ. Coders agreed strongly on whether System 2 was engaged and less strongly on which sub-process it was, which is the expected pattern for theory-driven qualitative coding. It is easier to tell that someone is deliberating than to name what kind of deliberating they are doing. Cross-cutting moments appeared only in the friction modes, never in the baseline.

κ

0.70

top level

Onboarding

Onboarding

Onboarding

Home

Home

Home

Bookings

Bookings

Bookings

Community

Community

Community

Account

Account

Account

Reminders

Reminders

Reminders

Records

Records

Records

Add a parent as a Family Health Member in minutes, one profile the whole family shares.

Results

Both mechanisms changed how people thought. Only one changed how they felt.

System 1 coding fell from 79% in the baseline to 42% in Generative Labour and 33% in Antagonistic. The second coder read lower throughout and found the same pattern.

Turns used: 22 of 35 in the baseline, 35 of 35 in both friction modes. Four of five people quit before the tool did, and two after only three turns.

Asked why they stopped early, participants said they had “no more to discuss”. That is premature convergence, the exact failure this project was built to interrupt.

Psychological ownership rose from 3.4 to 5.2 in Generative Labour. In Antagonistic it moved to 3.6, which on five people is nothing.

Antagonistic cost 1.4 points of satisfaction and 0.6 of confidence against the baseline. It made people think harder and left them feeling worse about the result.

So “friction works” is too simple. System 2 engagement and psychological ownership are separate outcomes, and only one of the two mechanisms moved both.

The finding underneath the finding

One keeps you in the room. The other slows you down inside it.

Both mechanisms roughly doubled engagement, and the summary numbers make them look like two versions of the same idea. The timing logs say otherwise. Mean time per turn in Generative Labour was 1:42. In the baseline it was 1:37. Five seconds apart. Generative Labour never made a single turn slower, and every extra minute came from turns a baseline user would simply have skipped. Antagonistic is the opposite: the same number of turns, each one half again as long at 2:26, with the extra time spent inside the exchange, arguing. Generative Labour was fixed at six turns, so people could not leave early. What is interesting is that being held in place did not change their pace. They worked at a normal speed, for longer. Our own report files this question under future work. The answer was already sitting in the timing logs, unadded.

What we learned

Friction charges a toll, and not everyone has the same currency

Someone always pays

One participant scored 26 of 52 on the Metacognitive Awareness Inventory against 34, 38, 43 and 48. They disliked both friction modes and rated the plain chatbot highest. The deficit sits in planning and conditional knowledge, the subscales that operate before a task starts, which is exactly what both friction modes demand. Our own report lists this under future work.

Effort has to feel like it is going somewhere

The hardest design problem was not adding work, it was making the work feel meaningful rather than tedious. An open canvas made the labour feel arbitrary. Six turns walking abstract to concrete fixed it.

Cutting the feature protected the finding

Giving the antagonist a canvas would have made a better product and a worthless experiment, because no result could then say which mechanism did the work. Knowing which of the two you are building is most of the job.

A null result is still a result. The antagonist made people think harder and gave them no more psychological ownership than a plain chatbot, while costing them satisfaction. That is the most useful thing the study produced, and it is the part we nearly rounded away in the write-up.

What I’d do next: test whether the two mechanisms compound when combined or cancel; judge the ideas themselves blind, since we measured how people thought and what they felt they owned but never whether the output was any good; and let a person tune how much friction they get, instead of assuming everyone arrives able to pay for it. Five participants, one model, text ideation only. None of these differences would survive a significance test, and we ran none.