UXCon Vienna 2026 Journal

Day 1 · Wednesday 16 September

Intelligence reimagined - Interaction redefinedElizabeth Churchill · MBZUAI

“Design teams must review agent payloads with the exact same rigor as visual components.”

Intelligence originally meant "to choose between", and the talk traces how computing lost that meaning (Babbage's deterministic engine, IQ as a single number, the bicycle-for-the-mind era) before agents brought choice back into the machine. Once agents plan, run for hours and fail by semantic drift instead of crashing, fixed screens stop working. The design answer is supervisory UX: mission-control views with dry runs, live confidence, mid-flight steering and rollback. And every product now has a second user, the agent, which reads schemas instead of screens, so the tidy data layer carries "semantic debt" and needs the same design rigor as the visual one.

ai-agentsagentic-uxsemantic-debthistory-of-computing
The new makersJames Lang · YouTube

“Skills: problem framing, epistemic judgement, taste.”

The talk opened on a UX Magazine article, "The New Makers", and then showed a landscape cross-section of what makes a designer or researcher: tools on the surface, processes just below, then skills, values and personality as the deeper strata. The point of the picture is that tools and processes are the thin, changeable topsoil, while skills like problem framing, epistemic judgement and taste, and the values and personality underneath, are what actually carry the work.

researchdesigner-skills
Designing AI behavior: why prompts aren't enoughBarbara Kofler · eBay

“We design behavior (not outputs).”

The old mental model was a straight line owned by data science: requirements go into a prompt, output comes out. Kofler's replacement is a loop where the prompt is only the beginning: product requirements and experience goals feed behavior design context and data, the output goes through human review and an LLM judge working from shared evaluation criteria, and what was learned flows back into the goals. Designers own the definition of "good" (understanding, substance, clarity, usefulness, with safety as the baseline), humans calibrate the judgment, and the LLM judge applies it at scale so each evaluation result turns into a hypothesis and a change.

ai-behavior-designdesign-process
AI skeptics: Technophobes or truth-tellers?Christian Gonzalez · Google

“Lead with value, not the technology.”

Google's Gmail surveys split users into AI Innovators, AI Receptive and AI Skeptics, and the skeptics are the majority, growing, and skewing younger. Their top reason is forced presence: AI features shoved in their face with no control, which is a UX problem rather than an ideological one. The fix is to lead with value instead of the technology and to give real user agency; copy experiments that reframed "AI inbox" as "never miss another important email" and "AI overview" as "prep for this meeting" raised interest sharply, most of all among skeptics and young users.

ai-skepticismresearch
Whose English gets to be default?Michael Kibedi · Independent

“If you haven't evaluated a system on all language varieties, you can't be sure it will function as required.”

Speech technology is built around a default English, and everyone whose accent, dialect or life stage sits outside it gets worse results. Kibedi frames this through "technocreep", the slow accumulation of unseen technological relations that hide racialised and gendered histories, and through Bender and Hanna's warning about automatic speech processing in 911 call handling: a system not evaluated on all language varieties cannot be trusted to work as required.

ai-bias

Day 2 · Thursday 17 September

Design will never be the sameMeaghan Choi · Anthropic

“Your design judgement is more important than ever.”

When anyone can code, ideas are free and speed is the default, the old handoff pipeline with clean PM, design and engineering swimlanes collapses into a tangle. Choi's answer is that designers should claim the two decisions only they can make well: deciding what to build (is this a real problem for a real person, should it exist at all) and deciding when to ship (scoping, reviewing PRs, defining the launch bar, design review as a CI check). Amid the jargon, the thing that matters is taste; design judgement, meaning creative solutions, simplifying complexity, product cohesion and fit and finish, is more important than ever.

designer-skillsbuilding-with-aidesign-process
Redefining research in the AI eraMaria Rosala · Nielsen Norman Group

“The valuable stuff wasn't in the report in the first place.”

Research used to be slow and expensive, so every report was a special pebble; now teams point AI at everything and produce a cloud of decks, summaries and wikis that pollutes rather than informs. Rosala's argument is that the report was never where the value lived. Knowledge is built by the effort of doing the work (System 2, the self-generation effect) and by shared experience and stories, and AI's illusion of learning removes exactly that effort. Research therefore has to be both the artist who produces evidence and the gallery curator who decides what matters, how it is presented, how findings connect and how people experience them.

researchai-slop
The messy science of conversion rate optimizationMarcella Sullivan · Creative CX

“Moving the bar isn't sloppy. Refusing to think about it is.”

How long an experiment takes is set by three dials: traffic (nobody controls it), the size of the change (the designer controls it) and how sure you need to be (the experimenter controls it). A timid design at a 95% habit ran six weeks and came back flat; a bolder design at a deliberately chosen 90% ran four weeks and lifted conversion, revenue and account visits. The confidence bar should move with the cost of being wrong: a microcopy tweak can live with 90%, removing a payment method deserves 99%. Research before testing also matters: tests backed by research win far more often.

experimentationresearch
Bad typography kills UXOliver Schöndorfer · Pimp My Type

“Make critical actions unmistakable.”

Typography is usability. The Oscars envelope mix-up becomes a checklist: clear visual hierarchy, most important information first, never below 14 to 16 px, test under pressure with the squint test, and make critical actions unmistakable. A body typeface should be understated and get out of the reader's way; a before-and-after of a newsletter signup shows how small type changes clean up a component.

typography
Build like an architect: Advanced AI for designersBrian Greene · Stealth Startup

“Agents like blueprints too, especially when they're written in markdown.”

The double diamond used to be the designer's whole territory; now the last diamond, Deliver, is something anyone can build, so the designer's edge moves to the front. Greene borrows the architect's way of working: plan what to build and for whom, align the crew on the plan, know the site, then build in order from foundation to details. The plan lives in a markdown blueprint that carries your vision, judgment and taste plus the functionality of the thing you are building, and the first exercise is to shape that plan with the agent before building anything.

building-with-aiai-agentsdesign-process
How to create moments of delight for your usersGiles Colborne · Made Tech

“Remember that delight fades away - find new pain points!”

Delight is not decoration. It comes either from brand experiences, which suit niche audiences, or from fixing pain points, which has wide appeal and gives people a story to share. Colborne's recipe is one pattern repeated in variants: find a point of anxiety (experienced, remembered, or deliberately heightened), then resolve it effortlessly, cleverly, with a happy ending, or with a result superior to what peers get. Pick one pain point and fix it completely, measure through user tests and word of mouth, and expect the delight to fade, so keep hunting for the next pain point.

delightdesign-process

Day 1 · Wednesday 16 September · 11:00

Intelligence reimagined - Interaction redefined

Elizabeth Churchill · MBZUAI

“Design teams must review agent payloads with the exact same rigor as visual components.”

ai-agentsagentic-uxsemantic-debthistory-of-computing

Key idea

Intelligence originally meant "to choose between", and the talk traces how computing lost that meaning (Babbage's deterministic engine, IQ as a single number, the bicycle-for-the-mind era) before agents brought choice back into the machine. Once agents plan, run for hours and fail by semantic drift instead of crashing, fixed screens stop working. The design answer is supervisory UX: mission-control views with dry runs, live confidence, mid-flight steering and rollback. And every product now has a second user, the agent, which reads schemas instead of screens, so the tidy data layer carries "semantic debt" and needs the same design rigor as the visual one.

Summary

Churchill's keynote reframes intelligence as choosing, and argues that agentic products need supervisory interfaces and a designed, audited schema layer for their second user.

Slides

From Deterministic Computational Tools to Multi-Agent Ecosystems: Dual Choice Architectures, Semantic Debt, and Designing for the Second Interface.

Etymology: From Cicero to Cognition

What is "Intelligence"? – To Choose Between

The Rise of Psychometrics & IQ

Babbage & The Deterministic Engine

A Bicycle for the Human Mind

Augmentation Over Replacement

The Rise of Agentic AI

From reactive text completion models to persistent, goal-seeking agents that interact with tools and the real world.

The Exponential Curve of Agency

Line chart, rising: 2016: Scripts · 2019: NLP Models · 2022: Chatbots · 2024: Solo Agents · 2025: Swarms · 2026: Co-Intel

Autonomous operational horizons have scaled from turn-by-turn question answering to multi-day, self-correcting swarm workflows.

The exponential curve of agency, 2016 scripts to 2026 co-intelligence
The exponential curve of agency, 2016 scripts to 2026 co-intelligence

Tool vs Agent: Architectural Shift

Classical Tool (1970–2022)

Autonomous Agent (2024+)

Anatomy of an Autonomous Agent

Screenshot of a node-based agent workflow builder; node labels [unreadable].

Screenshot of an agent workflow builder
Screenshot of an agent workflow builder

From Solo Agents to Swarms

Distributed Division of Labor

Shifting Human Cognitive Load

Horizontal bar chart:

Human cognitive contribution is fundamentally transitioning from manual task actuation to high-level strategic policy steering.

Bar chart of shifting human cognitive load across four interface generations
Bar chart of shifting human cognitive load across four interface generations

Designing for Multi-Agent UX

Why 40 years of deterministic design principles collapse in an agentic world, and how to build supervisory user experiences.

Collapse of Deterministic UI

The Fallacy of Fixed Paths

The Ghost State Dilemma

Supervisory UX: Mission Control

The Flight Director Paradigm

Core Pillars of Agent HUDs

An agent HUD visualizes an AI agent's internal state, memory, and actions in real time without chat windows.

The Second Interface

A tale of 2 "users" – one product, two choice architectures: How products now serve both human sensemakers and synthetic agents—and why the invisible layer steers decisions.

Two Users, Two Pathways

The Human User

The Agent User

Legibility & Semantic Debt

Clean kitchen sink above, tangled plumbing below: the semantic debt metaphor
Clean kitchen sink above, tangled plumbing below: the semantic debt metaphor

Asymmetric Bias & Schema Auditing

The 5 Vectors of Agent Bias

Auditing the Agent Layer

The 200 OK status code is a standard HTTP response meaning that the server successfully processed the client's request.

The Horizon of Co-Intelligence

Closing question

What do you think is the frontier of agentic architectures, dual interfaces, and artificial cognition?

Swarm Topologies · Dual Choice Architectures · Asymmetric Legibility · Supervisory UX Guardrails · Semantic Debt

Day 1 · Wednesday 16 September · 11:45

The new makers

James Lang · YouTube

“Skills: problem framing, epistemic judgement, taste.”

researchdesigner-skills

Key idea

The talk opened on a UX Magazine article, "The New Makers", and then showed a landscape cross-section of what makes a designer or researcher: tools on the surface, processes just below, then skills, values and personality as the deeper strata. The point of the picture is that tools and processes are the thin, changeable topsoil, while skills like problem framing, epistemic judgement and taste, and the values and personality underneath, are what actually carry the work.

Summary

A layered-landscape model of the designer: tools and processes on top, skills, values and personality as the bedrock that outlasts them.

Slides

Landscape cross-section diagram, sky to bedrock:

Layered landscape: tools, processes, skills, values, personality
Layered landscape: tools, processes, skills, values, personality

Day 1 · Wednesday 16 September · 14:45

Designing AI behavior: why prompts aren't enough

Barbara Kofler · eBay

“We design behavior (not outputs).”

ai-behavior-designdesign-process

Key idea

The old mental model was a straight line owned by data science: requirements go into a prompt, output comes out. Kofler's replacement is a loop where the prompt is only the beginning: product requirements and experience goals feed behavior design context and data, the output goes through human review and an LLM judge working from shared evaluation criteria, and what was learned flows back into the goals. Designers own the definition of "good" (understanding, substance, clarity, usefulness, with safety as the baseline), humans calibrate the judgment, and the LLM judge applies it at scale so each evaluation result turns into a hypothesis and a change.

Summary

Prompt tweaking is whack-a-mole; design AI behavior by defining what good means, evaluating against it at scale, and iterating on the signal.

Slides

How we shape, evaluate, and improve generative experiences beyond the prompt.

The old mental model

Data Science: PRD → Agent prompt → AI output

The old mental model: PRD, agent prompt, AI output inside a data-science box
The old mental model: PRD, agent prompt, AI output inside a data-science box

New: the prompt is only the beginning

PRD + Experience goals → Behavior design context + data → AI output → Human review → LLM judge (Shared evaluation criteria) → What we learned → back to PRD and Experience goals

New loop: shared evaluation criteria, human review, LLM judge, what we learned
New loop: shared evaluation criteria, human review, LLM judge, what we learned

Good instructions still matter

Eventually, tweaking isn't enough

Output examples lead to tweaks:

A fix for one example can change behavior elsewhere.

The evaluation data set shapes what we learn

1 example output: Looks good → 50 outputs: Patterns emerge → 1,000s of outputs: Coverage determines what we learn.

The evaluation set shapes what we learn

Same output; different judgments.

"Too long" / "Too little detail." / "Doesn't answer what I actually need."

Disagreement isn't failure. It's information.

So what does "good" mean?

Safety is a non-negotiable baseline.

Four quality dimensions: understanding, substance, clarity, usefulness
Four quality dimensions: understanding, substance, clarity, usefulness

We define good. Evaluation tests against it.

Human reviewers: Define what good means → Shared evaluation criteria: Align on how we judge it → Independent evaluation: LLM judges test the system against those criteria.

Humans define and calibrate the judgment. The LLM judge applies it at scale.

Evaluation turns observations into iteration

Evaluation result: Clarity ↓ → Look at examples: Important information is often buried. → Form a hypothesis: Stronger response hierarchy may help. → Make a change: Prompt / context / behavior → Evaluate again: Did it improve without reducing something else?

Evaluation gives us a signal. Human judgment turns it into a hypothesis.

Evaluation loop from result to hypothesis, change and re-evaluation
Evaluation loop from result to hypothesis, change and re-evaluation

Designers bridge technical systems and human needs.

We shape behavior, define what good looks like, and help turn evaluation into better experiences.

Conclusion

  1. We design behavior (not outputs).
  2. We define quality and what we expect "good" to look like.
  3. We use evaluation to learn what to change next.

Day 1 · Wednesday 16 September · 15:15

AI skeptics: Technophobes or truth-tellers?

Christian Gonzalez · Google

“Lead with value, not the technology.”

ai-skepticismresearch

Key idea

Google's Gmail surveys split users into AI Innovators, AI Receptive and AI Skeptics, and the skeptics are the majority, growing, and skewing younger. Their top reason is forced presence: AI features shoved in their face with no control, which is a UX problem rather than an ideological one. The fix is to lead with value instead of the technology and to give real user agency; copy experiments that reframed "AI inbox" as "never miss another important email" and "AI overview" as "prep for this meeting" raised interest sharply, most of all among skeptics and young users.

Summary

Skeptics are most users; they respond to value-first framing and real control, not to AI branding.

Slides

Sizing the AI User Type population on Gmail

Columns: Size (% of surveyed pop), AI Sentiment (% Positive), AI Usage (Daily+), AI Proficiency (% Deploying Agentic Systems)

Gmail in-product surveys run June - August 2026; n= 5,009 consumers and business users

Sizing the AI user type population on Gmail
Sizing the AI user type population on Gmail

Why should we care?

  1. There's a lot of skeptics. 49% American adults align with the statement: "AI presents significant risks to my quality of life through potential job loss, loss of privacy or abuse of data." Research America public opinion survey n ≈ 1000 US Adults
  2. Skepticism is increasing not decreasing. Which of the following statements comes closer to your view: A) AI has tremendous potential to bring about advancements that will improve my life. B) AI presents significant risks to my quality of life through potential job loss, loss of privacy or abuse of data. Line chart, Percentage of Respondents, 2020Q3 to 2026: Statement A falls from about 53 to 47, Statement B rises from about 34 to 54, Not sure falls from about 16 to 12.
  3. Many skeptics are younger users. Statement B % by Age bucket: 18-24 rises from about 42 to 63 between 2020Q3 and 2026, 55-69 from about 37 to 56, 35-54 from about 27 to 50, 25-34 from about 30 to 39.
  4. Skeptics actually use AI less than non-Skeptics. 4x less likely to use standalone AI. 3x less likely to use AI in product. Gmail in-product surveys run June - August 2026; n = 5,009 consumers and business users
Statement A vs Statement B over time
Statement A vs Statement B over time
Statement B share by age bucket over time
Statement B share by age bucket over time

Why are people skeptical?

Forced presence is the top driver.

So... what should we do?

Lead with Value, not the technology.

Recall our Skeptical self-id: I avoid AI features when possible and only use them if there is a proven, undeniable benefit.

Gmail AI Inbox (Beta): "Hi Christian 👋 You have 10 to-dos and 6 topics to catch up on." Suggested to-dos: Block time for GRAD Pre-Review; Review Sam's updates on Tianxin's NLA; Align on Research Integration for UX Launch Process; Show 7 more.

Gmail AI Inbox screenshot with suggested to-dos
Gmail AI Inbox screenshot with suggested to-dos

Control: Current tech-first in-product promotion for Gmail AI Inbox: "Find what matters first with AI inbox. AI automatically bubbles up critical action items, summarizes long threads, and keeps your inbox organized. Learn more"

Value-first Treatment: "Never miss another important email. Get to more of the important stuff in your inbox with auto-generated 'to-do' checklists and email summaries. Learn more". +21% increase in interest, highest among youth. Framing experiment survey with n = 1000 US Adults

Control vs value-first framing for Gmail AI Inbox
Control vs value-first framing for Gmail AI Inbox

Calendar event "Product plan sync, Tuesday, January 21 · 11:00 – 12:00 PM". Control: "AI overview 🔒 Only visible to you" = Current tech-first CTA for Calendar meeting summaries. Treatment: "Prep for this meeting 🔒 Only visible to you" = Value-first Treatment: 1.8x increase in interest and feature fit. Framing experiment survey with n = 3000 US Adults

Control vs value-first call to action for Calendar meeting summaries
Control vs value-first call to action for Calendar meeting summaries

Side note: not a new concept...

Uber app: "Where to?" versus "How would you like to leverage Cold War military technology (GPS) today?" Why Uber shows this, not this.

Uber "Where to?" compared with a technology-first phrasing
Uber "Where to?" compared with a technology-first phrasing

Respect user agency.

Next: Tailor to AI user type.

Imagine you are offered new productivity tools at work designed to help you revise documents.

Framing experiment survey with n = 3000 US, IN, BR knowledge workers

What's next? Tailoring experiences to AI user type.

AI Innovator (n=1,647), AI Receptive (n=1,591), AI Skeptic (n=670)

Preferred framing by AI user type
Preferred framing by AI user type

Recap

Day 1 · Wednesday 16 September · 16:55

Whose English gets to be default?

Michael Kibedi · Independent

“If you haven't evaluated a system on all language varieties, you can't be sure it will function as required.”

ai-bias

Key idea

Speech technology is built around a default English, and everyone whose accent, dialect or life stage sits outside it gets worse results. Kibedi frames this through "technocreep", the slow accumulation of unseen technological relations that hide racialised and gendered histories, and through Bender and Hanna's warning about automatic speech processing in 911 call handling: a system not evaluated on all language varieties cannot be trusted to work as required.

Summary

Accent bias in voice interfaces is a hidden, compounding harm; evaluate on every language variety or assume it fails.

Slides

Shadow of a Mirror (2024) by Ludovic Nkoth, the artwork that opened the talk
Shadow of a Mirror (2024) by Ludovic Nkoth, the artwork that opened the talk

"... technocreep encompasses the evidence and politics of things not seen, [so] we contend that it can serve as a feminist methodology for foregrounding racialised and gendered histories of work and exploitation, as well as of care and resistance, that gradually accumulate, but tend to be obscured in present-day technological relations." (pg. 8)

Technocreep and the Politics of Things Not Seen, edited by Neda Atanasoski and Nassim Parvin, Duke University Press, 2025

"... automatic processing of speech does not work equally well for different people. If you haven't evaluated a system on all language varieties — different regional accents, racial or ethnic accents, second language accents, life stages, disabilities — you can't be sure it will function as required." (lightly edited for readability)

In Seattle, 911 Uses "AI" to Process Your Calls, by Emily M. Bender and Alex Hanna, Mystery AI Hype Theater 3000 Newsletter, 21 June 2026

Day 2 · Thursday 17 September · 09:45

Design will never be the same

Meaghan Choi · Anthropic

“Your design judgement is more important than ever.”

designer-skillsbuilding-with-aidesign-process

Key idea

When anyone can code, ideas are free and speed is the default, the old handoff pipeline with clean PM, design and engineering swimlanes collapses into a tangle. Choi's answer is that designers should claim the two decisions only they can make well: deciding what to build (is this a real problem for a real person, should it exist at all) and deciding when to ship (scoping, reviewing PRs, defining the launch bar, design review as a CI check). Amid the jargon, the thing that matters is taste; design judgement, meaning creative solutions, simplifying complexity, product cohesion and fit and finish, is more important than ever.

Notes

How it works at Anthropic:

Summary

With building commoditised, designers own deciding what to build and when to ship; judgement and taste are the job.

Slides

65% of designers said they're taking on more product or engineering responsibilities. AI in Design 2026, Designer Fund

A wall of AI jargon (streaming · batch · model routing · model cascade · mixture of experts · distillation · quantization · open weights · small language model · multimodal · vision · voice mode · real-time · copilot · assistant · agent swarm · browser agent · computer-use agent · autonomous agent · AGI · ASI · alignment · constitutional AI · red teaming · jailbreak · prompt injection · safety layer · eval harness · benchmarks · SWE-bench · leaderboard · arena · vibe check · "just prompt it" · prompt library · no-code · low-code · citizen developer · AI-first · thin wrapper · moat · GPU-poor · compute · tokens per second · cost per token · rate limits · latency budget · "the model is the product" · "chat is not the interface" · dynamic UI · malleable software · end-user programming · personal software · disposable software · software on demand · "software is dead" · design engineer · forward-deployed · "PM is dead" · 100x · cracked · locked in · ship it · one-person unicorn · "took my job" · upskilling · reskilling · "learn to prompt" · LLM · transformer · tokens · context window · inference · latency · hallucination · grounding · retrieval · embeddings · vector database · semantic search · few-shot · in-context learning · system prompt · temperature · sampling · function calling · structured outputs · JSON mode ...) with the word "taste" repeated large down the middle.

Wall of AI jargon with the word taste repeated through it
Wall of AI jargon with the word taste repeated through it

Discover > Define > Explore > Refine > Build > Ship

Build > Ship

The big shifts

anyone can code / ideas are free / speed is the default

[decide what to build] > [decide when to ship]

Two loops: decide what to build, decide when to ship
Two loops: decide what to build, decide when to ship

Discover > Define > Explore > Refine > Build > Ship, with swimlanes: PM spans Discover to Define and returns at Ship, DES spans Define to Refine, ENG spans Refine to Ship

Old pipeline with PM, design and engineering swimlanes
Old pipeline with PM, design and engineering swimlanes

Build > Ship, with PM, DES and ENG lines tangled together

Build to ship with tangled PM, design and engineering lines
Build to ship with tangled PM, design and engineering lines

Just because you can build it, doesn't make it a good idea

should this exist? / who is it for? / is that a real person? / what problem is this solving? / does this make sense? / is that a real problem?

Audience & Use cases

Is this only applicable to you? Is this a real use case or a real need?

PM / DES handshake

PM and design handshake drawing
PM and design handshake drawing

~~Hand off and build~~ Design defines when to ship

ENG / DES handshake

Coding is not enough.

NEW: Scoping implementation / Eng timelines / Reviewing PRs / Defining launch bar / Design review as CI checks

Decide what only you can do.

[decide what to build] > [decide when to ship]

Your design judgement is more important than ever

Design judgement

Creative solutions / Simplifying complexity / Product cohesion / Fit and Finish

Design will never be the same. It's more than ever before

Day 2 · Thursday 17 September · 10:15

Redefining research in the AI era

Maria Rosala · Nielsen Norman Group

“The valuable stuff wasn't in the report in the first place.”

researchai-slop

Key idea

Research used to be slow and expensive, so every report was a special pebble; now teams point AI at everything and produce a cloud of decks, summaries and wikis that pollutes rather than informs. Rosala's argument is that the report was never where the value lived. Knowledge is built by the effort of doing the work (System 2, the self-generation effect) and by shared experience and stories, and AI's illusion of learning removes exactly that effort. Research therefore has to be both the artist who produces evidence and the gallery curator who decides what matters, how it is presented, how findings connect and how people experience them.

Summary

Research is not about producing more information; it is about building collective knowledge and helping teams apply wisdom.

Slides

"Where is the wisdom we have lost in knowledge? Where is the knowledge we have lost in information?" T.S. Eliot, Choruses from The Rock

In the old days...

Research Used to Be Slow & Expensive.

"Every research report was like a special pebble"

The research specialist

Diagram: teams as circles of grey dots, one red research team sending dotted arrows out to each team.

The research specialist feeding several teams
The research specialist feeding several teams

Today →

Teams are pointing AI at everything! Emails, Sales calls, Social media comments, Product reviews, Customer surveys, In-app feedback, Support logs, Customer contact, AI-moderated interviews, Industry trends, Analytics, Digital twins, Synthetic users

Outputs above the arrow: Design.MD, Research summary, Presentation deck, HTML dashboard, Report, UX.MD, Wiki

Information Pollution

Information pollution: inputs feeding a cloud of AI-made outputs
Information pollution: inputs feeding a cloud of AI-made outputs

→ Then came AI

What once was special, is now common

Definition: AI Slop

Low quality, mass-produced digital content using generative AI. Source: The Guardian, "AI-Generated Slop is Slowly Killing the Internet"

The Era of Slop in Research

AI in UX Survey, NN/G

"The people who did the least about of work before and [sic] now churning out slop and getting praised for it. The people who actually care about what we make end up doing double work having to review and fix the slop of others AND do our own work right."

"The incentives are so obviously misaligned that people given [AI], produce slop and then go digging for AI plug-ins to make their slop look less like AI."

"We must be in an era of AI slop at work (...) and somehow leadership is letting it slide. The amount of s****y ChatGPT slides and decks I've seen where there is little effort ... [to] ensure the copy makes sense is baffling to me."

Jorge Luis Borges, The Library of Babel: "When it was announced that the Library contained all books, the first reaction was unbounded joy. All men felt themselves the possessors of an intact and secret treasure. There was no personal problem, no world problem, whose eloquent solution did not exist- somewhere in some hexagon."

AI in UX Survey, NN/G: "AI has increased the amount of output, but not necessarily reduced the amount of effort (...) It's like we have transitioned from a canal to an ocean, and hence cast our nets wider, but we eventually only need as many fish as we need." - Researcher

The good news

The valuable stuff wasn't in the report in the first place.

Information ≠ Knowledge

Information: Data that has been organized or interpreted to convey meaning, such as a fact, statistic, statement, or observation.

Knowledge: Information that has been understood and connected into a mental model, allowing it to be recalled and applied in different situations.

Overview: The DIKW Model

  1. Data: Raw facts, metrics, or observations.
  2. Information: Data organized to answer a question.
  3. Knowledge: Patterns across information, understood well enough to act on.
  4. Wisdom: Judgment about what to do next, and why it matters.
The DIKW model: data, information, knowledge, wisdom
The DIKW model: data, information, knowledge, wisdom

AI's greatest party trick: The illusion of learning

Definition: The Illusion of Learning

A cognitive trap where a person mistakes familiarity with learning.

AI Makes Research Too Easy

System 1: Fast, unconscious, energy-efficient. System 2: Slow, deliberate, effortful. Source: Thinking Fast and Slow, Daniel Kahneman

System 1 and System 2
System 1 and System 2

Definition: Self-Generation Effect

Phenomenon where information generated is retained better than information merely read.

AI in UX Survey, NN/G: "Struggling with the data is part of a great end result, and cements that study in my head. Overall, AI just makes me feel stupider when I rely on it. It doesn't feel like I've done the work." - Researcher

Some AI Research Platforms Remove all the Work

Traditional method (crossed out): In-depth interviews (IDIs), Focus groups, Ethnographic studies, IHUTs / In-home use tests, Surveys & trackers, Shop-alongs.

Conveo alternative (ticked): 1:1 async AI interviews (Attitudes & behavior studies), Parallel async interviews (Creative & comms testing), Audio/video journaling (Usage & experiences), Async longitudinal follow-ups (Usage & experiences), Remote shopper behavior capture (Quant & hybrid research), Quant + AI-powered qual probes (Consumer behavior). Source: Conveo.ai

Conveo's table of traditional methods versus AI alternatives
Conveo's table of traditional methods versus AI alternatives

You apply mental effort by:

2 ways research builds knowledge

#2. Shared experiences and stories

A long time ago...

Our Brain Loves Stories

"We were all there. We all saw it. We all heard it. We all lived it. And those [moments were] ... always the fodder for pushing our design forward and inspiring whatever we ended up designing."

What Is Needed Today?

Redefining the Role of Research

What's Needed: Research Must Do More Than Add Information

The Artist

The Gallery Curator

Research Needs to Do Both.

The artist | The gallery curator

Research isn't about producing more information.

It's about building collective knowledge and helping teams apply wisdom.

Day 2 · Thursday 17 September · 11:55

The messy science of conversion rate optimization

Marcella Sullivan · Creative CX

“Moving the bar isn't sloppy. Refusing to think about it is.”

experimentationresearch

Key idea

How long an experiment takes is set by three dials: traffic (nobody controls it), the size of the change (the designer controls it) and how sure you need to be (the experimenter controls it). A timid design at a 95% habit ran six weeks and came back flat; a bolder design at a deliberately chosen 90% ran four weeks and lifted conversion, revenue and account visits. The confidence bar should move with the cost of being wrong: a microcopy tweak can live with 90%, removing a payment method deserves 99%. Research before testing also matters: tests backed by research win far more often.

Summary

Set the confidence bar by the cost of being wrong, and turn the dials you actually control: design boldness and required certainty.

Slides

Without research

29% are winners. Tests that have a significant positive uplift

With research

44% are winners. Tests that have a significant positive uplift

Three dials

What decides how long a test takes, and who gets to turn what.

→ Test one ran 6 weeks. Flat. 1,100,000 sessions

Three dials for test one: fixed traffic, small change, 95% certainty
Three dials for test one: fixed traffic, small change, 95% certainty

Why 95% confidence?

Second time round, we turned two of the dials.

Same traffic. A bolder design, and the dial set at 90% before we started.

→ Test two ran 4 weeks. 850,000 sessions.

90% means one time in ten I'll be wrong about a result. At 95%, one in twenty.

Three dials for test two: fixed traffic, bold change, 90% certainty
Three dials for test two: fixed traffic, bold change, 90% certainty

Test one: design one

6 weeks. Dial set at 95%. Product card: RETAIL PRICE £43.92 EACH, £36.60 EX. VAT, SAVE 20% WITH A TRADE ACCOUNT, TRADE PRICE £36.60 INC. VAT, Qty 1. Flat. No change we could see.

Test two: design two

4 weeks. Dial set at 90%. Product card: RETAIL PRICE £43.92 EACH, £36.60 EX. VAT, UNLOCK 20% OFF £36.60 INC. VAT, ACCESS TRADE PRICE, Qty 1. +3.7% Conversion, +4.5% Revenue, +46% Account page visits

Test one and test two designs with their results
Test one and test two designs with their results

How long before I could call this test?

4 weeks: 90% sure. 6 weeks: 95% sure. 8 weeks: 99% sure. Weeks to see the same 3.7% lift on the same traffic, at each setting.

Weeks needed at 90, 95 and 99 percent certainty
Weeks needed at 90, 95 and 99 percent certainty

Did it clear the bar?

How sure we are the result is real, by metric. The bar I set: 90%. The habit: 95%.

Certainty per metric against the 90% bar and the 95% habit
Certainty per metric against the 90% bar and the 95% habit

Minimum Viable Process

One site area, three bars.

Same site area, same team. The dial moves with the cost of being wrong.

Moving the bar isn't sloppy. Refusing to think about it is.

Three changes, three confidence bars
Three changes, three confidence bars

Day 2 · Thursday 17 September · 12:55

Bad typography kills UX

Oliver Schöndorfer · Pimp My Type

“Make critical actions unmistakable.”

typography

Key idea

Typography is usability. The Oscars envelope mix-up becomes a checklist: clear visual hierarchy, most important information first, never below 14 to 16 px, test under pressure with the squint test, and make critical actions unmistakable. A body typeface should be understated and get out of the reader's way; a before-and-after of a newsletter signup shows how small type changes clean up a component.

Notes

Typography directs the attention.

Summary

Hierarchy, size, contrast and unmistakable critical actions: typography failures are usability failures.

Slides

Uai is legible, durable and has a human touch

This is the body text, ideally the typeface for this long reading text is understated. Its speciality should be that it does not seem special, except to some type nerds, of course. Here content is king, not the typeface.

"Hope you can hear them. Open the door. Delirium! Sooner or later." Man, I'm listening to the new Gorillaz album all the time ... taken from 'Delirium'.

In other words a text typeface should not draw much attention to itself. It's humble and its job is to get out of the reader's way and let the words speak. Very bluntly said [unreadable]

Before / After

Newsletter card "Typography Tips. That improve UX – weekly" shown before and after the type fixes.

Before and after of the Typography Tips signup card
Before and after of the Typography Tips signup card

What can we learn from The Oscars?

Typography Survival Kit

Examples: willhaben search results, an H&M product photo, WienMobil Ticketshop ticket details.

Typography Survival Kit with app examples
Typography Survival Kit with app examples

Day 2 · Thursday 17 September · 14:55

Build like an architect: Advanced AI for designers

Brian Greene · Stealth Startup

“Agents like blueprints too, especially when they're written in markdown.”

building-with-aiai-agentsdesign-process

Key idea

The double diamond used to be the designer's whole territory; now the last diamond, Deliver, is something anyone can build, so the designer's edge moves to the front. Greene borrows the architect's way of working: plan what to build and for whom, align the crew on the plan, know the site, then build in order from foundation to details. The plan lives in a markdown blueprint that carries your vision, judgment and taste plus the functionality of the thing you are building, and the first exercise is to shape that plan with the agent before building anything.

Brian's own notes for the workshop are at briangreene.me/workshops/architecting-ai.

Notes

What I built in this workshop is the journal you are reading. Brian's method, applied: I first shaped a blueprint.md with Claude Code by answering its questions about how I capture, how it processes, and how I want to see the result, and only then let it build.

The blueprint stays the contract: every rule change goes there first, and the "log" skill is the runbook that follows it.

Summary

Work like an architect: write the blueprint with the agent first, then build in order.

Slides

We're here because the world changed.

Double diamond: Discover, Define, Develop, Deliver. "Our expertise" spans Discover to Develop; "Anyone can build" covers Deliver.

Double diamond with Deliver marked as what anyone can build
Double diamond with Deliver marked as what anyone can build

How architects work

How architects work: plan, align, know the site, build in order
How architects work: plan, align, know the site, build in order

They use blueprints.

Agents like blueprints too, especially when they're written in markdown.

``` blueprint.md

My UXCon journal

The plan for what I want to build. ```

What goes in your blueprint.md?

Your vision · Your judgment · Your taste + Your journal's functionality

What goes in your blueprint.md
What goes in your blueprint.md

Your turn: Shape your plan together.

The plan needs to cover how I capture experiences, how you process and save them, and how I see them in the journal. Ask me questions to understand and define how I want it to work. Record our choices in blueprint.md. Don't build yet.

Day 2 · Thursday 17 September · 17:50

How to create moments of delight for your users

Giles Colborne · Made Tech

“Remember that delight fades away - find new pain points!”

delightdesign-process

Key idea

Delight is not decoration. It comes either from brand experiences, which suit niche audiences, or from fixing pain points, which has wide appeal and gives people a story to share. Colborne's recipe is one pattern repeated in variants: find a point of anxiety (experienced, remembered, or deliberately heightened), then resolve it effortlessly, cleverly, with a happy ending, or with a result superior to what peers get. Pick one pain point and fix it completely, measure through user tests and word of mouth, and expect the delight to fade, so keep hunting for the next pain point.

Summary

Moments of delight are anxiety resolved well; pick one pain point, fix it completely, and go find the next.

Slides

How can we delight our customers?

anxiety → resolved effortlessly → delight

Anxiety resolved effortlessly leads to delight: easyJet plane, telephone, smiling man
Anxiety resolved effortlessly leads to delight: easyJet plane, telephone, smiling man

enhanced anxiety → ending happily → delight

remembered anxiety → resolved cleverly → delight

anxiety → superior result to your peers → delight

Why?

anxiety (highlighted) → resolved effortlessly → delight

Journey map

Search, Emotion, Conversion lines plotted over the stages Prepare, Specify, Buy, Own. Rows: Tasks, What people do, Our service, Competitor, Competitor. One cell in Specify is marked green for our service and red for both competitors.

Journey map with search, emotion and conversion curves across prepare, specify, buy, own
Journey map with search, emotion and conversion curves across prepare, specify, buy, own