What usability testing is and when to test
Usability testing means watching real people attempt real tasks with your product and observing where they succeed, struggle, or fail. It's the single most effective method for finding usability problems — more reliable than expert review, more specific than analytics, and more honest than asking people what they think.
The core insight from decades of usability research: you don't need many participants. Jakob Nielsen's research shows that 5 users find approximately 85% of usability problems. Testing with 5 users, fixing the problems, and testing again with 5 more users produces better results than testing with 50 users once. Small, frequent tests beat large, rare ones every time.
When to test
Before you design: Test competitor products to understand baseline expectations. During design: Test prototypes at every fidelity level — paper, wireframes, interactive. Before launch: Test the production-ready product with real content. After launch: Test with analytics-informed tasks (where do users actually struggle?). The best time to test is always "earlier than you think." The worst time is never.
Test formats: moderated, unmoderated, remote, in-person
Format Comparison Framework
Use this to choose the right format for your budget, timeline, and research question.
Moderated in-person: Highest quality. You observe body language, ask follow-up questions, and catch nuances. Best for complex flows, early-stage prototypes, and sensitive topics. Cost: highest (facility, travel, time). Moderated remote: Nearly as good as in-person for screen-based products. Screen share + video call. Broader participant pool (geography doesn't matter). Best for most product usability testing. Unmoderated remote: Participants complete tasks on their own time using tools like UserTesting, Maze, or Lookback. Faster to collect data, larger sample sizes, but no follow-up questions. Best for validating specific flows, A/B comparisons, and quantitative benchmarking. Guerrilla testing: Grab 5 people at a coffee shop with a laptop. Fastest and cheapest. Best for quick gut-checks on early concepts. Not rigorous, but beats not testing at all.
Planning a usability test
Test Plan Structure Core Method
Use when: setting up any usability test.
1. Research questions: What specifically do you want to learn? "Can users complete checkout?" is better than "Is the site usable?" Write 3-5 specific questions. 2. Participant criteria: Who represents your actual users? Define 3-4 screener criteria (job role, product familiarity, demographic). 3. Tasks: 4-6 realistic tasks, ordered from easy to hard. Each task should map to a research question. 4. Metrics: What will you measure? Task success rate, time-on-task, error count, satisfaction rating. 5. Logistics: Format, tools, schedule, recording consent, compensation ($50-150 for 60-minute sessions is standard).
Writing effective test tasks
Task Writing Technique
Use when: preparing tasks for any usability test.
Good tasks are scenarios, not instructions. Bad: "Click the Account button and change your password." Good: "You've heard about a security breach and want to update your password. How would you do that?" The scenario version doesn't reveal the answer (the user might not know to look in Account), uses realistic motivation, and sounds like something a real person would actually want to do.
Rules: Don't use the same words that appear in the interface — if the button says "Account," don't use "Account" in the task. Don't reveal the path ("Go to Settings" is an instruction, not a task). Include context ("You just received an email saying..."). Start with an easier task to build confidence before harder ones.
Leading tasks invalidate results
"Use the filter panel on the left to narrow your search results" tells the user exactly where to look and what to do. You'll observe 100% success — and learn nothing. The task should describe the goal ("You want to find a hotel in Tokyo under $200/night for next weekend") and let the user figure out how to get there.
Moderating techniques
Think-Aloud Protocol Core Method
Use when: conducting any moderated usability test.
Ask participants to verbalize their thoughts as they work: "Tell me what you're thinking as you go through this." This reveals their decision process, expectations, and confusion in real time. When they go silent (which means they're concentrating or confused), prompt gently: "What are you looking at?" or "What's going through your mind right now?" Never ask "Why did you click that?" — it sounds judgmental and makes people defensive. Instead: "What were you expecting to happen?"
Staying Neutral Technique
Use when: a participant asks for help, gets frustrated, or gives opinions about the design.
When they ask "Am I doing this right?": "There's no right or wrong — I'm interested in how you'd naturally approach this." When they're stuck: Wait at least 15-20 seconds before intervening. Often they'll find a way. If they're truly stuck, give a minimal hint: "Where might you look for that?" When they criticize the design: "That's really useful feedback" (don't defend or explain). When they praise the design: "Glad to hear that" (don't celebrate — you'll bias subsequent feedback). Your job is to observe, not to guide or react.
The echo technique
When a participant says something interesting ("I expected this to be in the sidebar"), repeat the last few words as a question: "In the sidebar?" This prompts them to elaborate without leading them. It's the single most useful moderating technique — it generates deeper insights without introducing bias.
Unmoderated remote testing
Unmoderated Test Setup Core Method
Use when: you need results fast, want larger samples, or are testing straightforward task flows.
Participants receive tasks and complete them independently, with screen and audio recording. Tools like UserTesting, Maze, Hotjar, and Lookback handle the logistics. Key differences from moderated: you can't ask follow-up questions, so tasks must be completely self-explanatory. Include a post-task questionnaire (SUS, single-ease question) to capture subjective data. Review recordings for qualitative insights — don't rely solely on completion metrics.
Best for: A/B comparison of two designs, benchmark testing (measuring improvement over time), first-click testing (where do users click first?), navigation testing (can users find specific content?). Not ideal for: Early-stage concepts, complex flows with many decision points, or research questions that require probing.
Analyzing results
Rainbow Spreadsheet Core Method
Use when: synthesizing observations from moderated usability tests.
Create a spreadsheet with participants as columns and observations as rows. Color-code each cell: green (success), yellow (difficulty), red (failure), white (not applicable). Patterns emerge visually — a row of red across multiple participants is a critical issue. A column of green means that participant had no trouble. The rainbow spreadsheet makes it easy to distinguish between individual quirks (one red cell) and systemic problems (a row of red cells).
Severity Rating Core Method
Use when: prioritizing usability issues for the development team.
Rate each issue on a severity scale. Nielsen's scale: 0 — Not a usability problem. 1 — Cosmetic only, fix if time allows. 2 — Minor problem, low priority. 3 — Major problem, important to fix, high priority. 4 — Usability catastrophe, must fix before release. Severity combines frequency (how many users encountered it), impact (could they recover or were they blocked?), and persistence (one-time confusion vs. recurring struggle). A severity-3 issue affecting 4/5 users is more urgent than a severity-4 issue affecting 1/5 users.
Quantitative Metrics Technique
Use when: benchmarking or comparing designs with data.
Task success rate: Percentage of participants who completed the task. Below 78% is concerning (industry average). Time-on-task: How long the task took. Compare against expectations or previous versions. Error rate: Wrong clicks, wrong paths, backtracking. SUS (System Usability Scale): 10-question post-test survey, scored 0-100. Above 68 is above average; above 80 is good. Single Ease Question (SEQ): "How easy was this task?" on a 1-7 scale. Quick and reliable per-task metric.
Reporting findings
Usability Report Core Method
Use when: communicating results to stakeholders or the development team.
Structure: Executive summary (3-5 key findings, top-line metrics) → methodology (who, how many, what tasks) → findings by task (observations, severity, recommendations) → appendix (raw data, participant demographics, full task list). Lead with findings, not methodology — stakeholders care about what you learned, not how you set up the test. Include video clips of critical moments — 15-second clips of users struggling are more persuasive than any written description.
The highlight reel
A 5-minute video compilation of the most important moments from your test sessions is the most impactful deliverable you can produce. It brings stakeholders into the room with real users. Include moments of confusion, frustration, success, and surprise. One clip of a user saying "I have no idea what to do here" is worth twenty slides of analysis.
Templates and checklists
[What specific questions will this test answer?]
[Number, screener criteria, recruitment source, compensation]
[Scenario-based tasks, ordered easy → hard, no leading language]
[Task success rate, time-on-task, errors, SUS/SEQ, qualitative notes]
[Moderated/unmoderated, remote/in-person, recording tool, analysis method]
- Consent form and recording permission prepared
- Prototype or product tested end-to-end by the moderator
- Tasks written as scenarios (not instructions), tested for clarity
- Recording tool tested and working (screen + audio + camera if applicable)
- Note-taker assigned (moderator should not take notes during the session)
- Backup plan for technical failures (spare device, alternative link)
- Introduction script ready (purpose, think-aloud instruction, "no wrong answers")
- Post-task and post-test questionnaires prepared
- Compensation method confirmed with participants
Real-world examples
Case study
Steve Krug's "Rocket Surgery Made Easy" approach
Krug advocates for monthly usability testing as a team habit rather than a formal research project. One morning per month: recruit 3 participants, run 30-minute sessions, debrief over lunch, pick the top 3 issues to fix. No formal report. No extensive analysis. The team watches together and the findings are immediately actionable. This approach is effective because it builds usability testing into the team's rhythm rather than treating it as a special event that requires budget approval and a research plan.
Why it works: Low overhead means it actually happens. Frequent small tests catch issues early. Team observation builds empathy directly.
Case study
Government Digital Service (UK): testing with assisted digital users
GDS tests every government service with users across the full digital ability spectrum — including people who have never used a computer. Their testing protocol includes participants using screen magnifiers, screen readers, and speech recognition. Testing with these users consistently reveals problems that affect everyone: confusing labels, inconsistent navigation, unclear error messages. The extreme users surface issues that moderate users work around without noticing.
Why it works: Testing with the full ability spectrum produces more accessible AND more usable products for everyone.
Case study
Maze: unmoderated testing at scale
Product teams using Maze run unmoderated tests on Figma prototypes within hours of completing a design. The tool measures first-click accuracy, task completion paths (expected vs. actual), time-on-task, and drop-off points. A designer can test 20-50 participants overnight and have quantitative results by morning. This speed enables testing multiple design variations in the same sprint, turning usability testing from a milestone into a daily practice.
Why it works: Speed removes the excuse for not testing. Quantitative data from unmoderated tests complements qualitative insights from moderated sessions.
Common pitfalls
Testing to validate, not to learn
"We tested it and users liked it" is not a usability finding. If your test confirms every assumption and reveals no problems, either your design is perfect (unlikely) or your test was flawed — leading tasks, biased moderation, or participants who were too polite to criticize. Go in expecting to find problems. If you don't, adjust your method.
Asking users what they want instead of observing what they do
"Would you use this feature?" always gets a yes. "What would you change?" generates design-by-committee suggestions. Usability testing is about observation, not opinion. Watch what users do, note where they struggle, and draw your own conclusions about why. User behavior is reliable data; user opinions are unreliable hypotheses.
Testing too late to act on findings
Testing the finished product one week before launch produces a list of problems that can't be fixed. Test early with low-fidelity prototypes (when changes are cheap) and test often throughout development (when fixes can be prioritized alongside features). The value of testing is proportional to your ability to act on the results.