04 · What I do
Knowing which method answers the question, and when to stop.
Eight participants into a usability study, most of what you are going to find has already been found twice. Everything after that is confirmation you are paying for, and the budget it consumes is the budget that should have gone on fixing what was found.
The two failures either side of that are the ones I get called about: a study that stops at three and calls a coincidence a finding, and a study that runs to twenty because the number was in the proposal.
User research is finding out how people actually work before a product is designed for them, and usability testing is checking a specific design against real tasks: you give somebody something to do, watch where they stop, and fix that. Neither is asking people what they would like.
The number that surprises most people is how few participants it takes. Five to eight is usually enough to find the problems in a design, and the argument this page makes is that running twenty is not more rigorous, it is more expensive for the same answer.
Choosing, not defaulting
The method follows the question. It is never the other way round.
If a proposal names the method before the question, somebody is selling you their favourite technique. These are the ones I actually run, with what each is good for and what it cannot tell you.
How one actually runs
Four stages, and the second is where studies go wrong.
What decision is waiting on this, who will act on it, and what result would change their mind. If no answer would change anything, the study should not run.
Then the screener: who counts as a participant and, more importantly, who does not.
One session, thrown away. A task that is ambiguous, a prompt that leads, a build that breaks on step four: all of it surfaces here for the cost of one participant.
Skip it and you find out on session six, with five unusable sessions behind you.
Thirty to sixty minutes each. The hardest discipline is silence: helping a struggling participant destroys the finding you were there for.
Findings recorded separately from interpretations, in the moment, because the two blur within a day.
Issues clustered by cause rather than by screen, severity by consequence and frequency, and a call on each.
Ending on "further research is needed" is not a finding, it is a way of avoiding one.
The question every stakeholder asks
How many people, and what that number can honestly support.
The right answer depends entirely on whether you are looking for problems or measuring them, and conflating those two is the most common error in this field.
Qualitative usability testing. Enough to surface most severe issues, and the curve flattens after that.
What it cannot tell you: how many of your users hit this. A problem seen by four of six participants is not "67% of users", and reporting it that way is how research loses the room.
Interviews or contextual inquiry, run until new sessions stop producing new categories rather than to a number fixed in advance.
The signal is saturation, and it is worth writing down the session at which it arrived, because that is a finding in itself.
Surveys, analytics, A/B tests. This is where percentages become legitimate and where a confidence interval belongs.
Use qualitative work to find out what to measure, then quantitative work to size it. Running them the other way round produces precise answers to the wrong question.
Where this has actually happened
Including the largest single day of testing I have run.
Usability testing at organisational scale. The environment was seeded so the operator became a fictional network serving fictional customers, which let engineers test a real product without touching a real account.
Sixty usability tests in a single day, at an event of 1,200 people, and roughly 70% of usability issues were caught before they reached engineering.
Sat with underwriters while they worked rather than asking them to describe it afterwards. The tell was not anything they said.
They were exporting to spreadsheets, which is what people do when the tool cannot hold the comparison they need. No interview would have surfaced that.
Organisers were not asking for features. They were asking the same four set-up questions every single time, which is the signature of a job that should be a template.
The persona that came out of it was built from the sessions, not from a workshop, and it named her available time as the binding constraint.
What I run it with
Named, because a spec usually names some of it.
The other five