Test App Usability: With Real Users on Real Devices

Test App Usability by watching real people attempt real tasks on a real device, then recording where they hesitate, fail, or succeed.

The purpose is not to confirm that features exist. It is to find the moments where a person stops, guesses, or gives up, and to decide which of those moments matter enough to fix before release. A usability test produces observations, not opinions, and the observations only hold up if the tasks were written before the session and the sessions were run the same way for every participant.

Most teams can run this without a research department. The work is mostly preparation. deciding what to learn, writing tasks that force a decision, recruiting a handful of people who resemble the target user, and keeping a written record. The sections below cover how to test app usability from planning through prioritisation, including the trade-offs between moderated and unmoderated sessions and the constraints that shape what a small team can realistically run.

How to test app usability before a release

Testing before release works best when the build is stable enough to complete a task but unfinished enough that changes are still cheap. A clickable prototype, a staging build, or a feature-flagged release candidate all qualify. A build that crashes mid-task produces noise rather than insight, because participants spend the session recovering from failure instead of revealing how they think.

Timing matters more than polish. Testing a rough prototype surfaces structural problems — wrong labels, missing confirmation, confusing navigation — while those problems are still a design conversation. Testing a finished build surfaces the same problems as engineering tickets, which cost more to resolve.

The practical sequence for a single round looks like this:

  1. Write down the two or three decisions the round must inform, and the questions that would change those decisions.
  2. Choose the build. prototype, staging, or release candidate, and confirm it can complete every task end to end.
  3. Write each task as a goal the participant pursues, not a set of instructions to follow.
  4. Recruit participants who match the people the app is built for, and screen out anyone who has already seen the design.
  5. Run a pilot session with a colleague to catch broken tasks, unclear wording, and recording failures.
  6. Run the sessions, one participant at a time, with the same task order and the same prompts.
  7. Review the recordings and notes together, then group observations by the screen or step where they occurred.
  8. Rate each issue by how badly it blocks the task and how many participants hit it, then decide what changes before the next build.

A pilot is not optional. It is the cheapest way to discover that a task cannot be completed in the build, that the recording software did not capture the screen, or that a task wording accidentally tells the participant where to tap.

What Test App Usability measures in practice

A usability session measures behaviour, not preference. The useful outputs are whether the participant completed the task, where the attempt broke down, how much help was needed, and what the participant said while working. Preference questions asked at the end are secondary, because people are poor at predicting what they will do and generous about designs they have just been shown.

Task success is the clearest signal. A participant either completes the task unaided, completes it after struggling, or fails. Recording which of those three happened, per participant, per task, gives a picture that survives internal debate far better than a general impression that the app "felt confusing".

Alongside success, note the specific friction: a misread label, a missed button, a back-and-forth between two screens, a moment of silence before a guess. These are the raw material for fixes. A session that produces only a satisfaction score produces nothing actionable.

Accessibility belongs in the same round rather than a separate one. If a participant uses a screen reader, larger text, or one-handed grip, that is a real usage condition and it belongs in the task set. The supplied evidence does not establish which conformance standard applies to a given app, so treat accessibility observations as findings to investigate rather than as a pass or fail against a named standard.

Recruiting participants and writing tasks that reveal friction

Recruit for resemblance to the target user, not for convenience. The people who will use the app have a particular context — a job, a device, a level of familiarity with similar products — and a session with someone outside that context produces findings that do not transfer. Screening questions should confirm the relevant behaviour, such as whether the person has used a comparable app in the last few months, and should exclude anyone who worked on the design or saw it in a previous round.

The supplied evidence does not define a participant count, session length, or task count for app usability testing, so treat any specific number as a decision to make deliberately rather than a rule to follow. The trade-off is straightforward: more participants surface rarer problems, while fewer participants let a small team run rounds more often. For a team without a research function, running a small round and repeating it after each fix usually produces more progress than one large round followed by months of silence.

Task wording decides what the session can reveal. A task should state a goal and stop. "Find out whether the app supports refunds" invites the participant to explore; "Request a refund for the order you placed yesterday" forces a decision and exposes the path. Avoid naming the screen, the button, or the feature, because that removes the discovery the test exists to observe.

Write tasks in the participant's language, not the product's. Internal names for features rarely match what users call them, and a task built from internal vocabulary tests vocabulary rather than usability.

Running the session and recording what happens

Run sessions one participant at a time, in the same order, with the same opening script. Consistency is what makes the observations comparable across participants; if one session gets extra hints, its results cannot be weighed against the others.

Ask the participant to think aloud while working. This produces a running commentary on expectations and confusion that a recording alone cannot supply, because a silent participant who completes a task by accident looks identical to one who understood it. Thinking aloud does slow people down and can change how they work, which is a known cost of the method rather than a flaw in the session.

Stay quiet during the task. When a participant asks whether an action is correct, deflect back to the task rather than confirming or correcting. The moment a facilitator answers, the observation is lost. If a participant is genuinely stuck and distressed, offer a neutral prompt and note that help was given, because a task completed with help is not the same result as one completed unaided.

Record the screen and, where possible, the participant's voice. Screen recording captures the taps, the scrolls, and the dead ends; the audio captures the reasoning. Note the device and operating system for each session, because behaviour on a real device differs from behaviour on a desktop preview, and a finding tied to one device may not reproduce on another.

Moderated and unmoderated sessions suit different situations. The comparison below reflects the general trade-offs of each format rather than any specific tool.

ConsiderationModerated sessionUnmoderated session
Facilitator presenceA facilitator runs the session live and can probe follow-up questionsNo facilitator; the participant works through tasks alone
Setup effortHigher, because scheduling, scripts, and live note-taking are neededLower per participant once tasks are written and the study is configured
Findings surfacedReasoning behind confusion, including why a participant expected something differentBehaviour and completion patterns across more participants
Best fitEarly rounds, complex flows, and anything where the reason for failure is unclearLater rounds, straightforward flows, and confirming a fix worked

Unmoderated sessions scale further but lose the follow-up question, which is often where the real cause appears. A participant who fails a task in an unmoderated session leaves a recording of the failure without an explanation. Moderated sessions cost more time per participant and produce fewer sessions, but they answer the question the recording raises.

Turning observations into prioritised fixes

Analysis starts by watching the sessions together, not by reading a summary. Two people reviewing the same recording will catch different moments, and the disagreement itself is useful because it flags findings that depend on interpretation.

Group observations by the screen or step where they occurred, then rate each one. A simple severity rating combines how badly the issue blocks the task with how many participants hit it. An issue that stops most participants from completing a core task outranks a cosmetic problem that one participant mentioned in passing. The rating is a judgement, not a measurement, and its value is that it forces the team to state the reasoning rather than argue from impression.

Separate findings into three buckets: things to fix now, things to investigate further, and things to accept. Some friction is inherent to the task — a confirmation step that slows people down but prevents errors is not a defect. Naming that trade-off explicitly stops the same discussion from returning in the next round.

Then close the loop. A fix that is never re-tested is a guess. The iteration cycle is what turns a round of sessions into a verified improvement: change the design, run the same tasks with new participants, and compare where the friction moved. Reusing the same tasks across rounds is what makes the comparison meaningful.

Keep the record. Notes, recordings, and severity ratings from earlier rounds explain why the current design looks the way it does, and they prevent a settled question from being reopened by someone who was not in the room.

What to prepare before the first session

Preparation is where most of the value is created, because a session cannot recover from a task that was never defined or a build that cannot complete it.

A usability test plan does not need to be long. It needs to state the decisions the round will inform, the build under test, the tasks, the participant profile, the session format, and who will run and observe. Writing it down before recruiting keeps the round from drifting toward whatever is easiest to test.

Confirm the build first. Walk through every task on the exact device and operating system that participants will use, and check that the recording setup captures what it needs to. A real device is the default; a desktop preview or an emulator can hide touch behaviour, keyboard handling, and performance that participants will notice immediately.

Prepare the session materials. the opening script, the task list in participant-facing wording, a consent note covering recording, and a place to capture notes per task. Decide in advance who facilitates and who observes, because a facilitator taking detailed notes will miss what the participant is doing.

Finally, decide what happens to the findings before the sessions start. If no one has agreed to review the results or own the fixes, the round will produce a document rather than a change. Naming the owner and the review point in advance is what makes the difference between a test that informed a release and a test that was simply run.

Blackstone Intelligence, a Kuching-based AI systems and digital growth agency operated by Blackstone Consultancy Sdn Bhd, builds websites, search systems, and AI-supported workflows for Malaysian organisations. Its published case studies include local SEO work for Sinar Saredah Sdn Bhd and Eyonic Sdn Bhd, an AI-supported e-commerce course with University Technology Sarawak, and an AI agent for student support navigation at the Students Development Services Centre UTS. Those projects show the same delivery pattern this article describes: define the problem, build a focused version, and improve it against observed results.

how to test app usability: Practical Guide