Most EHR selections rely on vendor demos, and most demos are performed by a specialist who has clicked through the same workflow hundreds of times. The result looks effortless. Then the practice goes live and discovers that a medication reconciliation takes fourteen clicks and the referral screen hides the field everyone needs. Usability testing during selection replaces the demo's polished narrative with a simple question: how long does it take our people to do our tasks in this system, and how many mistakes do they make along the way?
Why demos are not usability tests
A demo is controlled by the vendor. The presenter chooses the patient, the order of tasks, and the path through the software. A usability test is controlled by the practice. The practice chooses the tasks, the people performing them, and the criteria for success. The difference matters because usability problems tend to hide in the tasks a vendor would not choose to show: the exception paths, the interruptions, the patient with twenty medications and three insurance plans.
Usability also has a direct safety dimension. Poorly designed order entry screens, ambiguous alerts, and confusing patient banners have all been implicated in documented patient harm. Federal certification requires vendors to apply user-centered design to specific safety-related capabilities and to publish the results, which gives buyers a starting point but not a substitute for their own testing.
What certified vendors already publish
ONC's certification criteria include a safety-enhanced design requirement. Vendors must apply a user-centered design process to a defined list of capabilities, including medication ordering, medication allergy list, drug-drug and drug-allergy interaction checks, clinical decision support, and electronic prescribing, and must submit summative usability test reports describing the participants, tasks, and results. Those reports are public through the Certified Health IT Product List.
Read each finalist's report before designing your own test. Look at who the participants were (clinicians in your specialty, or the vendor's own staff), how many tasks were tested, what the task completion rates and error rates were, and what the report identifies as areas for improvement. A vendor whose report shows low completion rates on a task central to your practice has told you where to look first.
The published reports test the vendor's chosen configuration with the vendor's recruited participants. Your test will use your configuration priorities and your staff. Both are informative; only yours reflects how the system will feel on a Tuesday afternoon in your clinic.
Writing scripted scenarios
A scenario is a short, realistic description of a clinical or administrative situation that requires the tester to complete several tasks in sequence. Good scenarios come from the practice's own high-volume and high-risk workflows. A primary care practice might use:
- A new patient with a complex medication list arrives; reconcile the list, document an allergy, and prescribe a new medication that triggers an interaction alert.
- A returning patient with diabetes needs a visit note, a lab order set, a referral to a specialist, and a portal message with instructions.
- A front-desk user checks in a patient whose insurance has changed, collects a copay, and reschedules a follow-up.
- A nurse works the results in-basket, routes an abnormal result to the physician, and documents a phone call to the patient.
- A clinician receives an outside record through the health information exchange and incorporates a problem and an immunization into the chart.
For each scenario, write the tasks as the practice would describe them, not as the vendor names its screens. "Find out whether this patient had a colonoscopy in the past ten years" tests the system's information design; "open the Health Maintenance tab" does not. Define what counts as successful completion before the session, and decide which errors are critical, such as ordering the wrong medication, and which are minor.
Running the sessions
Ask each finalist vendor for a sandbox loaded with the same test patients and a configuration that approximates what the practice would use. Recruit four to six testers per role from the practice's own staff: clinicians, nurses or medical assistants, and front-desk or billing users. Give each tester a brief orientation of no more than thirty minutes; the point is to measure learnability as well as efficiency, so do not train them to mastery.
During the session, a facilitator reads the scenario, starts a timer, and observes. A second observer counts clicks or screens, notes errors, and records where the tester hesitates or asks for help. Testers should think aloud so observers can hear where the interface confuses them. After each scenario, ask the tester to rate difficulty on a simple scale and to name the single most frustrating step. Keep the vendor out of the room during testing; a vendor representative can answer questions afterward.
Run the same scenarios with the same testers on each finalist, ideally on different days to limit fatigue, and randomize the order in which testers see the systems so that familiarity with the scenarios does not favor the last vendor tested.
Scoring and comparing vendors
| Measure | How to capture | Interpretation |
|---|---|---|
| Task completion rate | Percent of tasks completed without assistance | Below roughly 80 percent on a core task is a serious concern |
| Time on task | Seconds from start to completion | Compare across vendors for the same task; absolute times matter less |
| Errors | Count, split into critical and minor | Any critical error on a safety-related task deserves follow-up with the vendor |
| Assists | Number of times the facilitator had to help | Indicates learnability and likely training burden |
| Perceived difficulty | Tester rating after each scenario | Captures frustration that timing misses |
| Standardized satisfaction score | A short post-session questionnaire such as the System Usability Scale | Allows comparison to published benchmarks |
Aggregate results by role and by scenario. A system that scores well for clinicians and poorly for the front desk is a real finding, because front-desk throughput drives the schedule. Weight the scenarios according to volume and risk, and present the comparison as a table with the raw measures visible, not just a composite score, so the selection committee can see where the differences come from.
Feeding results into the decision
Usability results should sit alongside functionality, cost, interoperability, vendor stability, and references in the selection scorecard, with a weight that reflects how much of the practice's day is spent inside the system. For most ambulatory practices that weight is substantial. Share the findings with the vendors and ask how they would address specific problems; some issues are configuration choices that can be fixed before go-live, and the vendor's response tells you something about its support culture.
Finally, keep the scenarios and the results. They become the acceptance tests for the implementation and the baseline for measuring whether optimization work after go-live actually made the system faster.
Common questions
How many testers do we need for a useful usability test?
Usability research consistently finds that a small number of testers per role, often four to six, surfaces most significant problems. The goal during selection is to compare vendors on the same tasks, not to produce statistically precise measurements, so consistency of testers and scenarios across vendors matters more than sample size.
Where can we find a vendor's published usability test results?
Certified health IT vendors must submit safety-enhanced design reports as part of certification. These are available through the ONC Certified Health IT Product List, where each product listing links to its usability report and other certification documentation.
Will vendors provide a sandbox for testing?
Most vendors will provide a test environment for finalists, though it may take a few weeks to arrange and may not include every module. Ask early in the process and request that the sandbox be loaded with the same test patients for each vendor so results are comparable.
What if our results conflict with the vendor's published report?
That is common and useful. The vendor's report reflects its recruited participants and its preferred configuration. Your results reflect your staff and your priorities. Share the discrepancy with the vendor and ask whether configuration or training changes would close the gap before you decide.