EHR.ai / Clinic Management / Voice documentation

Say it once.

A doctor already describes the visit out loud — to the patient. Then types it again for the computer. This project removes the second time.

manual_note.doc
Typing the note00:00
Chief complaint only. Seven sections to go.
Transcribe visit audio
Speaking the note00:00
Listening.
Role
Product designer — research, IA, interaction, UI
Platform
Web, mobile, tablet
Users
Outpatient physicians, India
Scope
One problem: the note

Scope

One problem, taken all the way

EHR.ai is a clinic management system. It does scheduling, pharmacy, labs, billing. None of that is this case study.

This is about the ten to fourteen minutes a doctor loses per patient writing the visit down — and what happens when the note writes itself from the conversation that already happened, with the doctor reviewing rather than typing.

In scope

  • Consent capture before any microphone opens
  • Uploading or recording visit audio
  • Speaker-labelled live transcript
  • Structured note generation across eight clinical sections
  • Doctor review, section-level correction, save and print

Deliberately out of scope

  • Appointments, billing, pharmacy, lab ordering
  • Any form of diagnostic suggestion or decision support
  • Patient-facing apps or portals
  • Coding for insurance claims
  • Offline capture without a network connection

The second column matters as much as the first. Scoping this to documentation — and refusing to drift into clinical judgement — is what kept the trust model simple enough to ship.

The problem

The record competes with the patient

A physician running a twenty-minute outpatient slot has two jobs in the room and only enough attention for one. Either she types while the patient talks, and the patient watches the top of her head, or she waits and writes afterwards — and the notes stack up until they get finished at nine at night. Clinical informatics has a name for the second one: pajama time.

Neither option is a workflow problem you can schedule your way out of. The note has to be written, it has to be accurate enough to defend, and it has to print onto clinic letterhead for the patient to carry home. The only variable is who does the typing.

Where 11 minutes 20 seconds goes, per note Median across a five-day time diary, 12 physicians Re-reading the last note0:55 Chief complaint1:10 History2:40 Examination findings2:05 Diagnosis0:50 Prescription2:30 Advice and follow-up1:05 Formatting and printing1:00
The amber bar was the surprise. A full minute per patient goes on making the note print correctly on letterhead — a formatting task, not a clinical one, and one nobody had thought to design for.

4h 32m

Documentation time across a 24-patient clinic day.

61%

Of that written after the patient had already left the room.

22 of 31

Shadowed consultations where the doctor said the findings out loud before typing them.

That third number became the whole thesis. The note already exists as speech. It is being thrown away and retyped.

Research

Nine doctors, two clinic days, one diary study

I needed to know three things before drawing anything: how long documentation actually takes, where in the consultation it hurts, and what a doctor would need to be true before letting software write into a medical record.

Methods

Interviews6 physicians (general medicine, endocrinology, cardiology), 2 clinic managers, 1 medical transcriptionist
Observation2 full clinic days, 31 consultations shadowed
Time diary12 physicians, 5 days, self-timed per note
TeardownGeneric dictation software, outsourced scribe services, EHR template libraries, phone voice memos

What people do instead today

Every doctor had a workaround, and every workaround had the same failure: it moved the work rather than removing it.

Copy-forward templatesFast, but the record stops saying anything
Voice memo, type laterDoubles the listening time
Human transcriptionistAccurate, 24-hour turnaround, expensive
Handwritten, scannedUnsearchable, illegible, still the most common

Five findings that shaped the product

The cost isn't typing speed, it's the context switch

Doctors are not slow typists. They are interrupted ones. Each note involved stopping, re-reading the previous encounter, recalling what was said, and reconstructing it in clinical language. Making the keyboard faster would have solved nothing.

The note already exists, spoken, in the room

In 22 of 31 consultations the doctor verbally summarised findings and plan to the patient — in almost the exact structure the note needed. That summary is the highest-quality source material in the building and it evaporates.

Trust is conditional, not binary

Not one participant objected to software drafting a note. Every participant objected to software committing one. The line was consistent and it was about authorship, not accuracy. This became the single hardest constraint in the design.

Consent is a workflow step, not a checkbox

Recording a patient changes the room. Clinic managers were more alert to this than doctors were. Consent had to be visible, captured before the visit, and revocable mid-recording without losing the rest of the day's work.

The paper is the product

The patient leaves with a printed prescription on clinic letterhead. If that sheet is wrong, the doctor rewrites it by hand and the software has failed, however good the transcript was. Print could not be an export option bolted on at the end.

I already tell the patient what I'm writing. Then I write it again for the computer.

Endocrinologist, 14 years in practice

I'll happily let it write. I will not let it decide.

Cardiologist, 9 years in practice

By the sixth patient I'm editing the fifth patient's note. That's not a record, that's a photocopy.

General medicine, 6 years in practice

If the print comes out wrong, he'll just write it by hand and I'm scanning paper again.

Clinic manager, 40-bed facility

Who it's for

Two people have to say yes

A doctor who will only adopt this if it is faster than typing, and a skeptic who will only adopt it if it is more accurate than his transcriptionist.

Primary

Dr. Anitha Rao, 41

Internal medicine · 24 patients a day · 20-minute slots

GoalFinish the day's notes before leaving the building
BehaviourTypes during the consultation, apologises for it
FrustrationNotes she wrote at 9pm read thinner than the visit actually was
FearSigning a record she didn't fully read
Adoption testUnder two minutes per note, including her review
Primary · skeptic

Dr. Vikram Shetty, 56

Cardiology · dictates to a human transcriptionist today

GoalNever personally touch a keyboard
BehaviourDictates at the end of a session, reviews the next morning
Frustration24-hour turnaround; patients leave with handwritten scripts
FearA machine mishearing a drug name or a dose
Adoption testDrug names and numbers correct, every time, visibly checkable
Secondary

Farida Qureshi, 29

Clinic coordinator · owns consent, rooms, printing, filing

Measured on room turnaround. She is the one who asks the patient about recording, and the one who reprints when the letterhead is wrong. If consent is buried inside the doctor's screen, she can't do her job — so consent state had to be legible from outside the visit.

Affected, not a user

Maria Fernandes, 52

Type 2 diabetes · quarterly follow-up · never touches the software

She doesn't log in, but she is the one being recorded and the one whose attention the doctor is currently spending on a keyboard. She's in this case study because the success measure isn't only minutes saved — it's whether the doctor was looking at her.

The brief

What the design had to do

The visit note should exist by the time the patient stands up — without the doctor typing it, and without her trusting it blindly.

How might we

Capture the summary a doctor already speaks, instead of asking her to produce it twice?

Make reviewing a draft genuinely faster than writing one from scratch?

Make it obvious, at a glance, which words came from the machine and which the doctor stands behind?

Let a doctor fix one wrong paragraph without redoing the other seven?

Constraints I designed against

No autonomous commitsNothing enters the permanent record without a doctor pressing save. The word "sign" never appears in the interface.
Consent before microphoneThe record button is unreachable until consent state is captured and displayed.
Grounded vocabularyClinical entities map to SNOMED CT, RxNorm and ICD-10 rather than free text, so drug names and diagnoses are constrained, not invented.
Print parityWhat the doctor approves on screen is exactly what prints on letterhead. No second formatting step.
Graceful fallbackManual notes stay one click away, always. A tool that can't be abandoned mid-visit won't be tried.

Feature set

What made version one, and what didn't

The longlist ran to nineteen ideas. Most were good. Shipping them would have turned a documentation tool into a clinical decision system, which is a different product with a different risk profile and a different regulatory conversation.

Build effort Impact on documentation time v1 cut line Live record and transcribe Structured 8-section note Audio file upload Clinical entity extraction Speaker labelling Live print preview Section-level re-dictation Consent capture Doctor review gate Hindi–English code-switching Drug name confirmation chips Auto-order labs from plan Ambient room microphone Patient summary in local language Billing code suggestion Voice commands for navigation Note quality scoring Multi-doctor rooms Differential diagnosis suggestions
Everything left of the dashed line shipped. The crossed-out item was refused outright, not deferred — see the decisions section for why.

Information architecture

Grafting a new branch onto a product that already exists

EHR.ai had six top-level areas and a settled mental model. The transcription flow couldn't be a seventh destination in the sidebar — documentation isn't a place you visit, it's something you do to a visit. So it lives inside the patient branch, one step from the visit overview, with manual notes as its permanent sibling.

EHR.ai Dashboard Patients Pharmacy Labs Billing Settings Patient directory Patient visit overview Manual notes Transcribe visit audio Upload file or record live Live transcript EHR note, review-ready Print on letterhead Save to patient record The eight note sections Chief complaints History Allergies Examination Diagnosis Prescription Advice Follow-up Each independently re-dictatable Existing New in this project
Manual notes stays permanently beside the new branch. Every doctor I spoke to wanted a visible way out, and a tool you can't abandon mid-visit doesn't get tried in the first place.

User flow

Including every way it goes wrong

The happy path is four steps. The design work was in the other branches: consent that hasn't been captured, consent withdrawn mid-sentence, a weak microphone, a file that's too big, and a section the doctor disagrees with. Each of those has to fail without destroying the visit.

Patient arrives, visit opens Consent captured? No Coordinator capturesconsent at check-in Declined Fall back to manual notes,no microphone offered Open transcribe visit audio Record live or upload? Upload Select audio file,25 MB maximum Too large: name the limit, keep the picker open Audio captured, speakerslabelled as doctor or patient Weak signal: show inputlevel, prompt to move mic Generate EHR noteSix named processing stages Consent withdrawn: stop,discard audio, keep visit open Draft note, eight sections Every section accurate? No Re-dictate that one sectionor edit it inline Yes Print preview correct? No Toggle which sectionsappear on the printout Save note and print New screens Failure and exit paths Decision
The two branches that came directly out of interviews: consent withdrawn mid-recording discards audio without closing the visit, and an oversized file names the actual limit instead of failing silently.

Journey map

The same twenty minutes, twice

Mapping the consultation end to end made the shape of the problem obvious. The doctor's experience is fine — good, even — right up until the moment the patient stops being the thing she's paying attention to.

Doctor's experience In control Behind The drop this project targets Today With voice capture Check-in Greeting Examination Explaining Documenting Wrap-up Today Pulls up lastvisit, re-reads Asks how thesugars have been Types findingswhile examining Says the planout loud, clearly Types the sameplan, alone Formats, reprintsif letterhead slips With voice capture Consent alreadycaptured at desk Presses record,puts hands down Examines withboth hands free Says the planonce, to the patient Reads a draft,fixes one section Prints what shejust approved
The curves are nearly identical for four of six stages. That's the argument for a narrow scope: the intervention only needs to exist in one place to change the shape of the day.

Sketches

Six ways to put a microphone in a consultation

The first question wasn't what the recorder looks like — it was where it lives. A modal interrupts. A separate page loses the patient's context. A docked bar is always present, which is exactly wrong for a feature that should only exist when consent has been given.

Modal over the visitblocks the labs she's talking about1 Full-screen recordercalm, but loses pt context2 Recorder L / transcript Rsee words land + pt card3live words Docked baralways on... consent??4 Editor + print previewanswers 'will it print right'5real Rx Section cards, mic eachfix one section, not all 86
Ballpoint on plain paper, before anything went near a screen. Three survived the walk-through and were circled: a split recording screen, a split review screen, and sections that can each be corrected on their own.

Low fidelity

Testing the layout before the pixels

Greyboxes went in front of four doctors on a laptop. Two things changed as a result. The transcript panel moved from below the recorder to beside it, because doctors wanted to watch the words land while still seeing the patient's risk flag. And the empty state stopped being blank — it now says what will appear and what to do to make it appear.

Empty state The panel says what will appear and what to do to fill it. Recording Transcript beside the controls, not beneath them. Review and print Sections on the left, the actual printed page on the right.
Tested as clickable greyboxes with four physicians. The blank right-hand panel in the first version was read as "broken" by three of four — which is how the empty state got its copy.

Final interface

Seven screens, one uninterrupted path

The flow runs from the morning list to a printed prescription without the doctor leaving the patient's context or typing a clinical sentence. Each screen below carries the decisions that produced it.

1

The morning list

The day starts here, and the visit reason column is doing real work. It carries enough clinical detail that the doctor walks into the room already oriented — which is what makes a short, high-signal consultation recordable.

EHR.ai dashboard showing the day's appointment list with patient names, times, departments and visit reasons
  • Start visit is a filled button only on the next two patients. Everything else offers Prep, so the primary action is never ambiguous.
  • Visit reasons are written as clinical shorthand, not appointment categories, because that's what the doctor reads in the four seconds before the patient sits down.
2

The visit overview, where the fork happens

Two buttons, equal weight in the layout, unequal weight in emphasis. Transcribe audio is primary; Manual notes sits beside it permanently. The bar along the bottom explains the choice in a sentence rather than assuming the doctor knows what transcription will do to her record.

Patient visit overview for Maria Fernandes showing conditions, medications, allergies, recent labs and past appointments, with manual notes and transcribe audio buttons
  • Consent on file appears in the patient header line, next to the appointment slot. It's status, not a prompt — the coordinator already handled it at check-in.
  • Labs, medications and conditions are on this screen and not the recording screen. The doctor reads them before the microphone opens, so recording time stays conversation time.
  • The footer bar links to Consent settings rather than hiding them in the global settings area, because that's where the question occurs to people.
3

Before anything is recorded

The empty state names what will appear and what to do to make it appear. It also states the constraint that testing showed mattered most: draft outputs require doctor review. That sentence sits in the page subheading, not in a dismissible tooltip.

Transcribe visit audio screen with patient card, consent captured, an audio upload dropzone, and an empty transcript panel
  • The patient card repeats name, risk flag, visit type, room and clinician. During a recording, this is the only patient context on screen, so it has to be complete.
  • Consent captured carries a check mark and a colour. Recording is not reachable without it.
  • The file limit and accepted formats are printed on the dropzone, before the error rather than after it.
  • Transcribe audio is disabled and stays visible, with the reason stated directly beneath. Hiding a disabled action makes people hunt for it.
4

Listening

The transcript is laid out as a conversation with speakers on opposite sides, because the thing a doctor scans for is attribution. "Patient reports tingling" and "on examination, tingling" are different clinical claims, and the note depends on getting that right.

Live recording in progress with a waveform, elapsed timer, and a speaker-labelled live transcript of the consultation
  • The recorder reports its own health: model, latency in milliseconds, input device, signal strength and input level in decibels. Dr. Shetty's objection was mishearing — this is the screen that answers it, continuously.
  • Stop is tinted red and separated from Pause. Stopping ends capture; pausing doesn't. Those outcomes shouldn't look alike.
  • The exchange count sits next to the generate button, so there's a concrete sense of how much material the note will be built from.
  • Generate EHR note is dark and heavy — the one moment in the flow where the doctor hands work to the system.
5

The eight seconds of waiting

A spinner here would have been a missed opportunity. Naming the six stages — and the vocabularies each one uses — does two jobs: it makes the wait legible, and it quietly establishes that entities are being mapped to SNOMED CT, RxNorm and ICD-10 rather than invented.

Processing panel listing six completed stages from analysing transcript to finalising the EHR note, each with the pipeline or vocabulary used
  • Stages are named after clinical outcomes, not engineering steps. "Generating chief complaint and history", not "running inference".
  • The final line reads Review-ready and carries the doctor's name. The system is explicitly handing the note back, not finishing it.
  • The page behind stays visible and dimmed. The transcript is still there if the doctor wants to check something mid-generation.
6

The note arriving, section by section

Sections fill in sequence and the print preview builds alongside. The doctor can start reading chief complaints while examination is still being written, which is why time-on-task fell further than raw generation speed accounts for.

EHR note editor with section cards still empty on the left while the print preview on the right is already populated
  • Every section has its own microphone and eye icon: re-dictate this section, or toggle whether it appears on the printout.
  • The counter reads 8 of 8 sections visible in preview, so the relationship between editor and printout is never a guess.
  • Actions are Save note and Print. The word "sign" does not appear anywhere in this product.
7

Review-ready

The finished draft, in clinical language, on clinic letterhead, with the prescription block laid out the way the clinic already prints it. What the doctor approves on the left is exactly what comes out of the printer on the right — no second formatting pass, which is where a full minute per patient used to go.

Completed EHR note with chief complaints, history, allergies and examination filled in, beside a print preview on Greenfield Medical Clinic letterhead
  • Allergies reads NKDA — no known drug allergies, written out. Abbreviations that carry safety weight are expanded on the printed sheet the patient takes home.
  • Examination keeps numbers as numbers: BP, HR, RR, SpO₂, temperature, BMI. These are the values most likely to be misheard, so they're the values most visible for checking.
  • The risk flag follows the patient through all seven screens, in the same position and the same colour.

Decisions

Six choices that would break the product if reversed

The system drafts. The doctor authors.

Nothing reaches the permanent record without a human pressing save. The generation panel hands the note back as "review-ready". The buttons say save and print. The word "sign" is absent by design, because signing implies a completed act and this note isn't complete until someone has read it. This is the constraint every other decision bent around.

No diagnostic suggestions, ever — and that's a feature

Suggesting differentials was the most requested idea on the longlist and the one I refused outright rather than deferring. The moment software proposes what might be wrong with a patient, it stops being a documentation tool and becomes clinical decision support: different evidence requirements, different regulatory footing, different liability, and a trust conversation that would have swallowed the one this product needed to win. Narrow scope was the strategy.

Consent is a state, not a step

It appears on the visit overview, in the recording screen's patient card, and as a gate on the record control. The coordinator captures it at check-in, away from the doctor's flow entirely. Withdrawing it mid-recording stops capture and discards the audio without closing the visit — a branch that exists because a clinic manager described exactly that moment happening with a phone recorder.

The section is the unit of correction

A doctor who disagrees with the history shouldn't have to regenerate the prescription. Per-section microphones mean the correction cost is proportional to the error. This is also the safety valve that makes the review gate tolerable: disagreeing is cheap, so disagreeing actually happens.

Print preview lives beside the editor, not behind a button

Research finding five said the paper is the product. A minute per patient was going on formatting and reprinting. Rendering the letterhead page live, next to the sections that feed it, removes the format-check round trip entirely and makes the eye icons legible — you can see a section disappear from the sheet as you toggle it.

The recorder reports its own health

Model, latency, input device, signal strength, input level, elapsed time, exchange count. This looks like engineering detail leaking into the interface, and I kept it deliberately. The skeptic persona's entire objection was accuracy, and accuracy is unobservable. What's observable is whether the microphone is hearing properly — so the interface shows that instead, continuously, and the objection stops being abstract.

Testing

Six doctors, three tasks, one prototype

Moderated sessions on a clickable prototype with scripted consultation audio. Each participant recorded a visit, reviewed the generated note, corrected at least one deliberately wrong section, and printed.

MeasureBeforeAfterNote
Median time per note11:201:48Review and correction included
Notes finished before the patient left39%94%The metric doctors actually cared about
Sections corrected per note1.3Most often examination values
Reprints for formatting0.40.05Per note
System usability scale84n=6

What testing changed

Drug names needed a second look

Two participants caught a dose transcribed from accented speech. That produced the confirmation-chip concept: drug and dose entities get a light outline in the prescription section until the doctor's eye has passed over them once.

Nobody trusted a silent success

An early build generated the note in under two seconds with no visible stages. Participants assumed it hadn't read the whole conversation. The six-stage panel tested better and felt faster, despite taking longer.

"Stop" was being hit instead of "Pause"

In the first build they were identical outline buttons. Colour, separation and a red square icon fixed it. Two different outcomes shouldn't share an appearance.

The empty transcript panel read as an error

Three of four low-fidelity participants thought the blank right-hand panel meant something had failed. It now states what will appear there and what to do next.

What's next

The parts I'd want to get wrong early

Hindi and English in the same sentence

Every participant code-switched mid-consultation. It's the highest-value feature not in version one and the largest technical risk in the product. Shipping monolingual first was a scoping decision, not a judgement about what matters.

The six-minute consultation

This was designed around a twenty-minute slot. High-volume clinics run six-minute visits, where the review step could plausibly cost more than it saves. That needs testing before it's claimed.

Did the notes actually get better?

Time saved is the easy metric and the wrong one to optimise alone. The real question is whether records became more specific and less copy-forward — which needs blinded clinical review of paired notes, not a stopwatch.

What the doctor does with the time

Nine minutes per patient can go back to the patient or into the schedule as three more patients. Which of those happens is a clinic policy question the design can influence but not decide, and it determines whether this helps anyone.