Implemented across 12 commits on feat/voice-treatment-entry. Still blocked on
the Persian ASR spike before it is trustworthy in front of patients: nothing in
the implementation compensates for a bad transcript.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The microphone becomes the second segment of the Add detail button, built like
the detail chip's trash affordance in the same file — an overflow-hidden rounded
wrapper holding two raw <button>s divided by border-s — rather than two shared
Buttons, which each hardcode their own rounding and would fight a segmented
control. border-s puts the mic at the logical end: visually right in en/nl,
visually left in fa, on the same side as the chip's trash in both directions.
The two halves share a wrapper and nothing else. Add keeps its exact behaviour.
The control never changes size while recording; the timer and level meter live
in a bar between the header row and the chip strip, because the header is
sm:justify-between and growing the button would shove the row on every start and
stop. The meter exists to prove the microphone is actually hearing something —
silence and a dead mic look identical otherwise.
Voice reaches the editor as one optional `voice` prop, so its absence *is* the
unavailable state and the two cannot disagree.
Fixes from review of this commit:
- mountedRef was set false on unmount and never re-armed, so under StrictMode
the hook was permanently "unmounted" in dev and recording silently never
started.
- onStart guarded only on `phase`, which does not change until getUserMedia
resolves; a second click during the permission prompt orphaned the first
MediaStream, leaving the mic indicator lit.
- Week start is now per locale. "Next Thursday" is week-relative, and hardcoding
Saturday put an en/nl clinician's deadline a week out.
- A missing `which` on a weekday intent is read as "this" rather than failing —
a bare weekday carries no qualifier, and rejecting it discarded a real
deadline.
- durationMs is client-reported and so is a claim, not enforcement; the cap is
now also checked against the vendor's own usage.seconds.
- Blob type falls back to the recorder's actual mimeType before webm, so old
Safari's mp4/aac clips are not mislabelled.
Two review findings were rejected as incorrect, both re-verified against live
sources: google/gemini-3.7-flash does exist on OpenRouter (1M context,
$0.375/$1.875 per M), and base64 JSON input_audio is the documented primary
path for /audio/transcriptions, with multipart as the OpenAI-compatible
alternative. The spec's stale "unverified" note is corrected, and the provider
now has unit tests covering the request shape, usage parsing, and that a vendor
error body never reaches the thrown message.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Design spec for filling a TreatmentDetail by voice, settled across three
grilling sessions (30 decisions, logged in the spec).
Key shape:
- two-stage pipeline: OpenRouter whisper-1 -> gemini-3.7-flash
- the LLM emits *intents*, never FDI codes or ISO dates; pure Jest-tested
backend resolvers own quadrant mapping and Jalali conversion
- provider registry keyed by locale so fa can diverge from en/nl
- review sheet confirms before anything touches the form
- audio and transcripts are never persisted
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>