Consider a perfectly transcribed instruction: Put it over there. The words may be exact. The destination is still absent. Spoken work is full of references whose meaning lives partly outside the sentence, in what participants can see, what was said earlier, who is authorized to act, and how the next person responds. A useful voice system therefore has two separate jobs: recognize the words and represent what those words establish. Success at the first does not guarantee success at the second.
A correction is part of the record
Conversation has organized ways to handle problems in speaking, hearing, and understanding. Schegloff, Jefferson, and Sacks distinguished initiating repair from completing it, and documented a preference for speakers to correct their own talk. That distinction is easy to flatten in a transcript-to-record pipeline. No, bay four, followed by sorry, bay five should not leave two equally current locations or silently preserve only the first recognized phrase. The record needs to represent that a correction occurred and which proposition it superseded.[3]
Repair can also fail. A listener may ask which gate and receive no answer, or hear a response that creates a second ambiguity. A system should not convert the presence of a correction sequence into evidence that the exchange ended with agreement. The defensible state may be location unresolved or correction attempted, pending confirmation. This is an editorial and product principle drawn from the literature, not a measured claim about Memo or any particular workplace.[3][1]
More context is not automatically better
Previous turns can help a recognizer resolve an acoustically uncertain name or complete a reference. Recent dialogue ASR research uses accumulated conversational context for that purpose. The same work identifies the opposing mechanism: when earlier context contains recognition errors, those errors can degrade later recognition. The researchers introduced noise-aware training to make context use more robust, which is evidence that simply feeding more history into a model is not a sufficient control.[4]
Operational context introduces further hazards that the dialogue study did not test. A previous location may be stale. Two jobs may use similar equipment names. A speaker may deliberately reverse an earlier instruction. Future turns can make a past utterance look obvious even though a live system could not have known them at the time. Our proposed evaluation would compare isolated speech, bounded prior context, irrelevant or shuffled context, and structured facts with known freshness. Future information belongs in retrospective analysis, not in a live-input condition.[4]
Represent what the exchange establishes
A useful output separates layers that are often collapsed: the audio that was received, the words attributed to it, the referents linked from context, and the operational state inferred from the exchange. Put it over there may be transcribed with high confidence while its destination remains unknown. Copy may establish audibility in one team, acceptance of responsibility in another, or merely a conventional response. The participants' account and the subsequent action are evidence; an outside model's fluent interpretation is not a substitute for either.[1][2]
Our proposed evaluation would score these layers separately. It can measure transcript accuracy, entity and slot precision, correction handling, unsupported assertions, and appropriate abstention. It should include unresolved references and misleading context as deliberate negative controls. If adding context raises recall while increasing false certainty, the tradeoff should be visible rather than hidden inside one aggregate score. The result may show that a simple rule or an explicit request for clarification is safer than a larger contextual model.[4]
Evidence boundary
Limits & open questions
- The foundational grounding and repair studies examined conversation and controlled reference tasks, not operational two-way radio.
- The dialogue ASR study does not establish performance under push-to-talk clipping, channel contention, local jargon, or industrial noise.
- This review proposes evaluation principles. It reports no new field observations, model measurements, or Memo results.
- Local conventions can make terse language meaningful to familiar participants, so an outside interpretation should not be treated as ground truth without participant or outcome evidence.
How this article was researched
Proposal based on a bounded review of primary conversation research and a peer-reviewed dialogue ASR study. Claims are limited to what those studies examined. A Memo-specific follow-up should use consented, independently annotated episodes with isolated, prior-context, irrelevant-context, and structured-context conditions; forbid future-context leakage; preserve unresolved labels; and report extraction precision, recall, unsupported assertions, abstention, correction handling, latency, and denominators by condition.
Article type: research review. The diagrams are explanatory schematics. No new Memo field trial, corpus measurement or model benchmark is reported here.
Sources & further reading
Sources reviewed on 23 September 2026. “n.d.” indicates no fixed publication date for the linked version.
- Grounding in communicationAmerican Psychological Association · 1991
- Referring as a collaborative processCognition · 1986
- The preference for self-correction in the organization of repair in conversationLanguage · 1977
- Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation LearningAssociation for Computational Linguistics · 2024