Skip to content
Memo
← All research

Voice & Meaning / Research review

When the transcript is right and the meaning is wrong

A transcript can preserve every spoken word and still lose the shared context, correction, and uncertainty that made the exchange intelligible.

Memo Research · · 4 min read
“Meet at the north gate.”“I’m heading there.”“Correction. South gate.”
Figure 01 Illustrative exchange: a preceding location gives meaning to “there”; a later correction changes the current destination. The later turn must not be used as if a live system already knew it.

Consider a perfectly transcribed instruction: Put it over there. The words may be exact. The destination is still absent. Spoken work is full of references whose meaning lives partly outside the sentence, in what participants can see, what was said earlier, who is authorized to act, and how the next person responds. A useful voice system therefore has two separate jobs: recognize the words and represent what those words establish. Success at the first does not guarantee success at the second.

Words depend on shared ground

Communication researchers use grounding to describe how participants establish enough mutual understanding for the activity at hand. Clark and Brennan argue that coordination depends on common ground that is updated as an interaction unfolds. The standard is practical rather than absolute: participants need sufficient evidence of understanding for the current purpose. That evidence may come from an explicit acknowledgement, a relevant next action, a clarification, or the absence of trouble when a next turn begins.[1]

This matters because a transcript is only one contribution to that process. Clark and Wilkes-Gibbs showed in a controlled reference task that speakers and addressees developed descriptions together, repairing or replacing them until both accepted a workable reference. Their experiment used abstract figures, not radios or industrial work, so it does not prove how a crew behaves. It does establish a broader warning: a noun phrase is not always a self-contained label whose meaning can be recovered without the interaction that shaped it.[2]

A correction is part of the record

Conversation has organized ways to handle problems in speaking, hearing, and understanding. Schegloff, Jefferson, and Sacks distinguished initiating repair from completing it, and documented a preference for speakers to correct their own talk. That distinction is easy to flatten in a transcript-to-record pipeline. No, bay four, followed by sorry, bay five should not leave two equally current locations or silently preserve only the first recognized phrase. The record needs to represent that a correction occurred and which proposition it superseded.[3]

Repair can also fail. A listener may ask which gate and receive no answer, or hear a response that creates a second ambiguity. A system should not convert the presence of a correction sequence into evidence that the exchange ended with agreement. The defensible state may be location unresolved or correction attempted, pending confirmation. This is an editorial and product principle drawn from the literature, not a measured claim about Memo or any particular workplace.[3][1]

More context is not automatically better

Previous turns can help a recognizer resolve an acoustically uncertain name or complete a reference. Recent dialogue ASR research uses accumulated conversational context for that purpose. The same work identifies the opposing mechanism: when earlier context contains recognition errors, those errors can degrade later recognition. The researchers introduced noise-aware training to make context use more robust, which is evidence that simply feeding more history into a model is not a sufficient control.[4]

Operational context introduces further hazards that the dialogue study did not test. A previous location may be stale. Two jobs may use similar equipment names. A speaker may deliberately reverse an earlier instruction. Future turns can make a past utterance look obvious even though a live system could not have known them at the time. Our proposed evaluation would compare isolated speech, bounded prior context, irrelevant or shuffled context, and structured facts with known freshness. Future information belongs in retrospective analysis, not in a live-input condition.[4]

Represent what the exchange establishes

A useful output separates layers that are often collapsed: the audio that was received, the words attributed to it, the referents linked from context, and the operational state inferred from the exchange. Put it over there may be transcribed with high confidence while its destination remains unknown. Copy may establish audibility in one team, acceptance of responsibility in another, or merely a conventional response. The participants' account and the subsequent action are evidence; an outside model's fluent interpretation is not a substitute for either.[1][2]

Our proposed evaluation would score these layers separately. It can measure transcript accuracy, entity and slot precision, correction handling, unsupported assertions, and appropriate abstention. It should include unresolved references and misleading context as deliberate negative controls. If adding context raises recall while increasing false certainty, the tradeoff should be visible rather than hidden inside one aggregate score. The result may show that a simple rule or an explicit request for clarification is safer than a larger contextual model.[4]

Evidence boundary

Limits & open questions

  • The foundational grounding and repair studies examined conversation and controlled reference tasks, not operational two-way radio.
  • The dialogue ASR study does not establish performance under push-to-talk clipping, channel contention, local jargon, or industrial noise.
  • This review proposes evaluation principles. It reports no new field observations, model measurements, or Memo results.
  • Local conventions can make terse language meaningful to familiar participants, so an outside interpretation should not be treated as ground truth without participant or outcome evidence.

How this article was researched

Proposal based on a bounded review of primary conversation research and a peer-reviewed dialogue ASR study. Claims are limited to what those studies examined. A Memo-specific follow-up should use consented, independently annotated episodes with isolated, prior-context, irrelevant-context, and structured-context conditions; forbid future-context leakage; preserve unresolved labels; and report extraction precision, recall, unsupported assertions, abstention, correction handling, latency, and denominators by condition.

Article type: research review. The diagrams are explanatory schematics. No new Memo field trial, corpus measurement or model benchmark is reported here.

Sources & further reading

Sources reviewed on 23 September 2026. “n.d.” indicates no fixed publication date for the linked version.

  1. Grounding in communicationAmerican Psychological Association · 1991
  2. Referring as a collaborative processCognition · 1986
  3. The preference for self-correction in the organization of repair in conversationLanguage · 1977
  4. Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation LearningAssociation for Computational Linguistics · 2024