Local transcription can insert words during silence because the decoder must resolve weak acoustic evidence using learned language patterns and inherited context.
A supposedly silent home recording still contains microphone self-noise, fans, room hum, compression artifacts, and reverberant tails. If the pipeline sends that segment to an autoregressive recognizer, the model receives features rather than an explicit instruction to output nothing. Previous text, language selection, decoding thresholds, and chunk boundaries can turn uncertain sound into plausible phrases.
Silence Produces Features Even Without Speech
Digital silence may be zeros, but real captured silence contains low-level energy and structured noise. Feature extraction converts that spectrum into frames that can resemble weak phonetic patterns, especially after automatic gain control raises the noise floor or a codec creates periodic artifacts.
An investigation of non-speech hallucinations deliberately exposes Whisper to non-speech sounds and finds recurring hallucinated phrases. The repetition indicates that ambiguous acoustics interact with learned output patterns rather than producing arbitrary characters. This distinction remains visible during later household testing.
A voice-activity detector can block clearly non-speech regions before decoding, but its own false positives rise with television audio, music, breathing, and appliances. Silence detection and transcription are separate classifiers with separate thresholds and failure modes.
Autoregressive Decoding Fills Acoustic Uncertainty With Language Priors
Speech decoders score text from both acoustic evidence and learned sequence likelihood. When the acoustic term is weak, common phrases, punctuation, subtitle conventions, or text supplied as a previous prompt can dominate the next-token decision.
A study of weakly supervised ASR describes large-scale weakly supervised speech recognition across varied data and tasks. That breadth provides strong language regularities, yet it also means decoding can continue with plausible text when a local segment contributes little speech evidence.
Long windows can include a few real words followed by silence, letting the model extend their pattern. Short windows can remove context needed to reject noise. Chunk overlap, prompt carryover, temperature fallback, and no-speech thresholds therefore influence which failure appears.
Post-Processing Can Hide Frequency but Not Establish Truth
A blacklist can remove phrases that recur during silence, and confidence filters can suppress low-probability segments. These safeguards reduce visible errors but can also delete genuine quiet speech containing the same words. The intermediate result must remain inspectable before automation follows.
A study of unsupported transcribed phrases reports complete invented phrases and highlights that fluent output can look credible even when unsupported by audio. The operational lesson is to retain timing and acoustic evidence for verification rather than trusting grammaticality.
The failure boundary is calling every word over a quiet waveform a hallucination. Distant speech, headphone bleed, ultrasonic interference folded into the band, or aggressive noise suppression may make speech hard to hear but still present. Inspect raw audio and synchronized levels before labeling the model.
Create a Silence Challenge Set Before Changing Thresholds
Record true digital silence, room tone, fans, HVAC, keyboard noise, television bleed, music, breathing, and quiet real speech at several gains. Preserve raw audio, then run identical chunks with and without voice gating, prompt carryover, and temperature fallback.
Measure false words per silent minute alongside missed quiet speech, using the gating mechanism described in voice-activity gating. Store no-speech probability, VAD decision, average energy, decoder confidence, language choice, chunk boundary, and emitted tokens for every failure.
Set thresholds where silent insertions fall without unacceptable quiet-speech loss. For consequential notes, retain timestamps and audio review; if a recurring phrase survives multiple controls, treat it as a decoder artifact rather than evidence that speech occurred.
Tech & AI HUB
More to Read

How Does Time-Series Downsampling Affect Smart Home Anomaly Detection?
See how bucket width, aggregation, anti-aliasing, missing data, event duration, and multiscale retention change smart home anomaly recall.

How Does an Occupancy Grid Combine Weak Smart Home Signals?
Learn how spatial cells, sensor models, log-odds updates, decay, correlated evidence, and thresholds turn weak home signals into occupancy estimates.

How Does Photometric Normalization Affect Private Face Clustering?
See how illumination correction changes face crops, embeddings, cluster distances, thresholds, over-normalization, and private photo-search evaluation.

