Whisper ASR — synthetic-data demo · conversion · generated ambience · noisy mixes

1 · Normal speech → synthetic whisper

Same speaker, converted to a true (unvoiced) whisper. This is the synthetic training signal.

2 · Generated ambient sound

Real-world backgrounds synthesized from text (aircraft cabin, café/crowd, bar, train) — the noise library we mix in.

3 · Whisper + noise (P.56-calibrated SNR)

Synthetic whisper placed into each scene at a declared SNR. Speech level measured by ITU-T P.56 active speech level, not naive RMS.

4 · Normal speech + noise (P.56-calibrated SNR)

The same sentences in normal voice, mixed into the same scenes — the everyday-call baseline.