1 · Normal speech → synthetic whisper
Same speaker, converted to a true (unvoiced) whisper. This is the synthetic training signal.
2 · Generated ambient sound
Real-world backgrounds synthesized from text (aircraft cabin, café/crowd, bar, train) — the noise library we mix in.
3 · Whisper + noise (P.56-calibrated SNR)
Synthetic whisper placed into each scene at a declared SNR. Speech level measured by ITU-T P.56 active speech level, not naive RMS.
4 · Normal speech + noise (P.56-calibrated SNR)
The same sentences in normal voice, mixed into the same scenes — the everyday-call baseline.