5 Interspeech papers
7 July 2026, by David Mosteller
We are happy to share that our Signal Processing Group at the University of Hamburg will present five papers at Interspeech 2026 in Sydney this September. Congratulations to all co-authors involved.
A brief overview:
* A Fast Solver for Interpolating Stochastic Differential Equation Diffusion Models for Speech Restoration (Interspeech long paper)
Bunlong Lay, Timo Gerkmann
We develop a unified formalism of interpolating SDEs (including SGMSE+) and derive a fast solver for them, enabling high-quality speech restoration with as few as 10 network evaluations across noise reduction, dereverberation, declipping, MP3 decoding, and bandwidth extension.
[arxiv]
* Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
Danilo de Oliveira, Tal Peer, Timo Gerkmann
ASR-based word error rate is a popular metric for speech enhancement, but depends heavily on the ASR model and text normalization. We find that models trained on large-scale noisy data correlate best with human recognition, yet can be uninformative for acoustics-focused evaluation—so the ASR setup should be stated and motivated clearly.
* Your U-Net Dereverberation Model is Secretly an RIR Encoder
Sina Khanagha, Timo Gerkmann
We show that NCSN++ U-Net dereverberation models implicitly encode room-impulse-response information in their deeper layers, with this correlating with performance. Explicitly conditioning on pre-trained RIR embeddings improves quality, speeds up convergence, and reduces the reverse diffusion steps needed at inference.
[arxiv]
* EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase
Renzheng Shi, Simon Welker, Timo Gerkmann, Tim Fingscheidt
We propose EffVOC, an efficient low-delay vocoder that reconstructs wideband or fullband speech from either amplitude spectra or Mel coefficients within one unified framework. A collaboration with TU Braunschweig.
* Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Rostislav Makarov, Timo Gerkmann
Classifier-guided diffusion normally needs two separate models. We instead freeze a conventionally trained speech classifier and train only a lightweight subnetwork on its representations, yielding a single-backbone generator that matches a standard U-Net at lower memory and compute, especially in low-data and zero-shot regimes.
[arxiv]
