Accelerating Diffusion-Based Speech Denoising with DDIM: A Cross-Lingual Evaluation of DiffWave on English and Arabic Speech

Main Article Content

Abderrahmane Adjila, Maamar Ahfir, Meriem Laouar, Sarra Zahouani

Abstract

Speech denoising is essential for reliable human-machine interaction, yet diffusion-based generative models, despite strong reconstruction quality, remain limited by high inference latency and near-exclusive evaluation on English speech. This study adapts the DiffWave architecture to speech denoising, compares two diffusion-step choices (T=50, T=200), accelerates inference with DDIM, and evaluates cross-lingual generalization to Arabic speech. Models were trained on the SC09 dataset and evaluated on SC09, Arabic Speech Commands, and LJSpeech, using 14 noise types (6 synthetic, 8 real) at three SNR levels (-5, 0, 5 dB), and six metrics: SNR, SI-SDR, PESQ, STOI, MOS, and RTF. The results showed that T=50 outperformed T=200 on all six metrics (SI-SDR: 10.93 vs 9.94 dB) due to longer training convergence. Under the full 14-noise/3-SNR benchmark, the model achieved its best overall performance on Arabic speech (SI-SDR = 3.94 dB, success rate = 71.2%), ahead of SC09 (3.37 dB, 69.8%) and LJSpeech (1.63 dB, 59.0%). DDIM reduced sampling from 50 to 3 steps, reaching an RTF of 0.0732 (13.7× real-time) with only a -0.64 dB SI-SDR loss on Arabic, while matching or exceeding DDPM on PESQ and STOI. As a conclusion DDIM-accelerated DiffWave offers a practical balance between denoising quality and computational efficiency, and generalizes to Arabic despite training exclusively on English digits, supporting its use in real-time, cross-lingual speech enhancement applications.

Article Details

Section
Articles