An Adaptive Kalman Speech Enhancement Method Driven by Burst Noise Suppression and Dual-Time-Scale Perception
-
摘要: 针对传统基于自回归(AR)模型的卡尔曼滤波语音增强算法在非平稳噪声环境下噪声模型难以自适应更新、突发噪声易导致模型失配以及协方差参数调节缺乏环境感知机制等问题,本文提出一种用于突发噪声抑制与双时间尺度感知驱动的自适应卡尔曼语音增强方法。首先,本文引入基于频谱能量熵(EER)构建的环境偏离比,用于刻画当前帧与背景环境统计特性的偏离程度。在此基础上,建立短时与长时双时间尺度EER跟踪机制,分别用于表示瞬时变化与稳定环境统计,并利用两者之间差异构造自适应偏离强度因子,从而实现对语音过程噪声协方差与观测噪声协方差的联合调节。其次,结合突发噪声的核心特性,提出一种基于能量突变阈值和频谱平坦度结合的突发噪声帧判断策略,用于有效区分时域能量突变且带有弱有色或白噪声特性的突发噪声,另外,提出一种基于线性预测残差方差比的语音可建模性因子判决准则,用于进一步区分语音结构与突发噪声成分,避免中等能量强度突发噪声误判导致语音模型污染,从而提升噪声建模的鲁棒性与可靠性。本文算法实验采用NOIZEUS数据库上的音频样本,结果表明,本文所用方法在短时客观可懂度(STOI)、感知语音质量评价指标(PESQ)和分段信噪比(SegSNR)等客观指标上优于传统的AR卡尔曼滤波,以及基于$ {J}_{1} $灵敏度的增广卡尔曼滤波,尤其在非平稳噪声和突发噪声条件下表现出更强的自适应能力与增强稳定性。Abstract:
Objective The traditional Auto-Regressive (AR) Kalman speech enhancement algorithm has three critical drawbacks under non-stationary noise: difficult adaptive noise model update, easy model mismatch caused by burst noise, and lack of environment-aware covariance adjustment. These defects degrade enhancement performance and cannot satisfy practical speech communication demands. To tackle the above issues, this paper proposes an improved adaptive Kalman method for better speech quality and intelligibility under complex noise. Methods First, an environment deviation ratio based on Energy Entropy Ratio (EER) is constructed to measure statistical deviation between the current frame and background. A dual-time-scale EER tracker is built to capture instantaneous fluctuations and steady background statistics respectively, and their difference generates an adaptive intensity factor for joint adjustment of process and observation noise covariances. Second, combining burst noise features, a two-stage discrimination scheme is presented: energy mutation threshold combined with spectral flatness detects impulsive noise; a Speech Modelling Metric derived from linear prediction residual variance further separates speech from medium-energy burst noise and avoids AR model contamination. Results and Discussions Experiments on NOIZEUS dataset show the proposed method outperforms classic AR-Kalman and J1-sensitivity based improved Kalman in STOI, PESQ and SegSNR. It gains better adaptability and stability under non-stationary and burst noise. Dual-time-scale perception and accurate burst identification accelerate environmental adaptation and reduce speech distortion. Conclusions The burst-suppressed dual-time-scale adaptive Kalman algorithm solves inherent defects of traditional AR-Kalman in complex noise. EER-based tracking realizes environment-aware covariance tuning, while the two-stage judgment greatly improves burst noise detection accuracy. Objective results verify its strong robustness and provide a feasible scheme for practical non-stationary speech enhancement. -
表 1 :不同SNR下STOI、PESQ、SegSNR评分
SNR STOI PESQ SegSNR SS KF AKF 本文方法 SS KF AKF 本文方法 SS KF AKF 本文方法 0 dB 0.64 0.64 0.705 0.70 1.89 1.95 1.82 1.87 –3.72 –4.77 –2.60 –1.96 5 dB 0.76 0.76 0.791 0.80 2.10 2.10 2.02 2.05 –0.94 –2.02 –0.33 –0.33 10 dB 0.89 0.881 0.87 0.89 2.31 2.35 2.35 2.42 2.65 1.42 2.01 2.83 15 dB 0.95 0.951 0.92 0.94 2.52 2.69 2.50 2.67 5.69 5.04 6.05 5.92 注:SS:基础谱减;KF:传统卡尔曼滤波;AKF:基于$ {J}_{1} $的增广卡尔曼滤波. 表 2 不同SNR下STOI、PESQ、SegSNR评分
Noise Type AKF 本文方法 Babble 0.66 0.66 Car 0.82 0.84 Restaurant 0.79 0.79 Station 0.79 0.780 -
[1] 王帅. 实时语音增强人工耳蜗的技术研究[D]. [硕士论文], 中国科学院大学(中国科学院沈阳计算技术研究所), 2017.WANG Shuai. Study on real time speech enhancement of cochlear implant[D]. [Master dissertation], Shenyang Institute of Computing Technology, Chinese Academy of Sciences, 2017. [2] 张殿熙, 乔兆亮. 语音增强技术及应用[C]. 天津市电子工业协会2025年年会论文集, 天津, 2025: 29–31. doi: 10.26914/c.cnkihy.2025.026058.ZHANG Dianxi and QIAO Zhaoliang. Speech enhancement technology and applications[C]. 2025 Annual Conference of Tianjin Electronic Industry Association, Tianjin, China, 2025: 29–31. doi: 10.26914/c.cnkihy.2025.026058. (查阅网上资料,未找到标黄信息,请确认). [3] 曹丽静. 语音增强技术研究综述[J]. 河北省科学院学报, 2020, 37(2): 30–36. doi: 10.16191/j.cnki.hbkx.2020.02.006.CAO Lijing. Overview of speech enhancement algorithms[J]. Journal of the Hebei Academy of Sciences, 2020, 37(2): 30–36. doi: 10.16191/j.cnki.hbkx.2020.02.006. [4] 杜扶遥, 姜囡, 刘浠辰. 涉案语音的降噪处理分析研究[J]. 广东公安科技, 2024, 32(4): 30–35.DU Fuyao, JIANG Nan, and LIU Xichen. Research on noise reduction processing of involved speech[J]. Guangdong Public Security Science and Technology, 2024, 32(4): 30–35. (查阅网上资料, 未找到对应的英文翻译, 请确认). [5] 王涛, 鲁怀伟, 刘宝成. 基于AR模型的Kalman语音增强算法[J]. 青岛大学学报(自然科学版), 2018, 31(2): 48–53. doi: 10.3969/j.issn.1006-1037.2018.05.09.WANG Tao, LU Huaiwei, and LIU Baocheng. Kalman speech enhancement algorithm based on AR model[J]. Journal of Qingdao University (Natural Science Edition), 2018, 31(2): 48–53. doi: 10.3969/j.issn.1006-1037.2018.05.09. [6] JODWAL M, KUMAR S, COLNEY L, et al. Performance analysis of speech enhancement techniques[C]. 2024 First International Conference on Electronics, Communication and Signal Processing, New Delhi, India, 2024: 1–7. doi: 10.1109/ICECSP61809.2024.10698182. [7] 王华朋, 冯嘉琪. 基于深度学习的语音增强方法综述[J]. 科学技术与工程, 2025, 25(20): 8331–8346. doi: 10.12404/j.issn.1671-1815.2404954.WANG Huapeng and FENG Jiaqi. Review of speech enhancement methods based on deep learning[J]. Science Technology and Engineering, 2025, 25(20): 8331–8346. doi: 10.12404/j.issn.1671-1815.2404954. [8] NIAN Zhaoxu, TU Yanhui, DU Jun, et al. A progressive learning approach to adaptive noise and speech estimation for speech enhancement and noisy speech recognition[C]. 2021 IEEE International Conference on Acoustics, Speech and Signal Processing, Toronto, Canada, 2021: 6913–6917. doi: 10.1109/ICASSP39728.2021.9413395. [9] WEI Haimeng and ZHANG Xiaobo. Time-frequency conformer and Kalman filter-based speech enhancement method[C]. 2025 7th International Conference on Intelligent Control, Measurement and Signal Processing, Xian, China, 2025: 325–329. doi: 10.1109/ICMSP68723.2025.11407752. [10] GEORGE A E W, SO S, GHOSH R, et al. Robustness metric-based tuning of the augmented Kalman filter for the enhancement of speech corrupted with coloured noise[J]. Speech Communication, 2018, 105: 62–76. doi: 10.1016/j.specom.2018.10.002. [11] ROY S K and PALIWAL K K. Sensitivity metric-based tuning of the augmented Kalman filter for speech enhancement[C]. 2020 14th International Conference on Signal Processing and Communication Systems, Adelaide, Australia, 2020: 1–6. doi: 10.1109/ICSPCS50536.2020.9310005. [12] XU Xiaodong, FLYNN R, and RUSSELL M. Speech intelligibility and quality: A comparative study of speech enhancement algorithms[C]. 2017 28th Irish Signals and Systems Conference, Killarney, Ireland, 2017: 1–6. doi: 10.1109/ISSC.2017.7983599. [13] VASEGHI S V. Linear prediction models[M]. VASEGHI S V. Advanced Digital Signal Processing and Noise Reduction. 2nd ed. Chichester: John Wiley & Sons, 2001: 227–262. doi: 10.1002/0470841621.ch8. (查阅网上资料,未能确认年份和版本修改是否正确,请确认). [14] KOO B, GIBSON J D, and GRAY S D. Filtering of colored noise for speech enhancement and coding[C]. International Conference on Acoustics, Speech, and Signal Processing, Glasgow, UK, 1989: 349–352. doi: 10.1109/ICASSP.1989.266437. [15] 王文益, 伊雪. 基于改进语音存在概率的自适应噪声跟踪算法[J]. 信号处理, 2020, 36(1): 32–41. doi: 10.16798/j.issn.1003-0530.2020.01.005.WANG Wenyi and YI Xue. An adaptive noise tracking algorithm using improved speech presence probability[J]. Journal of Signal Processing, 2020, 36(1): 32–41. doi: 10.16798/j.issn.1003-0530.2020.01.005. [16] DUBNOV S. Generalization of spectral flatness measure for non-Gaussian linear processes[J]. IEEE Signal Processing Letters, 2004, 11(8): 698–701. doi: 10.1109/LSP.2004.831663. [17] 兰朝凤, 蒋朋威, 陈欢, 等. 基于双路径递归网络与Conv-TasNet的多头注意力机制视听语音分离[J]. 电子与信息学报, 2024, 46(3): 1005–1012. doi: 10.11999/JEIT230260.LAN Chaofeng, JIANG Pengwei, CHEN Huan, et al. Multi-head attention time domain audiovisual speech separation based on dual-path recurrent network and Conv-TasNet[J]. Journal of Electronics & Information Technology, 2024, 46(3): 1005–1012. doi: 10.11999/JEIT230260. -
下载: