| Citation: | ZHOU Tong, CHEN Hongzhi, XU Jialong, SUN Peng, JIANG Dajie, LIU Jiankang. Research on AI-Enabled Real-Time Audio Joint Source-Channel Coding[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260379 |
| [1] |
周通, 姜大洁, 谭俊杰, 等. 语义通信的用例、挑战与标准化影响浅析[J]. 移动通信, 2025, 49(7): 2–11. doi: 10.3969/j.issn.1006-1010.20250522-0001.
ZHOU Tong, JIANG Dajie, TAN Junjie, et al. A survey of semantic communication: Use cases, open challenges, and standardization trends[J]. Mobile Communications, 2025, 49(7): 2–11. doi: 10.3969/j.issn.1006-1010.20250522-0001.
|
| [2] |
OPPO. 基于信源信道联合编码的CSI反馈[R]. IMT-2030-Semantic_2024004, 2024. (查阅网上资料, 未找到本条文献信息, 请确认).
|
| [3] |
OPPO. 信道与无线环境语义通信_基于联合信源信道编码的CSI反馈[R]. IMT-2030-Semantic_2024036, 2024. (查阅网上资料, 未找到本条文献信息, 请确认).
|
| [4] |
小米. 面向CSI反馈的信道估计与联合信源信道编码的联合设计[R]. IMT-2030-Semantic_2024056, 2024. (查阅网上资料, 未找到本条文献信息, 请确认).
|
| [5] |
YANG Ke, WANG Sixian, DAI Jincheng, et al. SwinJSCC: Taming swin transformer for deep joint source-channel coding[J]. IEEE Transactions on Cognitive Communications and Networking, 2025, 11(1): 90–104. doi: 10.1109/TCCN.2024.3424842.
|
| [6] |
GUO Jiangyuan, CHEN Wei, SUN Yuxuan, et al. VideoQA-SC: Adaptive semantic communication for video question answering[J]. IEEE Journal on Selected Areas in Communications, 2025, 43(7): 2462–2477. doi: 10.1109/JSAC.2025.3559160.
|
| [7] |
TUNG T Y and GÜNDÜZ D. DeepWiVe: Deep-learning-aided wireless video transmission[J]. IEEE Journal on Selected Areas in Communications, 2022, 40(9): 2570–2583. doi: 10.1109/JSAC.2022.3191354.
|
| [8] |
BOURTSOULATZE E, KURKA D B, and GÜNDÜZ D. Deep joint source-channel coding for wireless image transmission[J]. IEEE Transactions on Cognitive Communications and Networking, 2019, 5(3): 567–579. doi: 10.1109/TCCN.2019.2919300.
|
| [9] |
YILMAZ S F, NIU Xueyan, BAI Bo, et al. High perceptual quality wireless image delivery with denoising diffusion models[C]. IEEE Conference on Computer Communications Workshops, Vancouver, Canada, 2024: 1–5. doi: 10.1109/INFOCOMWKSHPS61880.2024.10620904.
|
| [10] |
DAI Jincheng, WANG Sixian, TAN Kailin, et al. Nonlinear transform source-channel coding for semantic communications[J]. IEEE Journal on Selected Areas in Communications, 2022, 40(8): 2300–2316. doi: 10.1109/JSAC.2022.3180802.
|
| [11] |
WENG Zhenzi, QIN Zhijin, and LI G Y. Robust semantic communications for speech transmission[C]. 2025 IEEE International Conference on Acoustics, Speech and Signal Processing, Hyderabad, India, 2025: 1–5. doi: 10.1109/ICASSP49660.2025.10889160. (查阅网上资料,未能确认本条文献修改是否正确,请确认).
|
| [12] |
HAN Tianxiao, YANG Qianqian, SHI Zhiguo, et al. Semantic-preserved communication system for highly efficient speech transmission[J]. IEEE Journal on Selected Areas in Communications, 2023, 41(1): 245–259. doi: 10.1109/JSAC.2022.3221952.
|
| [13] |
WENG Zhenzi, QIN Zhijin, TAO Xiaoming, et al. Deep learning enabled semantic communications with speech recognition and synthesis[J]. IEEE Transactions on Wireless Communications, 2023, 22(9): 6227–6240. doi: 10.1109/TWC.2023.3240969.
|
| [14] |
CHEN Xiaojiao, WANG Jing, XU Liang, et al. A perceptually motivated approach for low-complexity speech semantic communication[J]. IEEE Internet of Things Journal, 2024, 11(12): 22054–22065. doi: 10.1109/JIOT.2024.3378779.
|
| [15] |
3GPP. 3GPP TS 23.501 System architecture for the 5G System (5GS)[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认).
|
| [16] |
3GPP. 3GPP TS 38.323 Packet data convergence protocol (PDCP) specification[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认).
|
| [17] |
IMT-2030(6G)推进组. 面向6G的语义通信技术研究报告[M]. 北京: IMT-2030(6G)推进组, 2025. (查阅网上资料, 未找到对应的英文翻译, 请补充).
|
| [18] |
Descript. Descript-audio-codec[EB/OL]. https://github.com/descriptinc/descript-audio-codec, 2026.
|
| [19] |
3GPP. 3GPP TR 26.940 Study on ultra low bit rate speech codecs[S]. 2026. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认).
|
| [20] |
3GPP. 3GPP TS 36. 212 Multiplexing and channel coding[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认).
|
| [21] |
ETSI. ETSI TS 136 213 V19.1. 0 (2025-10) Physical layer procedures[S]. 2025. (查阅网上资料, 未找到出版信息, 请补充)(查阅网上资料, 未能确认本条文献修改是否正确, 请确认).
|
| [22] |
OpenDataLab. VCTK[EB/OL]. https://opendatalab.com/OpenDataLab/VCTK, 2026.
|
| [23] |
Common voice[EB/OL]. https://commonvoice.mozilla.org/zh-CN, 2026.(查阅网上资料,请补充作者信息).
|
| [24] |
Librispeech[EB/OL]. https://huggingface.co/datasets/openslr/librispeech_asr, 2026. (查阅网上资料,请补充作者信息).
|
| [25] |
ITU. ITU-T P. 863-2018 Perceptual objective listening quality prediction[S]. 2018. (查阅网上资料, 未找到出版信息, 请补充).
|