高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

CSR稀疏神经网络软错误敏感性分析与字段分级容错设计

王晨旭 杨善强 沈乐翔 李剑锋 徐天亮 吕振斌

王晨旭, 杨善强, 沈乐翔, 李剑锋, 徐天亮, 吕振斌. CSR稀疏神经网络软错误敏感性分析与字段分级容错设计[J]. 电子与信息学报. doi: 10.11999/JEIT260621
引用本文: 王晨旭, 杨善强, 沈乐翔, 李剑锋, 徐天亮, 吕振斌. CSR稀疏神经网络软错误敏感性分析与字段分级容错设计[J]. 电子与信息学报. doi: 10.11999/JEIT260621
WANG Chenxu, YANG Shanqiang, SHEN Lexiang, LI Jianfeng, XU Tianliang, LV Zhenbin. Soft Error Sensitivity Analysis and Field-Hierarchical Fault-Tolerant Design for CSR-Based Sparse Neural Networks[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260621
Citation: WANG Chenxu, YANG Shanqiang, SHEN Lexiang, LI Jianfeng, XU Tianliang, LV Zhenbin. Soft Error Sensitivity Analysis and Field-Hierarchical Fault-Tolerant Design for CSR-Based Sparse Neural Networks[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260621

CSR稀疏神经网络软错误敏感性分析与字段分级容错设计

doi: 10.11999/JEIT260621 cstr: 32379.14.JEIT260621
基金项目: 山东省重点研发计划(重大科技创新工程)项目(2022ZLGX04),国家自然科学基金(U2106202),山东省自然科学基金项目(ZR2023MA074)
详细信息
    作者简介:

    王晨旭:男,教授,研究方向为超大规模集成电路设计与安全加固、容错设计

    杨善强:男,博士生,研究方向为集成电路可靠性设计、抗辐照加固设计

    沈乐翔:男,本科生,研究方向为硬件加速器设计

    李剑锋:男,副教授,研究方向多模态遥感图像处理、边缘计算与安全加固

    徐天亮:男,助理研究员,研究方向为人工智能与海洋遥感数据处理、集成电路设计与硬件加速

    吕振斌:男,博士生,研究方向为集成电路硬件安全技术、忆阻器设计及应用

    通讯作者:

    王晨旭 wangchenxu@hit.edu.cn

  • 中图分类号: TN406

Soft Error Sensitivity Analysis and Field-Hierarchical Fault-Tolerant Design for CSR-Based Sparse Neural Networks

Funds: Major scientific and technological innovation projects of Shandong Province of China (No.2022ZLGX04), The National Natural Foundation of China (U2106202), Shandong Provincial Natural Science Foundation (ZR2023MA074)
  • 摘要: 压缩稀疏行(Compressed Sparse Row, CSR)格式广泛用于稀疏神经网络加速器的权重存储,其values、col_indices和row_ptr分别参与数值计算、激活寻址和行边界控制,单粒子翻转(Single-Event Upset, SEU)引起的单比特错误因而可能产生不同的数值、地址和结构后果。为确定相应的字段保护范围,该文依据CSR执行过程和合法性约束建立故障分析方法,在翻转目标逻辑位后检查CSR合法性,并对可执行案例完成推理;根据分析结果,采用单错误纠正、双错误检测(Single-Error Correction and Double-Error Detection, SECDED)码比较不同保护范围与码字组织。LeNet-5/MNIST实验全量遍历227264个逻辑存储位,values案例均可执行,col_indices字段内地址失效率和row_ptr字段内结构失效率分别为5.83%和61.20%;VGG16/CIFAR-10分层定额样本中亦观察到相应的字段故障后果。在每个码字至多发生1位错误的条件下,同时保护row_ptr和col_indices可消除结构失效与地址失效,元素级和64 bit存储字级附加存储开销分别为31.35%和6.35%;覆盖三个字段后,固定可执行案例集合的故障后推理准确率与无故障基线一致。模块级独立布局布线结果显示,64 bit流水译码核将最长内部数据路径由9.38 ns缩短至6.49 ns。以CSR结构完整性和访问地址合法性为首要目标时,可优先保护row_ptr和col_indices。
  • 图  1  CSR三个存储字段的执行通路与单比特翻转后果

    图  2  CSR存储字段单比特故障案例的评估流程

    图  3  LeNet-5中col_indices单比特翻转的地址失效分布

    图  4  LeNet-5中row_ptr内部边界单比特翻转的结构失效分布

    图  5  不同CSR存储保护配置的附加存储开销与故障影响

    表  1  CSR字段故障及其判定规则

    注入字段后果类型判断条件是否进入推理
    values数值扰动权重编码改变,索引与边界保持合法
    col_indices错误连接翻转后索引仍在当前层合法范围内
    col_indices地址失效翻转后索引超出当前层输入范围
    row_ptr错误边界完整边界序列仍满足式(2)
    row_ptr结构失效内部边界超出范围或破坏非递减关系
    下载: 导出CSV

    表  2  实验模型、CSR规模与故障注入配置

    模型数据集CSR目标
    层数
    非零权重数values/col_indices/
    row_ptr位宽
    values/col_indices/
    row_ptr评估案例数
    抽样方式完整测试集
    CSR准确率
    LeNet-5MNIST5139738/8/16 bit111784/111784/3696全量遍历99.05%
    VGG16CIFAR-101630054598/16/32 bit4096/16384/30944分层定额抽样94.02%
    下载: 导出CSV

    表  3  LeNet-5与VGG16的CSR字段级故障结果

    模型字段案例数可执行案例数地址失效数
    (比例)
    结构失效数(比例)准确率下降(pp)
    平均最大
    LeNet-5values1117841117840(0%)0(0%)$ 6.5\times {10}^{-3} $0.70
    col_indices1117841052696515(5.83%)0(0%)$ 7.8\times {10}^{-3} $0.60
    row_ptr369614340(0%)2262(61.20%)0.1210.50
    VGG16values409640960(0%)0(0%)0.022.70
    col_indices16384103406044(36.89%)0(0%)$ 9.1\times {10}^{-3} $0.90
    row_ptr3094476910(0%)23253(75.15%)0.1282.60
    下载: 导出CSV

    表  4  本文保护配置与代表性神经网络存储ECC方法的比较

    方法保护对象主要故障条件保护范围及选择依据ECC或码字组织
    全字段SECDED基线(P3)CSR三个字段单比特翻转全字段统一覆盖字段独立SECDED码字
    NN-ECC[28]权重参数权重存储故障训练约束线性分组码嵌入权重
    PoP-ECC[29]加速器存储数据多比特翻转多比特错误模式两级ECC
    Stegano-ECC[30]权重重要位权重位错误位位置与数据类型重要位采用SEC,校验信息嵌入低重要性位
    本文P1—P3及PTCSR三个字段单比特翻转字段功能及故障后果P1—P3字段独立;PT联合values与col_indices
    下载: 导出CSV

    表  5  8/16 bit元素级SECDED读路径封装的OOC资源与时序结果

    方案并行译码支路LUTFF最长内部数据路径/ns内部WNS/ns
    P00721.118.77
    P11×16 bit72786.713.22
    P21×8 bit+1×16 bit119836.503.44
    P32×8 bit+1×16 bit166886.823.20
    PT2×16 bit147846.863.11
    下载: 导出CSV
  • [1] BOLCHINI C, CASSANO L, and MIELE A. Resilience of deep learning applications: A systematic literature review of analysis and hardening techniques[J]. Computer Science Review, 2024, 54: 100682. doi: 10.1016/j.cosrev.2024.100682.
    [2] GUAN Hui, NING Lin, LIN Zhen, et al. In-place zero-space memory protection for CNN[C]. Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, Canada, 2019: 515.
    [3] SANTOS F F D, PIMENTA P F, LUNARDI C, et al. Analyzing and increasing the reliability of convolutional neural networks on GPUs[J]. IEEE Transactions on Reliability, 2019, 68(2): 663–677. doi: 10.1109/TR.2018.2878387.
    [4] SUN Wenhao, ZOU Zhiwei, LIU Deng, et al. Bit-balance: Model-hardware codesign for accelerating NNs by exploiting bit-level sparsity[J]. IEEE Transactions on Computers, 2024, 73(1): 152–163. doi: 10.1109/TC.2023.3324477.
    [5] HAN Song, LIU Xingyu, MAO Huizi, et al. EIE: Efficient inference engine on compressed deep neural network[C]. Proceedings of the 43rd Annual International Symposium on Computer Architecture, Seoul, South Korea, 2016: 243–254. doi: 10.1109/ISCA.2016.30.
    [6] TROMMER E, WASCHNECK B, and KUMAR A. dCSR: A memory-efficient sparse matrix representation for parallel neural network inference[C]. Proceedings of 2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD), Munich, Germany, 2021: 1–9. doi: 10.1109/ICCAD51958.2021.9643506.
    [7] ZHANG Chen, GAO Shijie, DAI Guohao, et al. Fine-grained structured sparse computing for FPGA-based AI inference[J]. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2025, 44(7): 2544–2557. doi: 10.1109/TCAD.2024.3524356.
    [8] LEE J H, PARK B, KONG J, et al. Row-wise product-based sparse matrix multiplication hardware accelerator with optimal load balancing[J]. IEEE Access, 2022, 10: 64547–64559. doi: 10.1109/ACCESS.2022.3184116.
    [9] Li Baoting, Zhang Danqing, Zhao Pengfei, et al. DQ-STP: An efficient sparse on-device training processor based on low-rank decomposition and quantization for DNN[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2024, 71(4): 1665–1678. doi: 10.1109/TCSI.2024.3364093.
    [10] 闫爱斌, 李坤, 黄正峰, 等. 两种面向宇航应用的高可靠性抗辐射加固技术静态随机存储器单元[J]. 电子与信息学报, 2024, 46(10): 4072–4080. doi: 10.11999/JEIT240082.

    YAN Aibin, LI Kun, HUANG Zhengfeng, et al. Two highly reliable radiation hardened by design static random access memory cells for aerospace applications[J]. Journal of Electronics & Information Technology, 2024, 46(10): 4072–4080. doi: 10.11999/JEIT240082.
    [11] 柏娜, 李钢, 许耀华, 等. 应用于航空航天领域的低功耗多节点抗辐射静态随机存取存储器设计[J]. 电子与信息学报, 2025, 47(3): 850–858. doi: 10.11999/JEIT240294.

    BAI Na, LI Gang, XU Yaohua, et al. Low-power multi-node radiation-hardened SRAM design for aerospace applications[J]. Journal of Electronics & Information Technology, 2025, 47(3): 850–858. doi: 10.11999/JEIT240294.
    [12] 蔡烁, 帅威, 胡星, 等. 面向高速读写需求的宇航级抗辐射静态随机存储器加固单元设计[J]. 电子与信息学报, 2026, 48(5): 1894–1904. doi: 10.11999/JEIT251287.

    CAI Shuo, SHUAI Wei, HU Xing, et al. Design of an aerospace-grade radiation-hardened SRAM cell for high-speed read/write applications[J]. Journal of Electronics & Information Technology, 2026, 48(5): 1894–1904. doi: 10.11999/JEIT251287.
    [13] WANG Haibin, WANG Yangsheng, XIAO Jianhua, et al. Impact of single-event upsets on convolutional neural networks in Xilinx Zynq FPGAs[J]. IEEE Transactions on Nuclear Science, 2021, 68(4): 394–401. doi: 10.1109/TNS.2021.3062014.
    [14] MAHMOUD A, AGGARWAL N, NOBBE A, et al. PyTorchFI: A runtime perturbation tool for DNNs[C]. Proceedings of the 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops, Valencia, Spain, 2020: 25–31. doi: 10.1109/DSN-W50199.2020.00014.
    [15] MAHMOUD A, HARI S K S, FLETCHER C W, et al. Optimizing selective protection for CNN resilience[C]. Proceedings of the 32nd IEEE International Symposium on Software Reliability Engineering, Wuhan, China, 2021: 127–138. doi: 10.1109/ISSRE52982.2021.00025.
    [16] 陈子洋, 张萌, 张吉良. 一种星载在轨神经网络的容错设计方法[J]. 电子与信息学报, 2023, 45(9): 3234–3243. doi: 10.11999/JEIT230378.

    CHEN Ziyang, ZHANG Meng, and ZHANG Jiliang. A fault-tolerant design of spaceborne onboard neural network[J]. Journal of Electronics & Information Technology, 2023, 45(9): 3234–3243. doi: 10.11999/JEIT230378.
    [17] 张青, 刘成, 刘波, 等. 容错深度学习加速器跨层优化[J]. 计算机研究与发展, 2024, 61(6): 1370–1387. doi: 10.7544/issn1000-1239.202331005.

    ZHANG Qing, LIU Cheng, LIU Bo, et al. Cross-layer optimization for fault-tolerant deep learning accelerators[J]. Journal of Computer Research and Development, 2024, 61(6): 1370–1387. doi: 10.7544/issn1000-1239.202331005.
    [18] LEE S S and YANG J S. Value-aware parity insertion ECC for fault-tolerant deep neural network[C]. Proceedings of the Design, Automation & Test in Europe Conference & Exhibition, Antwerp, Belgium, 2022: 724–729. doi: 10.23919/DATE54114.2022.9774543.
    [19] 柳姗姗, 金辉, 刘思佳, 等. 面向投票类AI分类器的零冗余存储器容错设计[J]. 集成电路与嵌入式系统, 2024, 24(6): 1–8.

    LIU Shanshan, JIN Hui, LIU Sijia, et al. Redundancy-free error-tolerant memory design for voting-based AI classifiers[J]. Integrated Circuits and Embedded Systems, 2024, 24(6): 1–8.
    [20] TRAIOLA M, KRITIKAKOU A, and SENTIEYS O. harDNNing: A machine-learning-based framework for fault tolerance assessment and protection of DNNs[C]. Proceedings of the 28th IEEE European Test Symposium, Venice, Italy, 2023: 1–6. doi: 10.1109/ETS56758.2023.10174178.
    [21] ZHAO Kai, DI Sheng, LI Sihuan, et al. FT-CNN: Algorithm-based fault tolerance for convolutional neural networks[J]. IEEE Transactions on Parallel and Distributed Systems, 2021, 32(7): 1677–1689. doi: 10.1109/TPDS.2020.3043449.
    [22] HARI S K S, SULLIVAN M B, TSAI T, et al. Making convolutions resilient via algorithm-based error detection techniques[J]. IEEE Transactions on Dependable and Secure Computing, 2022, 19(4): 2546–2558. doi: 10.1109/TDSC.2021.3063083.
    [23] GAO Zhen, QI Yanmao, SHI Jinchang, et al. Detect and replace: Efficient soft error protection of FPGA-based CNN accelerators[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2025, 33(1): 66–74. doi: 10.1109/tvlsi.2024.3443834.
    [24] LECUN Y, BOTTOU L, BENGIO Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11): 2278–2324. doi: 10.1109/5.726791.
    [25] SIMONYAN K and ZISSERMAN A. Very deep convolutional networks for large-scale image recognition[C]. Proceedings of the 3rd International Conference on Learning Representations, San Diego, USA, 2015.
    [26] KRIZHEVSKY A. Learning multiple layers of features from tiny images[R]. Toronto: University of Toronto, 2009.
    [27] GOLNARI P A and MALIK S. Evaluating matrix representations for error-tolerant computing[C]. Proceedings of the Design, Automation & Test in Europe Conference & Exhibition (DATE), Lausanne, Switzerland, 2017: 1659–1662. doi: 10.23919/DATE.2017.7927260.
    [28] AHMED S T, HEMARAM S, and TAHOORI M B. NN-ECC: Embedding error correction codes in neural network weight memories using multi-task learning[C]. Proceedings of 2024 IEEE 42nd VLSI Test Symposium (VTS), Tempe, USA, 2024: 1–7, doi: 10.1109/VTS60656.2024.10538886.
    [29] PARK T, GORGIN S, KIM D, et al. PoP-ECC: Robust and flexible error correction against multi-bit upsets in DNN accelerators[C]. Proceedings of the 62nd ACM/IEEE Design Automation Conference, San Francisco, USA, 2025: 1–7. doi: 10.1109/DAC63849.2025.11133373.
    [30] JO M J and LEE Y S. Stegano-ECC: Enhancing DNN fault tolerance with embedded parity for important bits[J]. Journal of Systems Architecture, 2026, 171: 103651. doi: 10.1016/j.sysarc.2025.103651.
  • 加载中
图(5) / 表(5)
计量
  • 文章访问数:  26
  • HTML全文浏览量:  4
  • PDF下载量:  1
  • 被引次数: 0
出版历程
  • 收稿日期:  2026-05-18
  • 修回日期:  2026-09-12
  • 录用日期:  2026-09-14
  • 网络出版日期:  2026-09-21

目录

    /

    返回文章
    返回