Soft Error Sensitivity Analysis and Field-Hierarchical Fault-Tolerant Design for CSR-Based Sparse Neural Networks
-
摘要: 压缩稀疏行(Compressed Sparse Row, CSR)格式广泛用于稀疏神经网络加速器的权重存储,其values、col_indices和row_ptr分别参与数值计算、激活寻址和行边界控制,单粒子翻转(Single-Event Upset, SEU)引起的单比特错误因而可能产生不同的数值、地址和结构后果。为确定相应的字段保护范围,该文依据CSR执行过程和合法性约束建立故障分析方法,在翻转目标逻辑位后检查CSR合法性,并对可执行案例完成推理;根据分析结果,采用单错误纠正、双错误检测(Single-Error Correction and Double-Error Detection, SECDED)码比较不同保护范围与码字组织。LeNet-5/MNIST实验全量遍历
227264 个逻辑存储位,values案例均可执行,col_indices字段内地址失效率和row_ptr字段内结构失效率分别为5.83%和61.20%;VGG16/CIFAR-10分层定额样本中亦观察到相应的字段故障后果。在每个码字至多发生1位错误的条件下,同时保护row_ptr和col_indices可消除结构失效与地址失效,元素级和64 bit存储字级附加存储开销分别为31.35%和6.35%;覆盖三个字段后,固定可执行案例集合的故障后推理准确率与无故障基线一致。模块级独立布局布线结果显示,64 bit流水译码核将最长内部数据路径由9.38 ns缩短至6.49 ns。以CSR结构完整性和访问地址合法性为首要目标时,可优先保护row_ptr和col_indices。Abstract:Objective The compressed sparse row (CSR) format is widely used in sparse neural-network accelerators to reduce weight storage and support row-wise computation. Alongside nonzero weights, column indices and row boundaries are stored as metadata. These fields serve different computational functions and can produce distinct consequences following a Single-Event Upset (SEU). Weight errors perturb arithmetic operands, whereas metadata errors can redirect activation accesses or disrupt row boundaries. Uniform protection does not exploit these differences and adds redundancy to all three fields. A field-specific fault analysis and selective protection strategy are therefore developed to relate fault consequences to field functions and balance inference reliability against storage and decoding costs. Methods A software fault-injection framework is established using CSR execution rules and validity constraints. The values, col_indices, and row_ptr fields provide nonzero weights, column indices for activation access, and row boundaries, respectively ( Fig. 1 ). A single logical storage bit is flipped in each trial. A corrupted column index produces an erroneous connection if it remains within the input range; otherwise, an address failure is recorded. A corrupted internal row boundary produces an erroneous boundary if the range and nondecreasing-order constraints remain satisfied; otherwise, a structural failure is recorded. Inference is performed for executable cases, and the accuracy decrease is calculated as the fault-free accuracy minus the post-fault accuracy. The original stored value is restored after each trial to maintain independent injections (Fig. 2 ). This procedure evaluates the conditional consequences of single-bit errors in stored CSR data. The first and last row_ptr endpoints of each layer are treated as configuration constants. Only internal boundaries are included in the row_ptr fault and protection spaces. LeNet-5/MNIST and VGG16/CIFAR-10 models are trained, pruned, and quantized to 8-bit signed integer (INT8) weights. All 227,264 logical storage bits in the five LeNet-5 target layers are examined, comprising 111,784 bits in each of values and col_indices and 3,696 row_ptr bits. Layer-stratified quota sampling is applied to the 16 VGG16 target layers to examine larger sparse matrices and wider metadata fields. The sampled set contains 4,096 fault cases for values, 16,384 for col_indices, and 30,944 for row_ptr (Table 2 ). For fault comparison, each model is evaluated with and without injected faults on the same fixed set of 1,000 test images. Field-level failure proportions are calculated over the corresponding field cases. Protection is implemented using Single-Error Correction and Double-Error Detection (SECDED) codes. P0 is unprotected; P1 protects internal row_ptr boundaries; P2 additionally protects col_indices; and P3 protects all fields independently. PT jointly encodes values and col_indices while protecting row_ptr separately. Protection effects are calculated over the original LeNet-5 CSR data-bit fault space, assuming correct SECDED operation and at most one erroneous bit per codeword. Accuracy comparisons use the fixed set of 218,487 cases executable under P0. Element-level encoding maps 8-bit and 16-bit data to 13-bit and 22-bit codewords, respectively; memory-word encoding maps 64-bit data to 72-bit codewords. Storage overhead is calculated relative to each organization's unprotected data capacity. Codewords are generated during model loading or memory writes; decoding and correction are applied on the inference read path. Decoder functionality is checked using C and Register-Transfer Level (RTL) co-simulation against an independent bit-level reference model. Module-level out-of-context (OOC) placement and routing are performed on a Xilinx XC7Z020 device with a 10 ns reference clock period.Results and Discussions All 111,784 LeNet-5 fault cases in values remain executable, whereas address and structural failures account for 5.83% of col_indices cases and 61.20% of row_ptr cases, respectively. The maximum accuracy decrease is 0.70 percentage points (pp) for executable weight faults and 10.50 pp for executable boundary faults ( Table 3 ). This difference reflects how an altered boundary can reassign multiple consecutive nonzero entries between adjacent output rows. Across the five layers, the aggregate col_indices address-failure rate reaches 31.52% at bit 7. No address failure occurs in the first fully connected layer because its 256-dimensional input covers the entire unsigned 8-bit index range (Fig. 3 ). For row_ptr, flips at bits 8–15 always violate the structural constraints and contribute 81.70% of structural failures (Fig. 4 ). These distributions reflect the interaction between bit significance, valid index ranges, and adjacent row boundaries. In the VGG16 samples, address and structural failures account for 36.89% and 75.15% of the corresponding field cases. The same field-specific failure types thus occur with larger sparse matrices and wider metadata. Within the LeNet-5 protection comparison, P1 eliminates the 2,262 structural failures, and P2 also eliminates the 6,515 address failures. P2 provides the lowest storage overhead among the evaluated configurations that eliminate both failure types. Its additional storage costs are 31.35% for element-level encoding and 6.35% for 64-bit memory-word encoding. P3 and PT also correct weight errors, restoring baseline accuracy throughout the fixed executable-case set. Joint encoding reduces full-field element-level overhead from 62.09% for P3 to 37.50% for PT. Under 64-bit organization, both configurations require 12.50% overhead because each 64-bit data word carries eight check bits (Fig. 5 ). Post-route results show that P1, P2, P3, and PT require 72, 119, 166, and 147 Lookup Tables (LUTs), respectively (Table 5 ). The four protected element-level read-path modules have longest internal data paths of 6.50–6.86 ns and Worst Negative Slack (WNS) values of 3.11–3.44 ns. The evaluated internal register-to-register paths satisfy the setup-time constraint for the 10 ns reference period. The decoder branches operate in parallel, so their delays do not accumulate along a single path. For the 64-bit decoder, pipelining shortens the longest internal data path from 9.38 ns to 6.49 ns, a 30.8% reduction. Pipelining also increases the flip-flop count from 140 to 379. The pipelined decoder core has a High-Level Synthesis (HLS) scheduling latency of three cycles. Its initiation interval of one cycle allows it to accept one new codeword per cycle.Conclusions CSR fault consequences depend jointly on field function, bit position, and CSR validity constraints. Protecting row boundaries and column indices eliminates structural and address failures in the evaluated LeNet-5 fault space, while full-field SECDED protection restores baseline accuracy over the fixed executable-case set. Joint codewords reduce parity overhead for element-level organization, but provide no additional storage savings for the evaluated 64-bit organization. Selective protection reduces decoding logic relative to P3, and pipelining shortens internal paths at the cost of additional registers and latency. These results support matching field coverage and codeword organization to reliability requirements and hardware resources. -
表 1 CSR字段故障及其判定规则
注入字段 后果类型 判断条件 是否进入推理 values 数值扰动 权重编码改变,索引与边界保持合法 是 col_indices 错误连接 翻转后索引仍在当前层合法范围内 是 col_indices 地址失效 翻转后索引超出当前层输入范围 否 row_ptr 错误边界 完整边界序列仍满足式(2) 是 row_ptr 结构失效 内部边界超出范围或破坏非递减关系 否 表 2 实验模型、CSR规模与故障注入配置
模型 数据集 CSR目标
层数非零权重数 values/col_indices/
row_ptr位宽values/col_indices/
row_ptr评估案例数抽样方式 完整测试集
CSR准确率LeNet-5 MNIST 5 13973 8/8/16 bit 111784 /111784 /3696 全量遍历 99.05% VGG16 CIFAR-10 16 3005459 8/16/32 bit 4096 /16384 /30944 分层定额抽样 94.02% 表 3 LeNet-5与VGG16的CSR字段级故障结果
模型 字段 案例数 可执行案例数 地址失效数
(比例)结构失效数(比例) 准确率下降(pp) 平均 最大 LeNet-5 values 111784 111784 0(0%) 0(0%) $ 6.5\times {10}^{-3} $ 0.70 col_indices 111784 105269 6515 (5.83%)0(0%) $ 7.8\times {10}^{-3} $ 0.60 row_ptr 3696 1434 0(0%) 2262 (61.20%)0.12 10.50 VGG16 values 4096 4096 0(0%) 0(0%) 0.02 2.70 col_indices 16384 10340 6044 (36.89%)0(0%) $ 9.1\times {10}^{-3} $ 0.90 row_ptr 30944 7691 0(0%) 23253 (75.15%)0.12 82.60 表 4 本文保护配置与代表性神经网络存储ECC方法的比较
表 5 8/16 bit元素级SECDED读路径封装的OOC资源与时序结果
方案 并行译码支路 LUT FF 最长内部数据路径/ns 内部WNS/ns P0 — 0 72 1.11 8.77 P1 1×16 bit 72 78 6.71 3.22 P2 1×8 bit+1×16 bit 119 83 6.50 3.44 P3 2×8 bit+1×16 bit 166 88 6.82 3.20 PT 2×16 bit 147 84 6.86 3.11 -
[1] BOLCHINI C, CASSANO L, and MIELE A. Resilience of deep learning applications: A systematic literature review of analysis and hardening techniques[J]. Computer Science Review, 2024, 54: 100682. doi: 10.1016/j.cosrev.2024.100682. [2] GUAN Hui, NING Lin, LIN Zhen, et al. In-place zero-space memory protection for CNN[C]. Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, Canada, 2019: 515. [3] SANTOS F F D, PIMENTA P F, LUNARDI C, et al. Analyzing and increasing the reliability of convolutional neural networks on GPUs[J]. IEEE Transactions on Reliability, 2019, 68(2): 663–677. doi: 10.1109/TR.2018.2878387. [4] SUN Wenhao, ZOU Zhiwei, LIU Deng, et al. Bit-balance: Model-hardware codesign for accelerating NNs by exploiting bit-level sparsity[J]. IEEE Transactions on Computers, 2024, 73(1): 152–163. doi: 10.1109/TC.2023.3324477. [5] HAN Song, LIU Xingyu, MAO Huizi, et al. EIE: Efficient inference engine on compressed deep neural network[C]. Proceedings of the 43rd Annual International Symposium on Computer Architecture, Seoul, South Korea, 2016: 243–254. doi: 10.1109/ISCA.2016.30. [6] TROMMER E, WASCHNECK B, and KUMAR A. dCSR: A memory-efficient sparse matrix representation for parallel neural network inference[C]. Proceedings of 2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD), Munich, Germany, 2021: 1–9. doi: 10.1109/ICCAD51958.2021.9643506. [7] ZHANG Chen, GAO Shijie, DAI Guohao, et al. Fine-grained structured sparse computing for FPGA-based AI inference[J]. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2025, 44(7): 2544–2557. doi: 10.1109/TCAD.2024.3524356. [8] LEE J H, PARK B, KONG J, et al. Row-wise product-based sparse matrix multiplication hardware accelerator with optimal load balancing[J]. IEEE Access, 2022, 10: 64547–64559. doi: 10.1109/ACCESS.2022.3184116. [9] Li Baoting, Zhang Danqing, Zhao Pengfei, et al. DQ-STP: An efficient sparse on-device training processor based on low-rank decomposition and quantization for DNN[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2024, 71(4): 1665–1678. doi: 10.1109/TCSI.2024.3364093. [10] 闫爱斌, 李坤, 黄正峰, 等. 两种面向宇航应用的高可靠性抗辐射加固技术静态随机存储器单元[J]. 电子与信息学报, 2024, 46(10): 4072–4080. doi: 10.11999/JEIT240082.YAN Aibin, LI Kun, HUANG Zhengfeng, et al. Two highly reliable radiation hardened by design static random access memory cells for aerospace applications[J]. Journal of Electronics & Information Technology, 2024, 46(10): 4072–4080. doi: 10.11999/JEIT240082. [11] 柏娜, 李钢, 许耀华, 等. 应用于航空航天领域的低功耗多节点抗辐射静态随机存取存储器设计[J]. 电子与信息学报, 2025, 47(3): 850–858. doi: 10.11999/JEIT240294.BAI Na, LI Gang, XU Yaohua, et al. Low-power multi-node radiation-hardened SRAM design for aerospace applications[J]. Journal of Electronics & Information Technology, 2025, 47(3): 850–858. doi: 10.11999/JEIT240294. [12] 蔡烁, 帅威, 胡星, 等. 面向高速读写需求的宇航级抗辐射静态随机存储器加固单元设计[J]. 电子与信息学报, 2026, 48(5): 1894–1904. doi: 10.11999/JEIT251287.CAI Shuo, SHUAI Wei, HU Xing, et al. Design of an aerospace-grade radiation-hardened SRAM cell for high-speed read/write applications[J]. Journal of Electronics & Information Technology, 2026, 48(5): 1894–1904. doi: 10.11999/JEIT251287. [13] WANG Haibin, WANG Yangsheng, XIAO Jianhua, et al. Impact of single-event upsets on convolutional neural networks in Xilinx Zynq FPGAs[J]. IEEE Transactions on Nuclear Science, 2021, 68(4): 394–401. doi: 10.1109/TNS.2021.3062014. [14] MAHMOUD A, AGGARWAL N, NOBBE A, et al. PyTorchFI: A runtime perturbation tool for DNNs[C]. Proceedings of the 50th Annual IEEE/IFIP International Conference on Dependable Systems and Networks Workshops, Valencia, Spain, 2020: 25–31. doi: 10.1109/DSN-W50199.2020.00014. [15] MAHMOUD A, HARI S K S, FLETCHER C W, et al. Optimizing selective protection for CNN resilience[C]. Proceedings of the 32nd IEEE International Symposium on Software Reliability Engineering, Wuhan, China, 2021: 127–138. doi: 10.1109/ISSRE52982.2021.00025. [16] 陈子洋, 张萌, 张吉良. 一种星载在轨神经网络的容错设计方法[J]. 电子与信息学报, 2023, 45(9): 3234–3243. doi: 10.11999/JEIT230378.CHEN Ziyang, ZHANG Meng, and ZHANG Jiliang. A fault-tolerant design of spaceborne onboard neural network[J]. Journal of Electronics & Information Technology, 2023, 45(9): 3234–3243. doi: 10.11999/JEIT230378. [17] 张青, 刘成, 刘波, 等. 容错深度学习加速器跨层优化[J]. 计算机研究与发展, 2024, 61(6): 1370–1387. doi: 10.7544/issn1000-1239.202331005.ZHANG Qing, LIU Cheng, LIU Bo, et al. Cross-layer optimization for fault-tolerant deep learning accelerators[J]. Journal of Computer Research and Development, 2024, 61(6): 1370–1387. doi: 10.7544/issn1000-1239.202331005. [18] LEE S S and YANG J S. Value-aware parity insertion ECC for fault-tolerant deep neural network[C]. Proceedings of the Design, Automation & Test in Europe Conference & Exhibition, Antwerp, Belgium, 2022: 724–729. doi: 10.23919/DATE54114.2022.9774543. [19] 柳姗姗, 金辉, 刘思佳, 等. 面向投票类AI分类器的零冗余存储器容错设计[J]. 集成电路与嵌入式系统, 2024, 24(6): 1–8.LIU Shanshan, JIN Hui, LIU Sijia, et al. Redundancy-free error-tolerant memory design for voting-based AI classifiers[J]. Integrated Circuits and Embedded Systems, 2024, 24(6): 1–8. [20] TRAIOLA M, KRITIKAKOU A, and SENTIEYS O. harDNNing: A machine-learning-based framework for fault tolerance assessment and protection of DNNs[C]. Proceedings of the 28th IEEE European Test Symposium, Venice, Italy, 2023: 1–6. doi: 10.1109/ETS56758.2023.10174178. [21] ZHAO Kai, DI Sheng, LI Sihuan, et al. FT-CNN: Algorithm-based fault tolerance for convolutional neural networks[J]. IEEE Transactions on Parallel and Distributed Systems, 2021, 32(7): 1677–1689. doi: 10.1109/TPDS.2020.3043449. [22] HARI S K S, SULLIVAN M B, TSAI T, et al. Making convolutions resilient via algorithm-based error detection techniques[J]. IEEE Transactions on Dependable and Secure Computing, 2022, 19(4): 2546–2558. doi: 10.1109/TDSC.2021.3063083. [23] GAO Zhen, QI Yanmao, SHI Jinchang, et al. Detect and replace: Efficient soft error protection of FPGA-based CNN accelerators[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2025, 33(1): 66–74. doi: 10.1109/tvlsi.2024.3443834. [24] LECUN Y, BOTTOU L, BENGIO Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11): 2278–2324. doi: 10.1109/5.726791. [25] SIMONYAN K and ZISSERMAN A. Very deep convolutional networks for large-scale image recognition[C]. Proceedings of the 3rd International Conference on Learning Representations, San Diego, USA, 2015. [26] KRIZHEVSKY A. Learning multiple layers of features from tiny images[R]. Toronto: University of Toronto, 2009. [27] GOLNARI P A and MALIK S. Evaluating matrix representations for error-tolerant computing[C]. Proceedings of the Design, Automation & Test in Europe Conference & Exhibition (DATE), Lausanne, Switzerland, 2017: 1659–1662. doi: 10.23919/DATE.2017.7927260. [28] AHMED S T, HEMARAM S, and TAHOORI M B. NN-ECC: Embedding error correction codes in neural network weight memories using multi-task learning[C]. Proceedings of 2024 IEEE 42nd VLSI Test Symposium (VTS), Tempe, USA, 2024: 1–7, doi: 10.1109/VTS60656.2024.10538886. [29] PARK T, GORGIN S, KIM D, et al. PoP-ECC: Robust and flexible error correction against multi-bit upsets in DNN accelerators[C]. Proceedings of the 62nd ACM/IEEE Design Automation Conference, San Francisco, USA, 2025: 1–7. doi: 10.1109/DAC63849.2025.11133373. [30] JO M J and LEE Y S. Stegano-ECC: Enhancing DNN fault tolerance with embedded parity for important bits[J]. Journal of Systems Architecture, 2026, 171: 103651. doi: 10.1016/j.sysarc.2025.103651. -
下载: