Decoupled Learning for Long-tailed Oracle Bone Character Recognition Based on Adaptive Difficulty Sampling
-
摘要: 针对甲骨文识别场景中普遍存在的类别长尾分布问题,以及现有方法难以兼顾头部类别判别力与尾部稀缺样本识别性能的技术瓶颈,该文提出一种两阶段解耦学习方法。第1阶段结合混合数据增强与标签分布感知间隔损失,学习全局泛化的特征空间并构建鲁棒决策边界;第2阶段冻结骨干网络,提出基于类别识别难度的自适应采样策略,结合类别加权损失优化分类器,通过动态分配采样权重聚焦尾部与困难类别,实现头部与尾部类别识别性能的协同提升。在OBC306公开数据集上的实验结果表明,该文方法总体识别准确率达94.34%,平均识别准确率达89.89%,综合性能优于主流方法,可为低资源古文字识别提供技术支撑。Abstract:
Objective Oracle Bone Character (OBC) recognition is challenged by an extreme long-tailed distribution and substantial intra-class variation. Conventional deep learning methods are often dominated by head classes, whereas existing approaches tend to overfit tail classes or fail to account for differences in learning difficulty across classes. To address these limitations, a two-stage decoupled learning framework is proposed to improve the recognition of tail and difficult classes while preserving the discriminative capability of head classes. Methods The proposed framework decouples feature representation learning from classifier optimization. In the first stage, the backbone network is trained using a mixed data augmentation strategy that combines CutMix and RandAugment with Label-Distribution-Aware Margin (LDAM) loss to learn robust feature representations and alleviate the effect of intra-class variation. In the second stage, the backbone network is frozen, and only the classifier is optimized. An adaptive difficulty sampling strategy is proposed to dynamically assign sampling weights according to historical and current class-level training difficulty. The classifier is further optimized using a Class-Balanced LDAM (CBL) loss, which combines class-balanced weighting with LDAM to refine decision boundaries for long-tailed classification. Results and Discussions Experiments on the highly imbalanced OBC306 dataset demonstrate that the proposed method achieves an overall accuracy of 94.34% and an average class accuracy of 89.89%. Compared with the Inception-v4 baseline, the proposed method improves the average class accuracy by 19.61%. Comparisons with representative long-tailed OBC recognition methods further demonstrate superior overall performance. Comprehensive ablation studies verify the effectiveness of the mixed data augmentation strategy and the adaptive difficulty sampling strategy in improving the recognition of rare and difficult characters. Parameter sensitivity analysis and qualitative error analysis further confirm the robustness and effectiveness of the proposed framework. Conclusions The proposed two-stage decoupled learning framework effectively addresses long-tailed OBC recognition by balancing the learning priorities of head, tail, and difficult classes. The mixed data augmentation strategy improves feature robustness, whereas the adaptive difficulty sampling strategy and the Class-Balanced LDAM loss jointly optimize classifier learning and refine decision boundaries without degrading head-class recognition performance. The proposed framework provides an effective solution for the digital recognition of Oracle Bone Characters and offers technical support for low-resource ancient character recognition. -
表 1 OBC306数据集上本文方法与其他甲骨文识别方法的对比
表 2 基于OBC306数据集的消融研究
CutMix RandAug DS 准确率(%) 总识别准确率 平均识别准确率 × × × 92.23 78.11 √ × × 94.40 81.85 × √ × 92.86 79.64 √ √ × 94.71 83.87 × × √ 93.01 80.57 × √ √ 94.60 82.21 √ √ √ 94.48 89.47 表 3 基于Oracle-AYNU数据集的消融研究
CutMix RandAug DS 准确率(%) 总识别准确率 平均识别准确率 × × × 77.32 65.14 √ × × 78.11 65.35 × √ × 75.32 63.29 √ √ × 76.89 64.06 × × √ 82.37 79.75 × √ √ 79.55 73.42 √ √ √ 80.56 75.48 表 4 比较不同$ {\alpha }_{c} $参数的性能
$ {\alpha }_{c} $参数 准确率(%) 总识别准确率 平均识别准确率 [0.1,0.1,0.1] 93.43 86.47 [0.5,0.5,0.5] 94.12 88.12 [0.9,0.9,0.9] 93.67 87.08 [0.1,0.5,0.9] 93.42 86.78 [0.9,0.5,0.1] 94.34 89.89 表 5 比较不同β参数的性能
β参数 准确率(%) 总识别准确率 平均识别准确率 0 94.96 86.02 0.9 94.90 87.00 0.99 94.80 88.03 0.999 94.48 89.47 0.9999 94.01 89.10 表 6 比较不同缩放因子s的性能
缩放因子s 准确率(%) 总识别准确率 平均识别准确率 5 94.48 89.47 10 94.60 88.79 20 94.81 88.77 30 94.17 89.11 表 7 比较不同类别间隔最大值的性能
类别间隔最大值 准确率(%) 总识别准确率 平均识别准确率 0.1 94.48 89.47 0.2 94.66 89.71 0.3 94.63 89.67 0.4 94.80 89.40 0.5 94.34 89.89 -
[1] 葛亮. 一百二十年来甲骨文材料的初步统计[J]. 汉字汉语研究, 2019(4): 33–54, 125. doi: 10.13513/j.cnki.41-1041/h.2019.04.006.GE Liang. Preliminary statistics of inscribed oracle bones excavated in the past 120 years[J]. The Study of Chinese Characters and Language, 2019(4): 33–54, 125. doi: 10.13513/j.cnki.41-1041/h.2019.04.006. [2] GUO Jun, WANG Changhu, ROMAN-RANGEL E, et al. Building hierarchical representations for oracle character and sketch recognition[J]. IEEE Transactions on Image Processing, 2016, 25(1): 104–118. doi: 10.1109/TIP.2015.2500019. [3] ZHANG Yikang, ZHANG Heng, LIU Yongge, et al. Oracle character recognition by nearest neighbor classification with deep metric learning[C]. 2019 International Conference on Document Analysis and Recognition, Sydney, Australia, 2019: 309–314. doi: 10.1109/ICDAR.2019.00057. [4] HUANG Shuangping, WANG Haobin, LIU Yongge, et al. OBC306: A large-scale oracle bone character recognition dataset[C]. International Conference on Document Analysis and Recognition, Sydney, Australia, 2019: 681–688. doi: 10.1109/ICDAR.2019.00114. [5] GUAN Haisu, WAN Jinpeng, LIU Yuliang, et al. An open dataset for the evolution of oracle bone characters: EVOBC[OL]. https://doi.org/10.48550/arXiv.2312.13631, 2024. [6] 韩佳艺, 刘建伟, 陈德华, 等. 深度长尾学习研究综述[J]. 自动化学报, 2025, 51(5): 985–1020. doi: 10.16383/j.aas.c240077.HAN Jiayi, LIU Jianwei, CHEN Dehua, et al. Survey on deep long-tailed learning[J]. Acta Automatica Sinica, 2025, 51(5): 985–1020. doi: 10.16383/j.aas.c240077. [7] LI Jing, WANG Qiufeng, ZHANG Rui, et al. Mix-up augmentation for oracle character recognition with imbalanced data distribution[C]. The 16th International Conference on Document Analysis and Recognition, Lausanne, Switzerland, 2021: 237–251. doi: 10.1007/978-3-030-86549-8_16. [8] HUANG Hongxiang, YANG Daihui, DAI Gang, et al. AGTGAN: Unpaired image translation for photographic ancient character generation[C]. The 30th ACM International Conference on Multimedia, Lisboa, Portugal, 2022: 5456–5467. doi: 10.1145/3503161.3548338. [9] LI Jing, WANG Qiufeng, WANG Siyuan, et al. Diff-Oracle: Deciphering oracle bone scripts with controllable diffusion model[OL]. https://doi.org/10.48550/arXiv.2312.13631, 2024. [10] LI Jing, DONG Bin, WANG Qiufeng, et al. Decoupled learning for long-tailed oracle character recognition[C]. The 17th International Conference on Document Analysis and Recognition, San José, USA, 2023: 165–181. doi: 10.1007/978-3-031-41685-9_11. [11] YUN S, HAN D, CHUN S, et al. CutMix: Regularization strategy to train strong classifiers with localizable features[C]. International Conference on Computer Vision, Seoul, South Korea, 2019: 6022–6031. doi: 10.1109/ICCV.2019.00612. [12] CUBUK E D, ZOPH B, SHLENS J, et al. Randaugment: Practical automated data augmentation with a reduced search space[C]. IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, USA, 2020: 3008–3017. doi: 10.1109/CVPRW50498.2020.00359. [13] LI Jing, CHI Xueke, WANG Qiufeng, et al. A comprehensive survey of oracle character recognition: Challenges, datasets, methodology, and beyond[J]. Pattern Recognition, 2026, 169: 111824. doi: 10.1016/j.patcog.2025.111824. [14] ZHOU Xinlun, HUA Xingcheng, and LI Feng. A method of Jia Gu Wen recognition based on a two-level classification[C]. 3rd International Conference on Document Analysis and Recognition, Montreal, Canada, 1995: 833–836. doi: 10.1109/ICDAR.1995.602030. [15] 栗青生, 杨玉星, 王爱民. 甲骨文识别的图同构方法[J]. 计算机工程与应用, 2011, 47(8): 112–114. doi: 10.3778/j.issn.1002-8331.2011.08.033.LI Qingsheng, YANG Yuxing, and WANG Aimin. Recognition of inscriptions on bones or tortoise shells based on graph isomorphism[J]. Computer Engineering and Applications, 2011, 47(8): 112–114. doi: 10.3778/j.issn.1002-8331.2011.08.033. [16] SZEGEDY C, IOFFE S, VANHOUCKE V, et al. Inception-v4, inception-ResNet and the impact of residual connections on learning[C]. The 31st AAAI Conference on Artificial Intelligence, San Francisco, USA, 2017: 4278–4284. doi: 10.1609/aaai.v31i1.11231. [17] MAI C, PENAVA P, and BUETTNER R. Oracle bone inscription character recognition based on a novel convolutional neural network architecture[J]. IEEE Access, 2024, 12: 197021–197034. doi: 10.1109/ACCESS.2024.3521319. [18] 毕晓君, 毛亚菲. 基于监督对比学习的小样本甲骨文字识别[J]. 智能系统学报, 2024, 19(1): 106–113. doi: 10.11992/tis.202309008.BI Xiaojun and MAO Yafei. Few-shot oracle bone character recognition based on supervised contrastive learning[J]. CAAI Transactions on Intelligent Systems, 2024, 19(1): 106–113. doi: 10.11992/tis.202309008. [19] 刘宗昊, 彭文杰, 代港, 等. 语义增强的零样本甲骨文字符识别[J]. 电子学报, 2024, 52(10): 3347–3358. doi: 10.12263/DZXB.20240286.LIU Zonghao, PENG Wenjie, DAI Gang, et al. Semantic-enhanced zero-shot oracle character recognition[J]. Acta Electronica Sinica, 2024, 52(10): 3347–3358. doi: 10.12263/DZXB.20240286. [20] WANG Wei, ZHANG Ting, ZHAO Yiwen, et al. Improving oracle bone characters recognition via a CycleGAN-based data augmentation method[C]. The 29th International Conference on Neural Information Processing, Virtual Event, 2022: 88–100. doi: 10.1007/978-981-99-1645-0_8. [21] LI Jing, WANG Qiufeng, HUANG Kaizhu, et al. Towards better long-tailed oracle character recognition with adversarial data augmentation[J]. Pattern Recognition, 2023, 140: 109534. doi: 10.1016/j.patcog.2023.109534. [22] CAO Kaidi, WEI C, GAIDON A, et al. Learning imbalanced datasets with label-distribution-aware margin loss[C]. The 33rd International Conference on Neural Information Processing Systems, Vancouver, Canada, 2019: 140. [23] YU Sihao, GUO Jiafeng, ZHANG Ruqing, et al. A re-balancing strategy for class-imbalanced classification based on instance difficulty[C]. Conference on Computer Vision and Pattern Recognition, New Orleans, USA, 2022: 70–79. doi: 10.1109/CVPR52688.2022.00017. [24] CUI Yin, JIA Menglin, LIN T Y, et al. Class-balanced loss based on effective number of samples[C]. Conference on Computer Vision and Pattern Recognition, Long Beach, USA, 2019: 9260–9269. doi: 10.1109/CVPR.2019.00949. [25] YU Weihao, ZHOU Pan, YAN Shuicheng, et al. InceptionNeXt: When inception meets ConvNeXt[C]. Conference on Computer Vision and Pattern Recognition, Seattle, USA, 2024: 5672–5683. doi: 10.1109/CVPR52733.2024.00542. -
下载: