Constraint Factors and Non-Compensatory Decision: A Multi-Agent Collaborative Method for Strong-Constraint Matching
-
摘要: 针对大语言模型在长文本强约束匹配任务中易出现条件遗漏和数值推理错误的问题,本文提出一种结合约束因子的多智能体协同匹配方法(FMAM)。该方法将约束条件和事实证据分别结构化为约束因子与事实因子,并以二者作为多阶段智能体之间的结构化信息传输载体。在本文构建的含5万个配对的政企匹配评测集上,FMAM的F1值达到90.17%,MCC为
0.8942 ,优于零样本LLM和少样本思维链等基线方法。消融实验结果显示,采用约束因子表示的方法在精确率、召回率和F1值上均优于自然语言中间表示方法。实验验证了在长文本匹配与复杂逻辑约束场景下,基于约束因子的结构化协同方式优于基于自然语言交互的多智能体模式,为强约束匹配任务中多智能体中间表示的选择提供了实证依据。Abstract:Objective Long-text strong-constraint matching is required in policy-to-enterprise recommendation, contract verification, and compliance screening. Categorical conditions, numerical thresholds, eligibility clauses, temporal scopes, and exclusion rules must be satisfied simultaneously, and violation of any hard constraint disqualifies a candidate. As large language models (LLMs) are applied to these decisions, attributes, comparison operators, and thresholds need to be preserved consistently over long contexts. Prevailing pipelines, including zero-shot prompting, retrieval-augmented generation (RAG), and chain-of-thought (CoT) reasoning, pass intermediate evidence mainly through semantic similarity or free-text steps, and therefore provide limited support for item-by-item verification of hard constraints. Open matching benchmarks likewise give limited coverage to continuous thresholds, exclusion clauses, and non-compensatory rules. A factorized multi-agent associative matching method, termed Factorized Multi-Agent Associative Matching (FMAM), is proposed. Constraint conditions and factual evidence are structured as constraint factors and fact factors, which serve as the information carriers among agents. Non-compensatory decision is realized by coupling hard-constraint gating with multiplicative aggregation. Methods FMAM separates factor extraction from deterministic evaluation ( Fig. 1 ). A Constraint Decomposition Agent (CDA) parses each source text once and writes every constraint into a field-level record, including the attribute, comparison operator, threshold, unit, type, and hard-constraint flag. A Feature Guidance Agent (FGA) collects attribute patterns from all source-side factors and extracts the corresponding fact factors from each candidate text once. The extraction is guided by these attribute patterns and is decoupled from pair-specific thresholds. The constraint-factor set of a source text is reused across all of its pairings, and the fact-factor set of a candidate text is reused across all of its pairings. For $ m $ source texts and $ n $ candidate texts, the number of long-text LLM calls is reduced from $ m\times n $ to $ m+n $, and the pairing stage requires no additional LLM calls. A Logic Evaluation Agent (LEA) binds compatible factors by attribute and type, normalizes units and temporal conditions, and computes a signed satisfaction margin for lower-bound, upper-bound, equality, set-inclusion, and set-exclusion operators. Identical factor records produce identical satisfaction scores. The margins are mapped to constraint-level satisfaction values, so that boundary cases remain distinguishable in the subsequent aggregation. A hard-constraint gate sets the matching score to zero when any necessary constraint is violated or the required evidence is absent. Satisfactions of all factors are aggregated by the geometric mean and multiplied by the gate. A Numerical Constraint Error Rate (NCER) is further defined to measure numerical-constraint extraction and comparison against the structured ground truth.Results and Discussions The method is evaluated on a constructed policy-enterprise matching set of 50,000 pairs ( Table 1 ), comprising 50 policy texts and 1,000 enterprise texts written in official-document and research-report styles. Field-level ground truth is determined from policy rules and enterprise indicators, and positive pairs account for 8.83% of the set. All methods use Gemini-2.5-Flash with the temperature set to 0. As shown inTable 2 andFig. 2 , FMAM obtains an accuracy of0.9835 , a precision of0.9523 , a recall of0.8562 , an F1-score of0.9017 , an AUC-PR of0.8567 , and an MCC of0.8942 , and achieves the highest accuracy, precision, F1-score, AUC-PR, and MCC among the compared methods. The F1-scores of few-shot CoT, Standard RAG, and GraphRAG are0.4559 ,0.1825 , and0.2452 , respectively. The NCER of FMAM is 4.87%, compared with 27.25% for the best baseline (Table 5 ). The average inference cost is 24.22 tokens per pair (Table 6 ), which is 1.27% of that of zero-shot LLM and 0.43% of that of ReAct. Standard RAG and GraphRAG consume 8.17 times and 28.83 times as many tokens as FMAM, respectively. Ablation studies show that replacing structured constraint factors with natural-language intermediate representations reduces the F1-score from0.9017 to0.3555 and raises the NCER from 4.87% to 77.57% (Table 4 ). Removing hard-constraint gating reduces the precision from0.9523 to0.7277 (Table 3 ).Conclusions Structured constraint factors combined with deterministic non-compensatory decision are effective for long-text strong-constraint matching. Field-level factors keep constraint information consistent across agents. Hard-constraint gating maintains a high precision under non-compensatory eligibility rules. Concentrating LLM calls in the extraction stage and performing pairing on structured factors reduces the inference cost from a multiplicative scale to a linear scale. Future work will extend the method to cross-institution texts and online calibration for broader compliance-screening applications. -
表 1 数据集统计信息
统计维度 统计项 数值 样本规模 政策样本数量 50 企业样本数量 1000 配对空间 政企候选配对总数 50,000 类别分布 正类匹配对数量 4,417 正类占比 8.83% 表 2 不同方法对比实验结果
评估方法 准确率 精确率 召回率 F1值 AUC-PR MCC FMAM 0.9835 0.9523 0.8562 0.9017 0.8567 0.8942 Standard RAG 0.9183 0.7876 0.1032 0.1825 0.1623 0.2666 LLM+Code 0.9432 0.7189 0.5855 0.6454 0.4636 0.6187 Zero-shot LLM 0.9543 0.7654 0.6957 0.7289 0.5826 0.7049 Few-shot CoT 0.8092 0.3047 0.9052 0.4559 0.4552 0.4548 ReAct 0.9370 0.6396 0.6582 0.6488 0.5363 0.6143 GraphRAG 0.8981 0.3545 0.1874 0.2452 0.1437 0.2076 表 3 决策机制消融结果
变体 准确率 精确率 召回率 F1值 FMAM 0.9835 0.9523 0.8562 0.9017 w/o Hard Gating 0.9590 0.7277 0.8562 0.7867 w/o Geometric Mean 0.9835 0.9523 0.8562 0.9017 w/o Both 0.8826 0.4293 0.9971 0.6002 表 4 中间表示消融结果
变体 精确率 召回率 F1值 NCER FMAM 0.9523 0.8562 0.9017 4.87% NL-Rep 0.2442 0.6534 0.3555 77.57% 表 5 数值约束审计结果
评估方法 NCER(%) 数值判定覆盖率(%) 已报告项错误率(%) FMAM 4.87 97.83 2.76 Standard RAG 94.13 7.37 20.33 LLM+Code 29.86 79.94 12.26 Zero-shot LLM 27.25 84.50 13.91 Few-shot CoT 33.47 69.61 4.43 ReAct 41.89 58.40 0.50 GraphRAG 89.95 10.21 1.58 表 6 Token用量分析结果
评估方法 输入Token 输出Token 总Token 平均Token/对 相对倍数 FMAM 1,025,616 185,567 1,211,183 24.22 1.00x Standard RAG 9,300,380 593,908 9,894,288 197.89 8.17x GraphRAG 34,208,380 711,817 34,920,197 698.40 28.83x Zero-shot LLM 89,823,800 5,441,482 95,265,282 1,905.31 78.65x LLM+Code 90,673,800 12,691,631 103,365,431 2,067.31 85.34x Few-shot CoT 106,723,800 11,040,077 117,763,877 2,355.28 97.23x ReAct 264,337,600 16,773,863 281,111,463 5,622.23 232.10x -
[1] GODBOLE A, GEORGE J G, and SHANDILYA S. Leveraging long-context large language models for multi-document understanding and summarization in enterprise applications[C]. Proceedings of the 1st International Conference on Business Intelligence, Computational Mathematics, and Data Analytics, Indore, India, 2024: 208–224. doi: 10.1007/978-3-031-87511-3_15. [2] HSU C C, WU I Z, and LIU S M. Decoding AI complexity: SHAP textual explanations via LLM for improved model transparency[C]. Proceedings of the 2024 International Conference on Consumer Electronics-Taiwan (ICCE-Taiwan), Taichung, China, 2024: 197–198. doi: 10.1109/ICCE-Taiwan62264.2024.10674465. [3] LEWIS P, PEREZ E, PIKTUS A, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks[C]. Proceedings of the 34th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2020: 793. [4] TRAN K T, DAO D, NGUYEN M D, et al. Multi-agent collaboration mechanisms: A survey of LLMs[EB/OL]. https://arxiv.org/abs/2501.06322, 2025. [5] LI Minghan, POPA D N, CHAGNON J, et al. The power of selecting key blocks with local pre-ranking for long document information retrieval[J]. ACM Transactions on Information Systems, 2023, 41(3): 73. doi: 10.1145/3568394. [6] ZHANG Ming, LU Jiyu, YANG Jiahao, et al. From coarse to fine: Enhancing multi-document summarization with multi-granularity relationship-based extractor[J]. Information Processing & Management, 2024, 61(3): 103696. doi: 10.1016/j.ipm.2024.103696. [7] BERTI L, GIORGI F, and KASNECI G. Emergent abilities in large language models: A survey[EB/OL]. https://arxiv.org/abs/2503.05788, 2025. [8] WEI J, WANG Xuezhi, SCHUURMANS D, et al. Chain-of-thought prompting elicits reasoning in large language models[C]. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 1800. [9] KOJIMA T, GU S S, REID M, et al. Large language models are zero-shot reasoners[C]. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, USA, 2022: 1613. [10] WANG Xuezhi, WEI J, SCHUURMANS D, et al. Self-consistency improves chain of thought reasoning in language models[C]. Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, 2023. [11] YAO Shunyu, YU Dian, ZHAO J, et al. Tree of thoughts: Deliberate problem solving with large language models[C]. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, 2023: 517. [12] BESTA M, BLACH N, KUBICEK A, et al. Graph of thoughts: Solving elaborate problems with large language models[C]. Proceedings of the 38th AAAI Conference on Artificial Intelligence, Vancouver, Canada, 2024: 17682–17690. doi: 10.1609/aaai.v38i16.29720. [13] YAO Shunyu, ZHAO J, YU Dian, et al. ReAct: Synergizing reasoning and acting in language models[C]. Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda, 2023. [14] GAO Luyu, MADAAN A, ZHOU Shuyan, et al. PAL: Program-aided language models[C]. Proceedings of the 40th International Conference on Machine Learning, Honolulu, USA, 2023: 435. [15] SCHICK T, DWIVEDI-YU J, DESSÌ R, et al. Toolformer: Language models can teach themselves to use tools[C]. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, 2023: 2997. [16] EDGE D, TRINH H, CHENG N, et al. From local to global: A graph RAG approach to query-focused summarization[EB/OL]. https://arxiv.org/abs/2404.16130, 2024. [17] YAN Shiqi, GU Jiachen, ZHU Yun, et al. Corrective retrieval augmented generation[EB/OL]. https://arxiv.org/abs/2401.15884, 2024. [18] ASAI A, WU Zeqiu, WANG Yizhong, et al. Self-RAG: Learning to retrieve, generate, and critique through self-reflection[C]. Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024. [19] JEONG S, BAEK J, CHO S, et al. Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity[C]. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, 2024: 7036–7050. doi: 10.18653/v1/2024.naacl-long.389. [20] RU Dongyu, QIU Lin, HU Xiangkun, et al. RAGChecker: A fine-grained framework for diagnosing retrieval-augmented generation[C]. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, Canada, 2024: 692. doi: 10.52202/079017-0692. [21] 李永斌, 刘楝, 郑杰. 一种面向特定信息领域的大模型命名实体识别方法[J]. 电子与信息学报, 2026, 48(2): 662–672. doi: 10.11999/JEIT250764.LI Yongbin, LIU Lian, and ZHENG Jie. A method for named entity recognition in military intelligence domain using large language models[J]. Journal of Electronics & Information Technology, 2026, 48(2): 662–672. doi: 10.11999/JEIT250764. [22] 徐祺坤, 刘娅汐, 韩淑娴, 等. 大语言模型文献挖掘驱动的网络指标体系与场景差异化分析[J]. 电子与信息学报, 2026, 48(7): 2865–2875. doi: 10.11999/JEIT251120.XU Qikun, LIU Yaxi, HAN Shuxian, et al. Network metric system and scenario-differentiated analysis driven by LLM literature mining[J]. Journal of Electronics & Information Technology, 2026, 48(7): 2865–2875. doi: 10.11999/JEIT251120. [23] LIU N F, LIN K, HEWITT J, et al. Lost in the middle: How language models use long contexts[J]. Transactions of the Association for Computational Linguistics, 2024, 12: 157–173. doi: 10.1162/tacl_a_00638. [24] LI Guohao, HAMMOUD H A A K, ITANI H, et al. CAMEL: Communicative agents for "mind" exploration of large language model society[C]. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, USA, 2023: 2264. [25] WU Qingyun, BANSAL G, ZHANG Jieyu, et al. AutoGen: Enabling next-gen LLM applications via multi-agent conversations[C]. Proceedings of the 4th Conference on Language Modeling, Philadelphia, USA, 2024. [26] HONG Sirui, ZHUGE Mingchen, CHEN J, et al. MetaGPT: Meta programming for a multi-agent collaborative framework[C]. Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024. [27] 夏维, 魏宏图, 程颖, 等. 面向卫星任务规划的专家链构建与优化方法[J]. 电子与信息学报, 2025, 47(12): 4986–4994. doi: 10.11999/JEIT251018.XIA Wei, WEI Hongtu, CHENG Ying, et al. An expert chain construction and optimization method for satellite mission planning[J]. Journal of Electronics & Information Technology, 2025, 47(12): 4986–4994. doi: 10.11999/JEIT251018. [28] BESOLD T R, D‘AVILA GARCEZ A, BADER S, et al. Neural-symbolic learning and reasoning: A survey and interpretation[M]. HITZLER P and SARKER M K. Neuro-Symbolic Artificial Intelligence: The State of the Art. Amsterdam: IOS Press, 2021: 1–51. doi: 10.3233/FAIA210348. [29] HAO Shibo, GU Yi, MA Haodi, et al. Reasoning with language model is planning with world model[C]. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, Singapore, 2023: 8154–8173. doi: 10.18653/v1/2023.emnlp-main.507. -
下载: