高级搜索

留言板

尊敬的读者、作者、审稿人, 关于本刊的投稿、审稿、编辑和出版的任何问题, 您可以本页添加留言。我们将尽快给您答复。谢谢您的支持!

姓名
邮箱
手机号码
标题
留言内容
验证码

CuSyn: 一种面向多用户、多次数据发布的可定制差分隐私数据合成框架

谭畅,  聂力海,  赵文宇,  李同

谭畅, 聂力海, 赵文宇, 李同. CuSyn: 一种面向多用户、多次数据发布的可定制差分隐私数据合成框架[J]. 电子与信息学报. doi: 10.11999/JEIT260503
引用本文: 谭畅, 聂力海, 赵文宇, 李同. CuSyn: 一种面向多用户、多次数据发布的可定制差分隐私数据合成框架[J]. 电子与信息学报. doi: 10.11999/JEIT260503
TAN Chang, NIE Lihai, ZHAO Wenyu, LI Tong. CuSyn: A Customizable Differential Privacy Data Synthesis Framework for Multi-User and Multiple Data Publishing Scenarios[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260503
Citation: TAN Chang, NIE Lihai, ZHAO Wenyu, LI Tong. CuSyn: A Customizable Differential Privacy Data Synthesis Framework for Multi-User and Multiple Data Publishing Scenarios[J]. Journal of Electronics & Information Technology. doi: 10.11999/JEIT260503

CuSyn: 一种面向多用户、多次数据发布的可定制差分隐私数据合成框架

doi: 10.11999/JEIT260503 cstr: 32379.14.JEIT260503
基金项目: 国家自然科学基金 (62272251, 62402248)
详细信息
    作者简介:

    谭畅:男,博士生在读,研究方向为差分隐私技术及其应用,邮箱 melonsistan@mail.nankai.edu.cn

    聂力海:男,副教授,研究方向为网络安全、机器学习

    赵文宇:男,高级工程师,研究方向为信息安全、数据安全、软件工程化等

    李同:男,副教授,研究方向为数据隐私保护、安全外包计算、安全机器学习,邮箱 tongli@nankai.edu.cn

    通讯作者:

    李同 tongli@nankai.edu.cn

  • 中图分类号: TP309.2; TP311.13

CuSyn: A Customizable Differential Privacy Data Synthesis Framework for Multi-User and Multiple Data Publishing Scenarios

Funds: National Natural Science Foundation of China (62272251, 62402248)
  • 摘要: 尽管现有许多隐私保护数据合成机制能够生成高可用性的合成数据集用于发布,但它们通常缺乏对多用户、多次发布场景下细粒度隐私保障的专门设计。这一局限性主要体现在两方面:一是缺乏系统性的隐私预算跟踪与管理方法,二是难以根据不同用户的具体分析需求进行定制化的数据合成。为此,该文提出了CuSyn,一种专为多用户、多次发布的数据分析环境而设计的新颖、可定制的差分隐私数据合成框架。CuSyn引入了一种可扩展的基于预算池的隐私管理策略,能够在不同用户与多次发布之间施加严格的隐私预算约束。同时,该框架还提供了一种定制化合成机制,能够为不同用户生成针对其查询需求优化的高可用合成数据。该文在真实数据集上对CuSyn进行了评估,实验结果表明,与现有数据合成机制相比,CuSyn能够在用户关注的查询上提供更高的数据效用,在NLTCS 数据集上用户兴趣查询误差较基线降低约 67%,同时在复杂的多用户发布场景下提供可靠的隐私保护。
  • 图  1  CuSyn合成框架工作流程

    图  2  不同长度用户兴趣查询序列下的工作负载误差

    图  3  不同隐私预算水平$ \epsilon =\{0.1,1,5,10\},\delta ={10}^{-9} $下的工作负载误差

    图  4  AIM算法下模拟多用户的兴趣序列误差改进比率

    图  5  AIM算法下模拟多用户的整体工作负载误差改进比率

    1  $ \text{BudgetEval}(u,D,{\epsilon }_{\text{need}}) $函数运行逻辑

     输入:用户标识: $ u $; 数据集标识: $ D $; 请求预算: $ {\epsilon }_{\text{need}} $
     输出:最终分配预算:$ {\epsilon }_{\text{alloc}} $
     (1) if请求超出单次上限$ {\epsilon }_{\text{need}} \gt \alpha $ 或预算池可用预算$ B_{D}^{\text{rem}}==0 $
     then
     (2)  返回0,分配失败
     (3) end if
     (4) 初始化最终分配预算为请求值$ {\epsilon }_{\text{alloc}}\leftarrow {\epsilon }_{\text{need}} $
     (5) 应用所有预算约束$ {\epsilon }_{\text{alloc}}\leftarrow \min \left({\epsilon }_{\text{alloc}}, B_{u}^{\text{single}}, B_{D}^{\text{rem}}, B_{u}^{\text{rem}}\right) $
     (6) if 所有约束下可分配预算$ {\epsilon }_{\text{alloc}}=0 $ then
     (7)  返回0,分配失败
     (8) else
     (9) 批准分配,并更新相关预算池$ B_{D}^{\text{rem}}\leftarrow B_{D}^{\text{rem}}-{\epsilon }_{\text{alloc}} $,
     $ B_{u}^{\text{rem}}\leftarrow B_{u}^{\text{rem}}-{\epsilon }_{\text{alloc}} $
     (10) 返回实际批准的预算$ {\epsilon }_{\text{alloc}} $
     (11) end if
    下载: 导出CSV

    2  CuSyn数据合成框架

     输入:原始数据集:$ D $;当前轮次隐私参数:$ ({\epsilon }_{\text{need}},\delta ) $;工作负
     载: $ W $;用户兴趣查询序列:$ {\mathcal{C}}^{u} $;迭代轮数:$ T $
     输出:定制合成数据集:$ \hat{D} $
     (1) 为当前轮次合成请求预算:$ {\epsilon }_{\text{alloc}} $$ \leftarrow $$ \text{BudgetEval}(u,D,{\epsilon }_{\text{need}}) $
     (2) if $ {\epsilon }_{\text{alloc}} $值为0(预算分配失败)then
     (3) 停止算法
     (4) end
     (5) 执行定制化数据合成获得合成数据$ \hat{D} $$ \leftarrow $
     $ \mathrm{CustomizableSynthesis}(D,{\epsilon }_{\text{alloc}},\delta ,W,{\mathcal{C}}^{u},T) $
    下载: 导出CSV

    3  以MWEM-PGM为例的$ \mathrm{CustomizableSynthesis} $函数演示

     输入:原始数据集:$ D $;已分配的隐私预算:$ ({\epsilon }_{\text{alloc}},\delta ) $;工作负
     载:$ W $;用户兴趣查询序列: $ {\mathcal{C}}^{u} $;迭代轮数:$ T $
     输出:定制合成数据集:$ \hat{D} $
     (1) 初始化概率分布$ \widehat{{p}_{0}} $ $ \leftarrow $ Uniform$ [\mathcal{X}] $
     (2) 转换隐私参数$ ({\epsilon }_{\text{alloc}},\delta ) $到zCDP的隐私参数$ \rho $: $ \rho $$ \leftarrow $
     $ \mathrm{EpsilonToRho}({\epsilon }_{\text{alloc}},\delta ) $
     (3) 获取用户兴趣查询序列$ {\mathcal{C}}^{u} $的长度$ L $: $ L=\mathrm{Length}({\mathcal{C}}^{u}) $
     (4) 为基于PGM的数据合成计算初始的$ \sigma $: $ \sigma \leftarrow \sqrt{(T+L)/(2\rho )} $
     (5) for $ t=1,2,\cdots,T $ do
     (6)  使用指数机制和隐私预算$ \rho $选择 $ {C}_{t}\in \mathbf{W} $:
      $ {q}_{C}(D)\leftarrow \| {\mathbf{Q}}_{C}(D)-{\mathbf{Q}}_{C}({\hat{p}}_{t-1}){\| }_{1}-{n}_{C} $
     (7)  对选定的$ {C}_{t} $进行测量:$ {\mathbf{y}}_{t}\leftarrow {\mathbf{Q}}_{{{C}_{t}}}(D)+\mathcal{N}\left(0,{\sigma }^{2}\mathbb{I}\right) $
     (8)  利用PGM算法估计一个合成数据分布$ {\hat{p}}_{t} $:
     $ {\hat{p}}_{t}\leftarrow {\mathrm{argmin}}_{p\in S}\sum \limits_{i=1}^{t}\left|\left|{\mathbf{Q}}_{{{C}_{i}}}(p)-{\mathbf{y}}_{i}\right|\right|_{2}^{2} $
     (9) end
     (10) for $ l=1,2,\cdots,L $ do
     (11) 每轮按顺序从$ {\mathcal{C}}^{u} $中(即第一轮选取序列第一项)选取一项
     $ {\mathcal{C}}_{l} $,使用与第7行同样的方法进行测量
     (12) 使用与第8行同样的方法估计一个合成数据分布$ {\hat{p}}_{l} $
     (13) end
     (14) 利用最后一轮获得的$ {\hat{p}}_{l} $,使用Private-PGM算法产生最终
     的定制合成数据集$ \hat{D} $
    下载: 导出CSV

    表  1  实验数据集参数

    数据集名属性数量属性域大小数据集条目数量
    NLTCS16$ 7\times {10}^{4} $21574
    COLORADO14$ 4\times {10}^{5} $661967
    TITANIC9$ 9\times {10}^{7} $1304
    SALARY9$ 1\times {10}^{13} $135727
    下载: 导出CSV

    表  2  实验设置

    实验名隐私预算兴趣序列长度合成轮次
    LEN1.05, 10, 151
    BGT0.1, 1.0, 5.0, 10.0101
    MUL10.0(单个模拟用户多轮总和预算)55, 10, 15
    下载: 导出CSV
  • [1] CHEN Chaochao, WU Huiwen, SU Jiajie, et al. Differential private knowledge transfer for privacy-preserving cross-domain recommendation[C]. Proceedings of the ACM Web Conference 2022, Lyon, France, 2022: 1455–1465. doi: 10.1145/3485447.3512192.
    [2] MCKENNA R, MULLINS B, SHELDON D, et al. AIM: An adaptive and iterative mechanism for differentially private synthetic data[J]. Proceedings of the VLDB Endowment, 2022, 15(11): 2599–2612. doi: 10.14778/3551793.3551817.
    [3] CAI Kuntai, LEI Xiaoyu, WEI Jianxin, et al. Data synthesis via differentially private Markov random fields[J]. Proceedings of the VLDB Endowment, 2021, 14(11): 2190–2202. doi: 10.14778/3476249.3476272.
    [4] DWORK C and ROTH A. The algorithmic foundations of differential privacy[J]. Foundations and Trends® in Theoretical Computer Science, 2014, 9(3/4): 211–407. doi: 10.1561/0400000042.
    [5] BUN M and STEINKE T. Concentrated differential privacy: Simplifications, extensions, and lower bounds[C]. Proceedings the 14th International Conference on Theory of Cryptography, Beijing, China, 2016: 635–658. doi: 10.1007/978-3-662-53641-4_24.
    [6] MCKENNA R, MIKLAU G, and SHELDON D. Winning the NIST Contest: A scalable and general approach to differentially private synthetic data[J]. Journal of Privacy and Confidentiality, 2021, 11(3): 1–30. doi: 10.29012/jpc.778.
    [7] MCKENNA R and LIU T. A simple recipe for private synthetic data generation[EB/OL]. https://differentialprivacy.org/synth-data-1/, 2022.
    [8] MCKENNA R, SHELDON D, and MIKLAU G. Graphical-model based estimation and inference for differential privacy[C]. International Conference on Machine Learning, Long Beach, USA, 2019: 4435–4444.
    [9] CORMODE G, MADDOCK S, ULLAH E, et al. Synthetic tabular data: Methods, attacks and defenses[C]. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Toronto, Canada, 2025: 5989–5998. doi: 10.1145/3711896.3736562.
    [10] FENG Shuya, MOHAMMADY M, WANG Han, et al. DPI: Ensuring strict differential privacy for infinite data streaming[C]. 2024 IEEE Symposium on Security and Privacy (SP), San Francisco, USA, 2024: 1009–1027. doi: 10.1109/SP54263.2024.00124.
    [11] WANG Xiujun, MO Lei, ZHENG Xiao, et al. Streaming histogram publication over weighted sliding windows under differential privacy[J]. Tsinghua Science and Technology, 2024, 29(6): 1674–1693. doi: 10.26599/TST.2023.9010083.
    [12] LI Xiaochen, LI Tianyu, CHENG Yitian, et al. Spas: Continuous release of data streams under w-event differential privacy[J]. Proceedings of the ACM on Management of Data, 2025, 3(1): 78a. doi: 10.1145/3714420.
    [13] LIU Xiang, GUO Yuchun, CHEN Yishuai, et al. Trajectory privacy protection on spatial streaming data with differential privacy[C]. 2018 IEEE Global Communications Conference (GLOBECOM), Abu Dhabi, UAE, 2018: 1–7. doi: 10.1109/GLOCOM.2018.8647918.
    [14] HU Yujia, DU Yuntao, ZHANG Zhikun, et al. Real-time trajectory synthesis with local differential privacy[C]. 2024 IEEE 40th International Conference on Data Engineering (ICDE), Utrecht, Netherlands, 2024: 1685–1698. doi: 10.1109/ICDE60146.2024.00137.
    [15] LIU Weiming, ZHENG Xiaolin, CHEN Chaochao, et al. Reducing item discrepancy via differentially private robust embedding alignment for privacy-preserving cross domain recommendation[C]. 41st International Conference on Machine Learning, Vienna, Austria, 2024: 1317. doi: 10.5555/3692070.3693387.
    [16] ALSHANTTI A, VARAGNOLO D, RASHEED A, et al. CasTGAN: Cascaded generative adversarial network for realistic tabular data synthesis[J]. IEEE Access, 2024, 12: 13213–13232. doi: 10.1109/ACCESS.2024.3356913.
    [17] KATO F, TAKAHASHI T, TAKAGI S, et al. HDPView: Differentially private materialized view for exploring high dimensional relational data[J]. Proceedings of the VLDB Endowment, 2022, 15(9): 1766–1778. doi: 10.14778/3538598.3538601.
    [18] 魏立斐, 张无忌, 张蕾, 等. 基于本地差分隐私的异步横向联邦安全梯度聚合方案[J]. 电子与信息学报, 2024, 46(7): 3010–3018. doi: 10.11999/JEIT230923.

    WEI Lifei, ZHANG Wuji, ZHANG Lei, et al. A secure gradient aggregation scheme based on local differential privacy in asynchronous horizontal federated learning[J]. Journal of Electronics & Information Technology, 2024, 46(7): 3010–3018. doi: 10.11999/JEIT230923.
    [19] 张朋飞, 程俊, 张治坤, 等. 满足本地差分隐私的混合噪音感知的模糊C均值聚类算法[J]. 电子与信息学报, 2025, 47(3): 739–757. doi: 10.11999/JEIT241067.

    ZHANG Pengfei, CHENG Jun, ZHANG Zhikun, et al. Fuzzy C-means clustering algorithm based on mixed noise-aware under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(3): 739–757. doi: 10.11999/JEIT241067.
    [20] 张朋飞, 安建隆, 程祥, 等. 本地差分隐私下基于混合分布的真值发现算法[J]. 电子与信息学报, 2025, 47(6): 1896–1910. doi: 10.11999/JEIT240936.

    ZHANG Pengfei, AN Jianlong, CHENG Xiang, et al. Mixture distribution-based truth discovery algorithm under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1896–1910. doi: 10.11999/JEIT240936.
    [21] 方贤进, 甄雅茹, 张朋飞, 等. 本地差分隐私下技能感知的任务分配算法研究[J]. 电子与信息学报, 2025, 47(11): 4429–4439. doi: 10.11999/JEIT250425.

    FANG Xianjin, ZHEN Yaru, ZHANG Pengfei, et al. Research on skill-aware task assignment algorithm under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4429–4439. doi: 10.11999/JEIT250425.
  • 加载中
图(5) / 表(5)
计量
  • 文章访问数:  40
  • HTML全文浏览量:  4
  • PDF下载量:  1
  • 被引次数: 0
出版历程
  • 修回日期:  2026-09-15
  • 录用日期:  2026-09-15
  • 网络出版日期:  2026-10-10

目录

    /

    返回文章
    返回