CuSyn: A Customizable Differential Privacy Data Synthesis Framework for Multi-User and Multiple Data Publishing Scenarios
-
摘要: 尽管现有许多隐私保护数据合成机制能够生成高可用性的合成数据集用于发布,但它们通常缺乏对多用户、多次发布场景下细粒度隐私保障的专门设计。这一局限性主要体现在两方面:一是缺乏系统性的隐私预算跟踪与管理方法,二是难以根据不同用户的具体分析需求进行定制化的数据合成。为此,该文提出了CuSyn,一种专为多用户、多次发布的数据分析环境而设计的新颖、可定制的差分隐私数据合成框架。CuSyn引入了一种可扩展的基于预算池的隐私管理策略,能够在不同用户与多次发布之间施加严格的隐私预算约束。同时,该框架还提供了一种定制化合成机制,能够为不同用户生成针对其查询需求优化的高可用合成数据。该文在真实数据集上对CuSyn进行了评估,实验结果表明,与现有数据合成机制相比,CuSyn能够在用户关注的查询上提供更高的数据效用,在NLTCS 数据集上用户兴趣查询误差较基线降低约 67%,同时在复杂的多用户发布场景下提供可靠的隐私保护。Abstract:
Objective With the rapid expansion of data-driven services and an increasing societal emphasis on privacy, the publication of privacy-preserving synthetic data has become a critical task for balancing analytical utility and individual confidentiality. While numerous differential privacy (DP) data synthesis mechanisms have been proposed to generate high-utility datasets, their design predominantly targets single-user, single-release scenarios. Deploying these conventional methods in environments characterized by multiple concurrent users and iterative data publishing introduces significant and unresolved challenges. The primary limitations are twofold: (1) the absence of systematic privacy budget tracking and fine-grained management mechanisms, which leaves the cumulative privacy risk across multiple releases uncontrolled and unauditable; and (2) the inability to customize synthesized data to meet the distinct and specific analytical queries of different users, resulting in suboptimal utility for targeted analyses. Consequently, there is a pressing need for a novel, integrated framework that can provide provable, hierarchical privacy budget control alongside on-demand data synthesis, thereby ensuring strong, auditable privacy protection and high data utility in complex, multi-user publishing ecosystems. Methods To systematically address these challenges, this paper proposes CuSyn, a customizable differential privacy data synthesis framework designed for multi-user and multiple data publishing environments. It is a systematic framework that integrates the customization requirement into the standard select-measure-generate paradigm through well-defined interfaces, and it comprises two synergistic components. First, a scalable, pool-based privacy budget management strategy provides strict, traceable control over privacy expenditure. For each original dataset, a dedicated budget pool is maintained and initialized with a predefined global limit, and a user-level pool may be maintained for each authorized user to enforce differentiated access control. A central BudgetEval function validates, negotiates, and allocates the budget of each synthesis request (Algorithm 1) through a configurable constraint system composed of a per-operation system limit, a per-request user limit, the remaining budget of the dataset pool, and the remaining budget of the user's pool. This hierarchical validation establishes control ranging from the global system level, through the individual user level, down to the dataset level, and ensures that cumulative privacy leakage across all users and releases remains strictly bounded even under worst-case collusion. When a request exceeds some constraint without being entirely unsatisfiable, the function does not reject it but adaptively negotiates the maximum admissible value as the approved budget. A formal composition analysis is provided in the zero-concentrated differential privacy (zCDP) framework, in which each approved budget is converted into a zCDP parameter and the privacy loss of all synthesis operations is bounded by the sum of these parameters, which cannot exceed the zCDP parameter induced by the preset global budget because the pool is never refilled. The guarantee therefore matches that of the global budget and depends only on the budget actually consumed rather than on user identity, so that it remains valid even when all users collude and pool their synthetic datasets. Second, a customizable synthesis mechanism extends the standard pipeline. Users submit a sequence of queries of interest; CuSyn first conducts the standard iterative query selection and measurement rounds to capture the general statistical structure, and then performs additional dedicated measurements on the user-specified queries (Algorithm 2 and Algorithm 3). The noisy results of both phases are jointly used to estimate the final probabilistic graphical model, from which the customized synthetic dataset is generated. The allocated budget is shared by the iterative rounds and the interest-query measurements, and the scale of the Gaussian noise is set accordingly, so that customization is obtained without additional privacy cost. This general-then-specific approach directly optimizes utility on user-specified queries without significantly degrading overall accuracy, and the number of standard iterations may be reduced for long interest sequences so as to avoid excessive dilution of the per-query budget. Results and Discussions The performance of CuSyn is evaluated on multiple real-world datasets with varying characteristics and scales ( Table 1 ) and compared against state-of-the-art baseline PGM-based mechanisms, namely AIM and MWEM-PGM. Three sets of experiments are conducted: one varying the length of the user interest query sequence (LEN experiments), one varying the privacy budget (BGT experiments), and one simulating multiple users publishing repeatedly under a shared budget pool (MUL experiments), with the settings summarized inTable 2 . The interest queries of each user are randomly drawn from the workload that contains all two-dimensional marginal queries of the dataset. In the LEN experiments (Fig. 2 ), instances of CuSyn consistently achieve significantly lower error on the user-specified interest queries than the baseline mechanisms across most dataset and sequence length configurations. For example, on the NLTCS dataset with a sequence length of 10, the error of CuSyn on the interest queries is reduced by approximately 67% with respect to the baseline, while its error on the entire evaluation workload remains comparable to that of the baselines, and in some configurations, such as the NLTCS dataset with a sequence length of 15, the overall workload error is also reduced substantially. A longer interest sequence, however, does not guarantee a lower error, because the budget is shared by the standard iterative phase and the interest-measurement phase; when the sequence occupies a considerable fraction of the workload, the standard phase is compressed and the overall workload error increases, so that the sequence length should balance the utility of the queries that a user cares about against the overall data availability. In the BGT experiments (Fig. 3 ), a similar trend is observed across different privacy budget levels. The average workload error of CuSyn over all attributes decreases as the budget increases, and its utility on the interest queries is higher than that of the baselines in most configurations. In the MUL experiments (Fig. 4 ,Fig. 5 ), four independent users holding different interest query sequences are simulated under a total budget limit, and CuSyn retains its utility advantage on the interest queries across 5, 10, and 15 rounds of multi-user publishing, while the overall workload error remains comparable to that of the single-round case on most datasets. The results confirm that CuSyn provides an effective and practical trade-off, substantially enhancing data utility for the queries specified by a particular user while maintaining competitive or superior utility for the general workload, which validates its core capability to customize synthesized data for targeted analysis. The current budget allocation strategy is static, and all synthesis operations are treated as new releases; future work could therefore explore budget recovery mechanisms for incremental updates under parallel composition rules.Conclusions This paper addresses the critical gap in privacy-preserving data publishing for multi-user and multiple-release scenarios by proposing the CuSyn framework. Its pool-based budget management system offers a practical and provably secure solution for fine-grained, hierarchical control over privacy expenditure, mitigating the risks of cumulative leakage. The customizable synthesis mechanism integrates user-defined interest queries into the data generation process, improving the utility of synthetic data for targeted analytical tasks. Extensive evaluations confirm that CuSyn outperforms state-of-the-art baselines on interest-query utility while preserving strong overall performance in both single-round and multi-round publishing. Future work will focus on adaptive budget allocation strategies based on data sensitivity and user query patterns, and on more efficient budget recovery for incremental publishing. -
Key words:
- Differential privacy /
- Data synthesis /
- Data publishing
-
1 $ \text{BudgetEval}(u,D,{\epsilon }_{\text{need}}) $函数运行逻辑
输入:用户标识: $ u $; 数据集标识: $ D $; 请求预算: $ {\epsilon }_{\text{need}} $ 输出:最终分配预算:$ {\epsilon }_{\text{alloc}} $ (1) if请求超出单次上限$ {\epsilon }_{\text{need}} \gt \alpha $ 或预算池可用预算$ B_{D}^{\text{rem}}==0 $
then(2) 返回0,分配失败 (3) end if (4) 初始化最终分配预算为请求值$ {\epsilon }_{\text{alloc}}\leftarrow {\epsilon }_{\text{need}} $ (5) 应用所有预算约束$ {\epsilon }_{\text{alloc}}\leftarrow \min \left({\epsilon }_{\text{alloc}}, B_{u}^{\text{single}}, B_{D}^{\text{rem}}, B_{u}^{\text{rem}}\right) $ (6) if 所有约束下可分配预算$ {\epsilon }_{\text{alloc}}=0 $ then (7) 返回0,分配失败 (8) else (9) 批准分配,并更新相关预算池$ B_{D}^{\text{rem}}\leftarrow B_{D}^{\text{rem}}-{\epsilon }_{\text{alloc}} $,
$ B_{u}^{\text{rem}}\leftarrow B_{u}^{\text{rem}}-{\epsilon }_{\text{alloc}} $(10) 返回实际批准的预算$ {\epsilon }_{\text{alloc}} $ (11) end if 2 CuSyn数据合成框架
输入:原始数据集:$ D $;当前轮次隐私参数:$ ({\epsilon }_{\text{need}},\delta ) $;工作负
载: $ W $;用户兴趣查询序列:$ {\mathcal{C}}^{u} $;迭代轮数:$ T $输出:定制合成数据集:$ \hat{D} $ (1) 为当前轮次合成请求预算:$ {\epsilon }_{\text{alloc}} $$ \leftarrow $$ \text{BudgetEval}(u,D,{\epsilon }_{\text{need}}) $ (2) if $ {\epsilon }_{\text{alloc}} $值为0(预算分配失败)then (3) 停止算法 (4) end (5) 执行定制化数据合成获得合成数据$ \hat{D} $$ \leftarrow $
$ \mathrm{CustomizableSynthesis}(D,{\epsilon }_{\text{alloc}},\delta ,W,{\mathcal{C}}^{u},T) $3 以MWEM-PGM为例的$ \mathrm{CustomizableSynthesis} $函数演示
输入:原始数据集:$ D $;已分配的隐私预算:$ ({\epsilon }_{\text{alloc}},\delta ) $;工作负
载:$ W $;用户兴趣查询序列: $ {\mathcal{C}}^{u} $;迭代轮数:$ T $输出:定制合成数据集:$ \hat{D} $ (1) 初始化概率分布$ \widehat{{p}_{0}} $ $ \leftarrow $ Uniform$ [\mathcal{X}] $ (2) 转换隐私参数$ ({\epsilon }_{\text{alloc}},\delta ) $到zCDP的隐私参数$ \rho $: $ \rho $$ \leftarrow $
$ \mathrm{EpsilonToRho}({\epsilon }_{\text{alloc}},\delta ) $(3) 获取用户兴趣查询序列$ {\mathcal{C}}^{u} $的长度$ L $: $ L=\mathrm{Length}({\mathcal{C}}^{u}) $ (4) 为基于PGM的数据合成计算初始的$ \sigma $: $ \sigma \leftarrow \sqrt{(T+L)/(2\rho )} $ (5) for $ t=1,2,\cdots,T $ do (6) 使用指数机制和隐私预算$ \rho $选择 $ {C}_{t}\in \mathbf{W} $:
$ {q}_{C}(D)\leftarrow \| {\mathbf{Q}}_{C}(D)-{\mathbf{Q}}_{C}({\hat{p}}_{t-1}){\| }_{1}-{n}_{C} $(7) 对选定的$ {C}_{t} $进行测量:$ {\mathbf{y}}_{t}\leftarrow {\mathbf{Q}}_{{{C}_{t}}}(D)+\mathcal{N}\left(0,{\sigma }^{2}\mathbb{I}\right) $ (8) 利用PGM算法估计一个合成数据分布$ {\hat{p}}_{t} $:
$ {\hat{p}}_{t}\leftarrow {\mathrm{argmin}}_{p\in S}\sum \limits_{i=1}^{t}\left|\left|{\mathbf{Q}}_{{{C}_{i}}}(p)-{\mathbf{y}}_{i}\right|\right|_{2}^{2} $(9) end (10) for $ l=1,2,\cdots,L $ do (11) 每轮按顺序从$ {\mathcal{C}}^{u} $中(即第一轮选取序列第一项)选取一项
$ {\mathcal{C}}_{l} $,使用与第7行同样的方法进行测量(12) 使用与第8行同样的方法估计一个合成数据分布$ {\hat{p}}_{l} $ (13) end (14) 利用最后一轮获得的$ {\hat{p}}_{l} $,使用Private-PGM算法产生最终
的定制合成数据集$ \hat{D} $表 1 实验数据集参数
数据集名 属性数量 属性域大小 数据集条目数量 NLTCS 16 $ 7\times {10}^{4} $ 21574 COLORADO 14 $ 4\times {10}^{5} $ 661967 TITANIC 9 $ 9\times {10}^{7} $ 1304 SALARY 9 $ 1\times {10}^{13} $ 135727 表 2 实验设置
实验名 隐私预算 兴趣序列长度 合成轮次 LEN 1.0 5, 10, 15 1 BGT 0.1, 1.0, 5.0, 10.0 10 1 MUL 10.0(单个模拟用户多轮总和预算) 5 5, 10, 15 -
[1] CHEN Chaochao, WU Huiwen, SU Jiajie, et al. Differential private knowledge transfer for privacy-preserving cross-domain recommendation[C]. Proceedings of the ACM Web Conference 2022, Lyon, France, 2022: 1455–1465. doi: 10.1145/3485447.3512192. [2] MCKENNA R, MULLINS B, SHELDON D, et al. AIM: An adaptive and iterative mechanism for differentially private synthetic data[J]. Proceedings of the VLDB Endowment, 2022, 15(11): 2599–2612. doi: 10.14778/3551793.3551817. [3] CAI Kuntai, LEI Xiaoyu, WEI Jianxin, et al. Data synthesis via differentially private Markov random fields[J]. Proceedings of the VLDB Endowment, 2021, 14(11): 2190–2202. doi: 10.14778/3476249.3476272. [4] DWORK C and ROTH A. The algorithmic foundations of differential privacy[J]. Foundations and Trends® in Theoretical Computer Science, 2014, 9(3/4): 211–407. doi: 10.1561/0400000042. [5] BUN M and STEINKE T. Concentrated differential privacy: Simplifications, extensions, and lower bounds[C]. Proceedings the 14th International Conference on Theory of Cryptography, Beijing, China, 2016: 635–658. doi: 10.1007/978-3-662-53641-4_24. [6] MCKENNA R, MIKLAU G, and SHELDON D. Winning the NIST Contest: A scalable and general approach to differentially private synthetic data[J]. Journal of Privacy and Confidentiality, 2021, 11(3): 1–30. doi: 10.29012/jpc.778. [7] MCKENNA R and LIU T. A simple recipe for private synthetic data generation[EB/OL]. https://differentialprivacy.org/synth-data-1/, 2022. [8] MCKENNA R, SHELDON D, and MIKLAU G. Graphical-model based estimation and inference for differential privacy[C]. International Conference on Machine Learning, Long Beach, USA, 2019: 4435–4444. [9] CORMODE G, MADDOCK S, ULLAH E, et al. Synthetic tabular data: Methods, attacks and defenses[C]. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Toronto, Canada, 2025: 5989–5998. doi: 10.1145/3711896.3736562. [10] FENG Shuya, MOHAMMADY M, WANG Han, et al. DPI: Ensuring strict differential privacy for infinite data streaming[C]. 2024 IEEE Symposium on Security and Privacy (SP), San Francisco, USA, 2024: 1009–1027. doi: 10.1109/SP54263.2024.00124. [11] WANG Xiujun, MO Lei, ZHENG Xiao, et al. Streaming histogram publication over weighted sliding windows under differential privacy[J]. Tsinghua Science and Technology, 2024, 29(6): 1674–1693. doi: 10.26599/TST.2023.9010083. [12] LI Xiaochen, LI Tianyu, CHENG Yitian, et al. Spas: Continuous release of data streams under w-event differential privacy[J]. Proceedings of the ACM on Management of Data, 2025, 3(1): 78a. doi: 10.1145/3714420. [13] LIU Xiang, GUO Yuchun, CHEN Yishuai, et al. Trajectory privacy protection on spatial streaming data with differential privacy[C]. 2018 IEEE Global Communications Conference (GLOBECOM), Abu Dhabi, UAE, 2018: 1–7. doi: 10.1109/GLOCOM.2018.8647918. [14] HU Yujia, DU Yuntao, ZHANG Zhikun, et al. Real-time trajectory synthesis with local differential privacy[C]. 2024 IEEE 40th International Conference on Data Engineering (ICDE), Utrecht, Netherlands, 2024: 1685–1698. doi: 10.1109/ICDE60146.2024.00137. [15] LIU Weiming, ZHENG Xiaolin, CHEN Chaochao, et al. Reducing item discrepancy via differentially private robust embedding alignment for privacy-preserving cross domain recommendation[C]. 41st International Conference on Machine Learning, Vienna, Austria, 2024: 1317. doi: 10.5555/3692070.3693387. [16] ALSHANTTI A, VARAGNOLO D, RASHEED A, et al. CasTGAN: Cascaded generative adversarial network for realistic tabular data synthesis[J]. IEEE Access, 2024, 12: 13213–13232. doi: 10.1109/ACCESS.2024.3356913. [17] KATO F, TAKAHASHI T, TAKAGI S, et al. HDPView: Differentially private materialized view for exploring high dimensional relational data[J]. Proceedings of the VLDB Endowment, 2022, 15(9): 1766–1778. doi: 10.14778/3538598.3538601. [18] 魏立斐, 张无忌, 张蕾, 等. 基于本地差分隐私的异步横向联邦安全梯度聚合方案[J]. 电子与信息学报, 2024, 46(7): 3010–3018. doi: 10.11999/JEIT230923.WEI Lifei, ZHANG Wuji, ZHANG Lei, et al. A secure gradient aggregation scheme based on local differential privacy in asynchronous horizontal federated learning[J]. Journal of Electronics & Information Technology, 2024, 46(7): 3010–3018. doi: 10.11999/JEIT230923. [19] 张朋飞, 程俊, 张治坤, 等. 满足本地差分隐私的混合噪音感知的模糊C均值聚类算法[J]. 电子与信息学报, 2025, 47(3): 739–757. doi: 10.11999/JEIT241067.ZHANG Pengfei, CHENG Jun, ZHANG Zhikun, et al. Fuzzy C-means clustering algorithm based on mixed noise-aware under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(3): 739–757. doi: 10.11999/JEIT241067. [20] 张朋飞, 安建隆, 程祥, 等. 本地差分隐私下基于混合分布的真值发现算法[J]. 电子与信息学报, 2025, 47(6): 1896–1910. doi: 10.11999/JEIT240936.ZHANG Pengfei, AN Jianlong, CHENG Xiang, et al. Mixture distribution-based truth discovery algorithm under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(6): 1896–1910. doi: 10.11999/JEIT240936. [21] 方贤进, 甄雅茹, 张朋飞, 等. 本地差分隐私下技能感知的任务分配算法研究[J]. 电子与信息学报, 2025, 47(11): 4429–4439. doi: 10.11999/JEIT250425.FANG Xianjin, ZHEN Yaru, ZHANG Pengfei, et al. Research on skill-aware task assignment algorithm under local differential privacy[J]. Journal of Electronics & Information Technology, 2025, 47(11): 4429–4439. doi: 10.11999/JEIT250425. -
下载: