Randomized trial presents an algorithm generating synthetic data, indicating improved privacy and accuracy.
We present a polynomial-time algorithm for online differentially private synthetic data generation. For a data stream within the hypercube <tex-math notation="LaTeX">[0,1]ᵈ</tex-math> and an infinite time horizon, we develop an online algorithm that generates a differentially private synthetic dataset at each time <tex-math notation="LaTeX">t</tex-math>. This algorithm achieves a near-optimal accuracy bound of <tex-math notation="LaTeX">O(log (t)t-1/d)</tex-math> for <tex-math notation="LaTeX">d≥ 2</tex-math> and <tex-math notation="LaTeX">O(log 4.5(t)t⁻¹)</tex-math> for <tex-math notation="LaTeX">$d=1$</tex-math> in the 1-Wasserstein distance. This result extends the previous work on the continual release model for counting queries to Lipschitz queries. Compared to the offline case, where the entire dataset is available at once (Boedihardjo et al., 2024), (He et al., 2023), our approach requires only an extra polylog factor in the accuracy bound.
No takes yet. Share an insight, caveat, or question.
He et al. (2024) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: