A Comparative Study of Fidelity Metric Variability in Synthetic Tabular Data Generated With and Without Differential Privacy Guarantees

avoin
Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.
Lataukset2

Verkkojulkaisu

DOI

Tiivistelmä

Synthetic tabular data generation is increasingly used when real datasets cannot be shared directly because of privacy, legal, or ethical concerns. However, synthetic data is useful when it preserves important statistical properties of the real data and produces reliable results across repeated generations. This thesis investigates the fidelity metric variability of differentially private (DP) and Non-Differential Private (Non-DP) synthetic tabular data generators. The main objective is to examine how closely synthetic data reproduces the univariate distributions of real data and how much run-to-run dispersion these fidelity measurements show. The study compares six synthetic data generation methods using three datasets: Adult, Bank Marketing, and Cardiovascular Disease. The Non-DP synthesizers include Gaussian Copula, CTGAN, TVAE, and RTVAE, while the DP synthesizers include AIM and PrivBayes. The experiments are conducted across three sample sizes, n = 500, n = 5000, and n = 10000. For DP methods, three privacy budget values are used: ϵ = 1, ϵ = 5, and ϵ = 10. Hellinger distance is used to capture the geometric difference between probability distributions, while Jensen Shannon distance is used to capture the information theoretic difference between each distribution and their shared midpoint distribution. The results show that the Gaussian Copula model generally achieves high synthetic data fidelity based on Hellinger and Jensen Shannon distances for categorical variables, especially for the Adult and Bank datasets. CTGAN and TVAE show moderate performance, while RTVAE shows higher run-to-run variability for several complex variables. The Cardio dataset is more difficult for all synthesizers, mainly because of variables with irregular numerical and clinical distributions. Among the differentially private synthesizers, AIM generally outperforms PrivBayes by producing lower distance values and lower-dispersion fidelity scores. Increasing the sample size usually improves stability, especially from n = 500 to n = 5000. Increasing the privacy budget from ϵ = 1 to ϵ = 5 generally improves fidelity and stability, although the improvement from ϵ = 5 to ϵ = 10 did not produce a uniform improvement across models and variables.

item.page.okmtext