A Comparative Study of Fidelity Metric Variability in Synthetic Tabular Data Generated With and Without Differential Privacy Guarantees

dc.contributor.authorUllah, Ghufran
dc.contributor.departmentfi=Tietotekniikan laitos|en=Department of Computing|
dc.contributor.facultyfi=Teknillinen tiedekunta|en=Faculty of Technology|
dc.contributor.studysubjectfi=Tietotekniikka|en=Information and Communication Technology|
dc.date.accessioned2026-08-03T19:31:37Z
dc.date.issued2026-07-14
dc.description.abstractSynthetic tabular data generation is increasingly used when real datasets cannot be shared directly because of privacy, legal, or ethical concerns. However, synthetic data is useful when it preserves important statistical properties of the real data and produces reliable results across repeated generations. This thesis investigates the fidelity metric variability of differentially private (DP) and Non-Differential Private (Non-DP) synthetic tabular data generators. The main objective is to examine how closely synthetic data reproduces the univariate distributions of real data and how much run-to-run dispersion these fidelity measurements show. The study compares six synthetic data generation methods using three datasets: Adult, Bank Marketing, and Cardiovascular Disease. The Non-DP synthesizers include Gaussian Copula, CTGAN, TVAE, and RTVAE, while the DP synthesizers include AIM and PrivBayes. The experiments are conducted across three sample sizes, n = 500, n = 5000, and n = 10000. For DP methods, three privacy budget values are used: ϵ = 1, ϵ = 5, and ϵ = 10. Hellinger distance is used to capture the geometric difference between probability distributions, while Jensen Shannon distance is used to capture the information theoretic difference between each distribution and their shared midpoint distribution. The results show that the Gaussian Copula model generally achieves high synthetic data fidelity based on Hellinger and Jensen Shannon distances for categorical variables, especially for the Adult and Bank datasets. CTGAN and TVAE show moderate performance, while RTVAE shows higher run-to-run variability for several complex variables. The Cardio dataset is more difficult for all synthesizers, mainly because of variables with irregular numerical and clinical distributions. Among the differentially private synthesizers, AIM generally outperforms PrivBayes by producing lower distance values and lower-dispersion fidelity scores. Increasing the sample size usually improves stability, especially from n = 500 to n = 5000. Increasing the privacy budget from ϵ = 1 to ϵ = 5 generally improves fidelity and stability, although the improvement from ϵ = 5 to ϵ = 10 did not produce a uniform improvement across models and variables.
dc.format.extent108
dc.identifier.urihttps://www.utupub.fi/handle/11111/62861
dc.identifier.urnURN:NBN:fi-fe20260803114548
dc.language.isoeng
dc.rightsfi=Julkaisu on tekijänoikeussäännösten alainen. Teosta voi lukea ja tulostaa henkilökohtaista käyttöä varten. Käyttö kaupallisiin tarkoituksiin on kielletty.|en=This publication is copyrighted. You may download, display and print it for Your own personal use. Commercial use is prohibited.|
dc.rights.accessrightsavoin
dc.subjectSynthetic data
dc.subjectdifferential privacy
dc.subjectdata fidelity
dc.subjectHellinger distance
dc.subjectJensen Shannon distance
dc.subjectprivacy budget
dc.titleA Comparative Study of Fidelity Metric Variability in Synthetic Tabular Data Generated With and Without Differential Privacy Guarantees
dc.type.ontasotfi=Diplomityö|en=Master's thesis|

Tiedostot

Näytetään 1 - 1 / 1
Ladataan...
Name:
Ullah_Ghufran_Thesis.pdf
Size:
12.43 MB
Format:
Adobe Portable Document Format