We found a match
Your institution may have access to this item. Find your institution then sign in to continue.
- Title
Differentially Private Distance Learning in Categorical Data.
- Authors
Battaglia, Elena; Celano, Simone; Pensa, Ruggero G.
- Abstract
Most privacy-preserving machine learning methods are designed around continuous or numeric data, but categorical attributes are common in many application scenarios, including clinical and health records, census and survey data. Distance-based methods, in particular, have limited applicability to categorical data, since they do not capture the complexity of the relationships among different values of a categorical attribute. Although distance learning algorithms exist for categorical data, they may disclose private information about individual records if applied to a secret dataset. To address this problem, we introduce a differentially private family of algorithms for learning distances between any pair of values of a categorical attribute according to the way they are co-distributed with the values of other categorical attributes forming the so-called context. We define different variants of our algorithm and we show empirically that our approach consumes little privacy budget while providing accurate distances, making it suitable in distance-based applications, such as clustering and classification.
- Subjects
CATEGORIES (Mathematics); DISTANCE education; MACHINE learning; ALGORITHMS; CLASSIFICATION algorithms
- Publication
Data Mining & Knowledge Discovery, 2021, Vol 35, Issue 5, p2050
- ISSN
1384-5810
- Publication type
Article
- DOI
10.1007/s10618-021-00778-0