Abstract:In crowdsourcing database, human operators are embedded into the database engine and collaborate with other conventional database operators to process the queries. Each human operator publishes small HITs (Human Intelligent Task) to the crowdsourcing platform, which consists of a set of database records and corresponding questions for human workers. The human workers complete the HITs and return the results to the crowdsourcing database for further processing. In practice, published records in HITs may contain sensitive attributes, probably causing privacy leakage so that malicious workers could link them with other public databases to reveal individual private information. Conventional privacy protection techniques, such as K-Anonymity , can be applied to partially solve the problem. However, after generalizing the data, the result of standard K-Anonymity algorithms may render uncontrollable information loss and affects the accuracy of crowdsourcing. In this paper, we first study the tradeoff between the privacy and accuracy for the human operator within data anonymization process. A probability model is proposed to estimate the lower bound and upper bound of the accuracy for general K-Anonymity approaches. We show that searching the optimal anonymity approach is NP-Hard and only heuristic approach is available. The second contribution of the paper is a general feedback-based K-Anonymity scheme. In our scheme, synthetic samples are published to the human workers, the results of which are used to guide the selection on anonymity strategies. We apply the scheme on Mondrian algorithm by adaptively cutting the dimensions based on our feedback results on the synthetic samples. We evaluate the performance of the feedback-based approach on U.S. census dataset, and show that given a predefined $K$ , our proposal outperforms standard K-Anonymity approaches on retaining the effectiveness of crowdsourcing.

(α, k)-anonymity based privacy preservation by lossy join

(&Alpha-Anonymity Based Privacy Preservation By Lossy Join

An Enhanced K-Anonymity Model Against Homogeneity Attack.

Privacy Inference Attacking And Prevention On Multiple Relative K-Anonymized Microdata Sets

A Data Privacy Preservation Method Based on Lossy Decomposition

Dissemination of Anonymized Streaming Data.

Clustering-Based k-anonymity

Supporting Pattern-Preserving Anonymization for Time-Series Data

A Dynamic Anonymization Privacy-Preserving Model Based on Hierarchical Sequential Three-Way Decisions

Privacy-enhancing K -Anonymization of Customer Data

K-Anonymity for Crowdsourcing Database

Maintaining K-Anonymity Against Incremental Updates

T-Closeness Slicing: A New Privacy-Preserving Approach for Transactional Data Publishing.

Multi-level personalized k-anonymity privacy-preserving model based on sequential three-way decisions

Towards an Anti-inference (K, ℓ)-Anonymity Model with Value Association Rules

A Personalized (a,k)-Anonymity Model

Differentially private data release for data mining

On Anonymization of Multi-graphs.

Towards the Diversity of Sensitive Attributes in k-Anonymity

(K,p)-Anonymity

Privacy Gain Based Multi-Iterative k-Anonymization to Protect Respondents Privacy