Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning

Hamza Harkous; Kassem Fawaz; Rémi Lebret; Florian Schaub; Kang G. Shin; Karl Aberer

Authors:

Hamza Harkous, École Polytechnique Fédérale de Lausanne (EPFL); Kassem Fawaz, University of Wisconsin-Madison; Rémi Lebret, École Polytechnique Fédérale de Lausanne (EPFL); Florian Schaub and Kang G. Shin, University of Michigan; Karl Aberer, École Polytechnique Fédérale de Lausanne (EPFL)

Abstract:

Privacy policies are the primary channel through which companies inform users about their data collection and sharing practices. These policies are often long and difficult to comprehend. Short notices based on information extracted from privacy policies have been shown to be useful but face a significant scalability hurdle, given the number of policies and their evolution over time. Companies, users, researchers, and regulators still lack usable and scalable tools to cope with the breadth and depth of privacy policies. To address these hurdles, we propose an automated framework for privacy policy analysis (Polisis). It enables scalable, dynamic, and multi-dimensional queries on natural language privacy policies. At the core of Polisis is a privacy-centric language model, built with 130K privacy policies, and a novel hierarchy of neural-network classifiers that accounts for both high-level aspects and fine-grained details of privacy practices. We demonstrate Polisis’ modularity and utility with two applications supporting structured and free-form querying. The structured querying application is the automated assignment of privacy icons from privacy policies. With Polisis, we can achieve an accuracy of 88.4% on this task. The second application, PriBot, is the first freeform question-answering system for privacy policies. We show that PriBot can produce a correct answer among its top-3 results for 82% of the test questions. Using an MTurk user study with 700 participants, we show that at least one of PriBot’s top-3 answers is relevant to users for 89% of the test questions.

Hamza Harkous, École Polytechnique Fédérale de Lausanne (EPFL)

Kassem Fawaz, University of Wisconsin-Madison

Rémi Lebret, École Polytechnique Fédérale de Lausanne (EPFL)

Florian Schaub, University of Michigan

Kang G. Shin, University of Michigan

Karl Aberer, École Polytechnique Fédérale de Lausanne (EPFL);

Open Access Media

USENIX is committed to Open Access to the research presented at our events. Papers and proceedings are freely available to everyone once the event begins. Any video, audio, and/or slides that are posted after the event are also free and open to everyone. Support USENIX and our commitment to Open Access.

BibTeX

@inproceedings {217470,
author = {Hamza Harkous and Kassem Fawaz and R{\'e}mi Lebret and Florian Schaub and Kang G. Shin and Karl Aberer},
title = {Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning},
booktitle = {27th USENIX Security Symposium (USENIX Security 18)},
year = {2018},
isbn = {978-1-939133-04-5},
address = {Baltimore, MD},
pages = {531--548},
url = {https://www.usenix.org/conference/usenixsecurity18/presentation/harkous},
publisher = {USENIX Association},
month = aug
}

Download