Through Their Eyes: User Perceptions on Sensitive Attribute Inference of Social Media Videos by Visual Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Shuning, Zhang, Gengrui, Meng, Yibo, Zhang, Ziyi, Zhao, Hantao, Yi, Xin, Li, Hewu
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911101150560256
author Zhang, Shuning
Zhang, Gengrui
Meng, Yibo
Zhang, Ziyi
Zhao, Hantao
Yi, Xin
Li, Hewu
author_facet Zhang, Shuning
Zhang, Gengrui
Meng, Yibo
Zhang, Ziyi
Zhao, Hantao
Yi, Xin
Li, Hewu
contents The rapid advancement of Visual Language Models (VLMs) has enabled sophisticated analysis of visual content, leading to concerns about the inference of sensitive user attributes and subsequent privacy risks. While technical capabilities of VLMs are increasingly studied, users' understanding, perceptions, and reactions to these inferences remain less explored, especially concerning videos uploaded on the social media. This paper addresses this gap through a semi-structured interview (N=17), investigating user perspectives on VLM-driven sensitive attribute inference from their visual data. Findings reveal that users perceive VLMs as capable of inferring a range of attributes, including location, demographics, and socioeconomic indicators, often with unsettling accuracy. Key concerns include unauthorized identification, misuse of personal information, pervasive surveillance, and harm from inaccurate inferences. Participants reported employing various mitigation strategies, though with skepticism about their ultimate effectiveness against advanced AI. Users also articulate clear expectations for platforms and regulators, emphasizing the need for enhanced transparency, user control, and proactive privacy safeguards. These insights are crucial for guiding the development of responsible AI systems, effective privacy-enhancing technologies, and informed policymaking that aligns with user expectations and societal values.
format Preprint
id arxiv_https___arxiv_org_abs_2508_07658
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Through Their Eyes: User Perceptions on Sensitive Attribute Inference of Social Media Videos by Visual Language Models
Zhang, Shuning
Zhang, Gengrui
Meng, Yibo
Zhang, Ziyi
Zhao, Hantao
Yi, Xin
Li, Hewu
Human-Computer Interaction
The rapid advancement of Visual Language Models (VLMs) has enabled sophisticated analysis of visual content, leading to concerns about the inference of sensitive user attributes and subsequent privacy risks. While technical capabilities of VLMs are increasingly studied, users' understanding, perceptions, and reactions to these inferences remain less explored, especially concerning videos uploaded on the social media. This paper addresses this gap through a semi-structured interview (N=17), investigating user perspectives on VLM-driven sensitive attribute inference from their visual data. Findings reveal that users perceive VLMs as capable of inferring a range of attributes, including location, demographics, and socioeconomic indicators, often with unsettling accuracy. Key concerns include unauthorized identification, misuse of personal information, pervasive surveillance, and harm from inaccurate inferences. Participants reported employing various mitigation strategies, though with skepticism about their ultimate effectiveness against advanced AI. Users also articulate clear expectations for platforms and regulators, emphasizing the need for enhanced transparency, user control, and proactive privacy safeguards. These insights are crucial for guiding the development of responsible AI systems, effective privacy-enhancing technologies, and informed policymaking that aligns with user expectations and societal values.
title Through Their Eyes: User Perceptions on Sensitive Attribute Inference of Social Media Videos by Visual Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2508.07658