Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Weidi, Lu, Tianyu, Zhang, Qiming, Liu, Xiaogeng, Hu, Bin, Zhao, Yue, Zhao, Jieyu, Gao, Song, McDaniel, Patrick, Xiang, Zhen, Xiao, Chaowei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917309990305792
author Luo, Weidi
Lu, Tianyu
Zhang, Qiming
Liu, Xiaogeng
Hu, Bin
Zhao, Yue
Zhao, Jieyu
Gao, Song
McDaniel, Patrick
Xiang, Zhen
Xiao, Chaowei
author_facet Luo, Weidi
Lu, Tianyu
Zhang, Qiming
Liu, Xiaogeng
Hu, Bin
Zhao, Yue
Zhao, Jieyu
Gao, Song
McDaniel, Patrick
Xiang, Zhen
Xiao, Chaowei
contents Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models enable impressive reasoning capabilities, they also introduce novel and underexplored privacy risks. In this paper, we identify a novel category of privacy leakage in MLRMs: Adversaries can infer sensitive geolocation information, such as a user's home address or neighborhood, from user-generated images, including selfies captured in private settings. To formalize and evaluate these risks, we propose a three-level visual privacy risk framework that categorizes image content based on contextual sensitivity and potential for location inference. We further introduce DoxBench, a curated dataset of 500 real-world images reflecting diverse privacy scenarios. Our evaluation across 11 advanced MLRMs and MLLMs demonstrates that these models consistently outperform non-expert humans in geolocation inference and can effectively leak location-related private information. This significantly lowers the barrier for adversaries to obtain users' sensitive geolocation information. We further analyze and identify two primary factors contributing to this vulnerability: (1) MLRMs exhibit strong reasoning capabilities by leveraging visual clues in combination with their internal world knowledge; and (2) MLRMs frequently rely on privacy-related visual clues for inference without any built-in mechanisms to suppress or avoid such usage. To better understand and demonstrate real-world attack feasibility, we propose GeoMiner, a collaborative attack framework that decomposes the prediction process into two stages: clue extraction and reasoning to improve geolocation performance while introducing a novel attack perspective. Our findings highlight the urgent need to reassess inference-time privacy risks in MLRMs to better protect users' sensitive information.
format Preprint
id arxiv_https___arxiv_org_abs_2504_19373
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
Luo, Weidi
Lu, Tianyu
Zhang, Qiming
Liu, Xiaogeng
Hu, Bin
Zhao, Yue
Zhao, Jieyu
Gao, Song
McDaniel, Patrick
Xiang, Zhen
Xiao, Chaowei
Cryptography and Security
Artificial Intelligence
Recent advances in multi-modal large reasoning models (MLRMs) have shown significant ability to interpret complex visual content. While these models enable impressive reasoning capabilities, they also introduce novel and underexplored privacy risks. In this paper, we identify a novel category of privacy leakage in MLRMs: Adversaries can infer sensitive geolocation information, such as a user's home address or neighborhood, from user-generated images, including selfies captured in private settings. To formalize and evaluate these risks, we propose a three-level visual privacy risk framework that categorizes image content based on contextual sensitivity and potential for location inference. We further introduce DoxBench, a curated dataset of 500 real-world images reflecting diverse privacy scenarios. Our evaluation across 11 advanced MLRMs and MLLMs demonstrates that these models consistently outperform non-expert humans in geolocation inference and can effectively leak location-related private information. This significantly lowers the barrier for adversaries to obtain users' sensitive geolocation information. We further analyze and identify two primary factors contributing to this vulnerability: (1) MLRMs exhibit strong reasoning capabilities by leveraging visual clues in combination with their internal world knowledge; and (2) MLRMs frequently rely on privacy-related visual clues for inference without any built-in mechanisms to suppress or avoid such usage. To better understand and demonstrate real-world attack feasibility, we propose GeoMiner, a collaborative attack framework that decomposes the prediction process into two stages: clue extraction and reasoning to improve geolocation performance while introducing a novel attack perspective. Our findings highlight the urgent need to reassess inference-time privacy risks in MLRMs to better protect users' sensitive information.
title Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2504.19373