Beyond Visual Appearances: Privacy-sensitive Objects Identification via Hybrid Graph Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Zhuohang, Tong, Bingkui, Du, Xia, Alhammadi, Ahmed, Zhou, Jizhe |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SHAN: Object-Level Privacy Detection via Inference on Scene Heterogeneous Graph
by: Jiang, Zhuohang, et al.
Published: (2024)
by: Jiang, Zhuohang, et al.
Published: (2024)
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023)
by: Ma, Xiaochen, et al.
Published: (2023)
IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization
by: Ma, Xiaochen, et al.
Published: (2024)
by: Ma, Xiaochen, et al.
Published: (2024)
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
M^3:Manipulation Mask Manufacturer for Arbitrary-Scale Super-Resolution Mask
by: Yang, Xinyu, et al.
Published: (2024)
by: Yang, Xinyu, et al.
Published: (2024)
Measuring Epistemic Humility in Multimodal Large Language Models
by: Tong, Bingkui, et al.
Published: (2025)
by: Tong, Bingkui, et al.
Published: (2025)
ARPGNet: Appearance- and Relation-aware Parallel Graph Attention Fusion Network for Facial Expression Recognition
by: Li, Yan, et al.
Published: (2025)
by: Li, Yan, et al.
Published: (2025)
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness
by: He, Yusheng, et al.
Published: (2026)
by: He, Yusheng, et al.
Published: (2026)
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
by: Xu, Yu, et al.
Published: (2026)
by: Xu, Yu, et al.
Published: (2026)
Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization
by: Zhu, Xuekang, et al.
Published: (2024)
by: Zhu, Xuekang, et al.
Published: (2024)
Generalizable Object Re-Identification via Visual In-Context Prompting
by: Huang, Zhizhong, et al.
Published: (2025)
by: Huang, Zhizhong, et al.
Published: (2025)
Transparent Visual Reasoning via Object-Centric Agent Collaboration
by: Teoh, Benjamin, et al.
Published: (2025)
by: Teoh, Benjamin, et al.
Published: (2025)
A Study of Commonsense Reasoning over Visual Object Properties
by: Kolari, Abhishek, et al.
Published: (2025)
by: Kolari, Abhishek, et al.
Published: (2025)
Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
by: Guo, Garvin, et al.
Published: (2026)
by: Guo, Garvin, et al.
Published: (2026)
Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph Reasoning
by: Wan, Xixi, et al.
Published: (2025)
by: Wan, Xixi, et al.
Published: (2025)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Transformer for Object Re-Identification: A Survey
by: Ye, Mang, et al.
Published: (2024)
by: Ye, Mang, et al.
Published: (2024)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
by: Guo, Guangfu, et al.
Published: (2026)
by: Guo, Guangfu, et al.
Published: (2026)
Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning
by: Cheng, Hanbo, et al.
Published: (2026)
by: Cheng, Hanbo, et al.
Published: (2026)
EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric Vision
by: Cao, Yifei, et al.
Published: (2025)
by: Cao, Yifei, et al.
Published: (2025)
Beyond Pixels: Vector-to-Graph Transformation for Reliable Schematic Auditing
by: Ma, Chengwei, et al.
Published: (2026)
by: Ma, Chengwei, et al.
Published: (2026)
Beyond Appearance: Transformer-based Person Identification from Conversational Dynamics
by: Chapariniya, Masoumeh, et al.
Published: (2025)
by: Chapariniya, Masoumeh, et al.
Published: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Unseen Object Reasoning with Shared Appearance Cues
by: Singh, Paridhi, et al.
Published: (2024)
by: Singh, Paridhi, et al.
Published: (2024)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
by: Cao, Min, et al.
Published: (2025)
by: Cao, Min, et al.
Published: (2025)
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
by: Shi, Fan, et al.
Published: (2025)
by: Shi, Fan, et al.
Published: (2025)
Visual Grounding for Object-Level Generalization in Reinforcement Learning
by: Jiang, Haobin, et al.
Published: (2024)
by: Jiang, Haobin, et al.
Published: (2024)
Research about the Ability of LLM in the Tamper-Detection Area
by: Yang, Xinyu, et al.
Published: (2024)
by: Yang, Xinyu, et al.
Published: (2024)
Why Settle for One? Text-to-ImageSet Generation and Evaluation
by: Jia, Chengyou, et al.
Published: (2025)
by: Jia, Chengyou, et al.
Published: (2025)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning
by: Lv, Guannan, et al.
Published: (2026)
by: Lv, Guannan, et al.
Published: (2026)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
by: Liu, Ye, et al.
Published: (2025)
by: Liu, Ye, et al.
Published: (2025)
UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Beyond Bare Queries: Open-Vocabulary Object Grounding with 3D Scene Graph
by: Linok, Sergey, et al.
Published: (2024)
by: Linok, Sergey, et al.
Published: (2024)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
$\mathrm{D}^\mathrm{3}$-Predictor: Noise-Free Deterministic Diffusion for Dense Prediction
by: Xia, Changliang, et al.
Published: (2025)
by: Xia, Changliang, et al.
Published: (2025)
Similar Items
-
SHAN: Object-Level Privacy Detection via Inference on Scene Heterogeneous Graph
by: Jiang, Zhuohang, et al.
Published: (2024) -
IML-ViT: Benchmarking Image Manipulation Localization by Vision Transformer
by: Ma, Xiaochen, et al.
Published: (2023) -
IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization
by: Ma, Xiaochen, et al.
Published: (2024) -
Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
by: Tong, Bingkui, et al.
Published: (2025) -
M^3:Manipulation Mask Manufacturer for Arbitrary-Scale Super-Resolution Mask
by: Yang, Xinyu, et al.
Published: (2024)