Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
Fuente:
arXiv
Guardado en:
| Autores principales: | Wen, Ziqi, Madinei, Parsa, Eckstein, Miguel P. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
por: Skaza, Jonathan, et al.
Publicado: (2025)
por: Skaza, Jonathan, et al.
Publicado: (2025)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
por: Murlidaran, Shravan, et al.
Publicado: (2026)
por: Murlidaran, Shravan, et al.
Publicado: (2026)
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
por: Madinei, Parsa, et al.
Publicado: (2025)
por: Madinei, Parsa, et al.
Publicado: (2025)
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
por: Madinei, Parsa, et al.
Publicado: (2026)
por: Madinei, Parsa, et al.
Publicado: (2026)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
por: Sun, Penglei, et al.
Publicado: (2025)
por: Sun, Penglei, et al.
Publicado: (2025)
Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction
por: Bejerano, Emily, et al.
Publicado: (2026)
por: Bejerano, Emily, et al.
Publicado: (2026)
Revealing the Semantic Selection Gap in DINOv3 through Training-Free Few-Shot Segmentation
por: Zakir, Hussni Mohd, et al.
Publicado: (2026)
por: Zakir, Hussni Mohd, et al.
Publicado: (2026)
Saliency Suppressed, Semantics Surfaced: Visual Transformations in Neural Networks and the Brain
por: Opiełka, Gustaw, et al.
Publicado: (2024)
por: Opiełka, Gustaw, et al.
Publicado: (2024)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
por: Liu, Hanqing, et al.
Publicado: (2026)
por: Liu, Hanqing, et al.
Publicado: (2026)
SUM: Saliency Unification through Mamba for Visual Attention Modeling
por: Hosseini, Alireza, et al.
Publicado: (2024)
por: Hosseini, Alireza, et al.
Publicado: (2024)
Generalist Multimodal LLMs Gain Biometric Expertise via Human Salience
por: Piland, Jacob, et al.
Publicado: (2026)
por: Piland, Jacob, et al.
Publicado: (2026)
Structure Your Data: Towards Semantic Graph Counterfactuals
por: Dimitriou, Angeliki, et al.
Publicado: (2024)
por: Dimitriou, Angeliki, et al.
Publicado: (2024)
Belief-Aware VLM Model for Human-like Reasoning
por: Nayak, Anshul, et al.
Publicado: (2026)
por: Nayak, Anshul, et al.
Publicado: (2026)
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
por: Wen, Ziqi, et al.
Publicado: (2025)
por: Wen, Ziqi, et al.
Publicado: (2025)
Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
por: Mao, Shunqi, et al.
Publicado: (2025)
por: Mao, Shunqi, et al.
Publicado: (2025)
InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement
por: Zou, Yude, et al.
Publicado: (2026)
por: Zou, Yude, et al.
Publicado: (2026)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
por: Chen, Chuang, et al.
Publicado: (2024)
por: Chen, Chuang, et al.
Publicado: (2024)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
por: Carvalho, Miguel, et al.
Publicado: (2025)
por: Carvalho, Miguel, et al.
Publicado: (2025)
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
por: Wang, Wentao, et al.
Publicado: (2025)
por: Wang, Wentao, et al.
Publicado: (2025)
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
por: Li, Tianqin, et al.
Publicado: (2025)
por: Li, Tianqin, et al.
Publicado: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
por: Wei, Zhihua, et al.
Publicado: (2026)
por: Wei, Zhihua, et al.
Publicado: (2026)
FaceSaliencyAug: Mitigating Geographic, Gender and Stereotypical Biases via Saliency-Based Data Augmentation
por: Kumar, Teerath, et al.
Publicado: (2024)
por: Kumar, Teerath, et al.
Publicado: (2024)
Privacy-Concealing Cooperative Perception for BEV Scene Segmentation
por: Wang, Song, et al.
Publicado: (2026)
por: Wang, Song, et al.
Publicado: (2026)
Pan-FM: A Pan-Organ Foundation Model with Saliency-Guided Masking for Missing Robustness
por: Wu, Qiangqiang, et al.
Publicado: (2026)
por: Wu, Qiangqiang, et al.
Publicado: (2026)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
por: Kang, Donggoo, et al.
Publicado: (2024)
por: Kang, Donggoo, et al.
Publicado: (2024)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
por: Cho, Jungbin, et al.
Publicado: (2025)
por: Cho, Jungbin, et al.
Publicado: (2025)
Diffusion Counterfactual Generation with Semantic Abduction
por: Rasal, Rajat, et al.
Publicado: (2025)
por: Rasal, Rajat, et al.
Publicado: (2025)
Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
por: Hermosilla, Pedro, et al.
Publicado: (2025)
por: Hermosilla, Pedro, et al.
Publicado: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
por: Pani, Anupam, et al.
Publicado: (2025)
por: Pani, Anupam, et al.
Publicado: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
por: Xun, Yuan, et al.
Publicado: (2024)
por: Xun, Yuan, et al.
Publicado: (2024)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
por: Nguyen, Vinh
Publicado: (2024)
por: Nguyen, Vinh
Publicado: (2024)
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
por: Wang, Kangkang, et al.
Publicado: (2026)
por: Wang, Kangkang, et al.
Publicado: (2026)
Salience Adjustment for Context-Based Emotion Recognition
por: Han, Bin, et al.
Publicado: (2025)
por: Han, Bin, et al.
Publicado: (2025)
Diffusion Features to Bridge Domain Gap for Semantic Segmentation
por: Ji, Yuxiang, et al.
Publicado: (2024)
por: Ji, Yuxiang, et al.
Publicado: (2024)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
por: Liu, Tianhui, et al.
Publicado: (2026)
por: Liu, Tianhui, et al.
Publicado: (2026)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
por: Tang, Tianci, et al.
Publicado: (2026)
por: Tang, Tianci, et al.
Publicado: (2026)
BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion
por: Kim, Gwanghyun, et al.
Publicado: (2024)
por: Kim, Gwanghyun, et al.
Publicado: (2024)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
por: Singh, Aditya Kumar, et al.
Publicado: (2026)
por: Singh, Aditya Kumar, et al.
Publicado: (2026)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
por: Qian, Chen, et al.
Publicado: (2025)
por: Qian, Chen, et al.
Publicado: (2025)
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
por: Zhong, Ziqi, et al.
Publicado: (2025)
por: Zhong, Ziqi, et al.
Publicado: (2025)
Ejemplares similares
-
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
por: Skaza, Jonathan, et al.
Publicado: (2025) -
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
por: Murlidaran, Shravan, et al.
Publicado: (2026) -
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
por: Madinei, Parsa, et al.
Publicado: (2025) -
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
por: Madinei, Parsa, et al.
Publicado: (2026) -
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
por: Sun, Penglei, et al.
Publicado: (2025)