Salvato in:
| Autori principali: | Wen, Ziqi, Madinei, Parsa, Eckstein, Miguel P. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.13047 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
di: Skaza, Jonathan, et al.
Pubblicazione: (2025)
di: Skaza, Jonathan, et al.
Pubblicazione: (2025)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
di: Murlidaran, Shravan, et al.
Pubblicazione: (2026)
di: Murlidaran, Shravan, et al.
Pubblicazione: (2026)
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
di: Madinei, Parsa, et al.
Pubblicazione: (2025)
di: Madinei, Parsa, et al.
Pubblicazione: (2025)
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
di: Madinei, Parsa, et al.
Pubblicazione: (2026)
di: Madinei, Parsa, et al.
Pubblicazione: (2026)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
di: Sun, Penglei, et al.
Pubblicazione: (2025)
di: Sun, Penglei, et al.
Pubblicazione: (2025)
Sim2Radar: Toward Bridging the Radar Sim-to-Real Gap with VLM-Guided Scene Reconstruction
di: Bejerano, Emily, et al.
Pubblicazione: (2026)
di: Bejerano, Emily, et al.
Pubblicazione: (2026)
Revealing the Semantic Selection Gap in DINOv3 through Training-Free Few-Shot Segmentation
di: Zakir, Hussni Mohd, et al.
Pubblicazione: (2026)
di: Zakir, Hussni Mohd, et al.
Pubblicazione: (2026)
Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps
di: Wen, Ziqi, et al.
Pubblicazione: (2025)
di: Wen, Ziqi, et al.
Pubblicazione: (2025)
Saliency Suppressed, Semantics Surfaced: Visual Transformations in Neural Networks and the Brain
di: Opiełka, Gustaw, et al.
Pubblicazione: (2024)
di: Opiełka, Gustaw, et al.
Pubblicazione: (2024)
RE-VLM: Event-Augmented Vision-Language Model for Scene Understanding
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
di: Liu, Hanqing, et al.
Pubblicazione: (2026)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
di: Carvalho, Miguel, et al.
Pubblicazione: (2025)
di: Carvalho, Miguel, et al.
Pubblicazione: (2025)
SUM: Saliency Unification through Mamba for Visual Attention Modeling
di: Hosseini, Alireza, et al.
Pubblicazione: (2024)
di: Hosseini, Alireza, et al.
Pubblicazione: (2024)
Generalist Multimodal LLMs Gain Biometric Expertise via Human Salience
di: Piland, Jacob, et al.
Pubblicazione: (2026)
di: Piland, Jacob, et al.
Pubblicazione: (2026)
Structure Your Data: Towards Semantic Graph Counterfactuals
di: Dimitriou, Angeliki, et al.
Pubblicazione: (2024)
di: Dimitriou, Angeliki, et al.
Pubblicazione: (2024)
Belief-Aware VLM Model for Human-like Reasoning
di: Nayak, Anshul, et al.
Pubblicazione: (2026)
di: Nayak, Anshul, et al.
Pubblicazione: (2026)
Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding
di: Mao, Shunqi, et al.
Pubblicazione: (2025)
di: Mao, Shunqi, et al.
Pubblicazione: (2025)
InfBaGel: Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement
di: Zou, Yude, et al.
Pubblicazione: (2026)
di: Zou, Yude, et al.
Pubblicazione: (2026)
UniEmoX: Cross-modal Semantic-Guided Large-Scale Pretraining for Universal Scene Emotion Perception
di: Chen, Chuang, et al.
Pubblicazione: (2024)
di: Chen, Chuang, et al.
Pubblicazione: (2024)
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
di: Li, Tianqin, et al.
Pubblicazione: (2025)
di: Li, Tianqin, et al.
Pubblicazione: (2025)
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
di: Wang, Wentao, et al.
Pubblicazione: (2025)
di: Wang, Wentao, et al.
Pubblicazione: (2025)
Diffusion Counterfactual Generation with Semantic Abduction
di: Rasal, Rajat, et al.
Pubblicazione: (2025)
di: Rasal, Rajat, et al.
Pubblicazione: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
FaceSaliencyAug: Mitigating Geographic, Gender and Stereotypical Biases via Saliency-Based Data Augmentation
di: Kumar, Teerath, et al.
Pubblicazione: (2024)
di: Kumar, Teerath, et al.
Pubblicazione: (2024)
Pan-FM: A Pan-Organ Foundation Model with Saliency-Guided Masking for Missing Robustness
di: Wu, Qiangqiang, et al.
Pubblicazione: (2026)
di: Wu, Qiangqiang, et al.
Pubblicazione: (2026)
Privacy-Concealing Cooperative Perception for BEV Scene Segmentation
di: Wang, Song, et al.
Pubblicazione: (2026)
di: Wang, Song, et al.
Pubblicazione: (2026)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
di: Kang, Donggoo, et al.
Pubblicazione: (2024)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
di: Cho, Jungbin, et al.
Pubblicazione: (2025)
di: Cho, Jungbin, et al.
Pubblicazione: (2025)
Masked Scene Modeling: Narrowing the Gap Between Supervised and Self-Supervised Learning in 3D Scene Understanding
di: Hermosilla, Pedro, et al.
Pubblicazione: (2025)
di: Hermosilla, Pedro, et al.
Pubblicazione: (2025)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
di: Liu, Tianhui, et al.
Pubblicazione: (2026)
di: Liu, Tianhui, et al.
Pubblicazione: (2026)
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
di: Zhong, Ziqi, et al.
Pubblicazione: (2025)
di: Zhong, Ziqi, et al.
Pubblicazione: (2025)
Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
di: Pani, Anupam, et al.
Pubblicazione: (2025)
di: Pani, Anupam, et al.
Pubblicazione: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
di: Xun, Yuan, et al.
Pubblicazione: (2024)
di: Xun, Yuan, et al.
Pubblicazione: (2024)
MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models
di: Wang, Kangkang, et al.
Pubblicazione: (2026)
di: Wang, Kangkang, et al.
Pubblicazione: (2026)
Salience Adjustment for Context-Based Emotion Recognition
di: Han, Bin, et al.
Pubblicazione: (2025)
di: Han, Bin, et al.
Pubblicazione: (2025)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
di: Nguyen, Vinh
Pubblicazione: (2024)
di: Nguyen, Vinh
Pubblicazione: (2024)
Learning User Embeddings from Human Gaze for Personalised Saliency Prediction
di: Strohm, Florian, et al.
Pubblicazione: (2024)
di: Strohm, Florian, et al.
Pubblicazione: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
Diffusion Features to Bridge Domain Gap for Semantic Segmentation
di: Ji, Yuxiang, et al.
Pubblicazione: (2024)
di: Ji, Yuxiang, et al.
Pubblicazione: (2024)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
di: Tang, Tianci, et al.
Pubblicazione: (2026)
di: Tang, Tianci, et al.
Pubblicazione: (2026)
Explaining Low Perception Model Competency with High-Competency Counterfactuals
di: Pohland, Sara, et al.
Pubblicazione: (2025)
di: Pohland, Sara, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DReX: Pure Vision Fusion of Self-Supervised and Convolutional Representations for Image Complexity Prediction
di: Skaza, Jonathan, et al.
Pubblicazione: (2025) -
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
di: Murlidaran, Shravan, et al.
Pubblicazione: (2026) -
INTERLACE: Interleaved Layer Pruning and Efficient Adaptation in Large Vision-Language Models
di: Madinei, Parsa, et al.
Pubblicazione: (2025) -
IRIS: Intent Resolution via Inference-time Saccades for Open-Ended VQA in Large Vision-Language Models
di: Madinei, Parsa, et al.
Pubblicazione: (2026) -
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
di: Sun, Penglei, et al.
Pubblicazione: (2025)