Birds of a Feather Flock Together: Background-Invariant Representations via Linear Structure in VLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zaazou, Youssef, Thomas, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
por: Burdisso, Sergio, et al.
Publicado: (2024)
por: Burdisso, Sergio, et al.
Publicado: (2024)
Same Answer, Different Representations: Hidden instability in VLMs
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
por: Wani, Farooq Ahmad, et al.
Publicado: (2026)
Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations
por: Mohammed, Fatma Youssef, et al.
Publicado: (2025)
por: Mohammed, Fatma Youssef, et al.
Publicado: (2025)
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
por: Kang, Inha, et al.
Publicado: (2025)
por: Kang, Inha, et al.
Publicado: (2025)
Rethink MAE with Linear Time-Invariant Dynamics
por: Wang, Zice
Publicado: (2026)
por: Wang, Zice
Publicado: (2026)
Fine-grained Background Representation for Weakly Supervised Semantic Segmentation
por: Yin, Xu, et al.
Publicado: (2024)
por: Yin, Xu, et al.
Publicado: (2024)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
por: Berman, Shmuel, et al.
Publicado: (2025)
por: Berman, Shmuel, et al.
Publicado: (2025)
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
por: Kang, Hyeonsu, et al.
Publicado: (2025)
por: Kang, Hyeonsu, et al.
Publicado: (2025)
Learning Content-Aware Multi-Modal Joint Input Pruning via Bird's-Eye-View Representation
por: Li, Yuxin, et al.
Publicado: (2024)
por: Li, Yuxin, et al.
Publicado: (2024)
Towards a Universal 3D Medical Multi-modality Generalization via Learning Personalized Invariant Representation
por: Tan, Zhaorui, et al.
Publicado: (2024)
por: Tan, Zhaorui, et al.
Publicado: (2024)
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
por: Chen, Yanlong, et al.
Publicado: (2026)
por: Chen, Yanlong, et al.
Publicado: (2026)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
por: Liang, Yue, et al.
Publicado: (2026)
por: Liang, Yue, et al.
Publicado: (2026)
Caption This, Reason That: VLMs Caught in the Middle
por: Weng, Zihan, et al.
Publicado: (2025)
por: Weng, Zihan, et al.
Publicado: (2025)
Line of Sight: On Linear Representations in VLLMs
por: Rajaram, Achyuta, et al.
Publicado: (2025)
por: Rajaram, Achyuta, et al.
Publicado: (2025)
Evaluating Compositional Generalisation in VLMs and Diffusion Models
por: Pearson, Beth, et al.
Publicado: (2025)
por: Pearson, Beth, et al.
Publicado: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
por: Xu, Zhengzhuo, et al.
Publicado: (2025)
por: Xu, Zhengzhuo, et al.
Publicado: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
por: Gambashidze, Alexander, et al.
Publicado: (2025)
por: Gambashidze, Alexander, et al.
Publicado: (2025)
WIDIn: Wording Image for Domain-Invariant Representation in Single-Source Domain Generalization
por: Ma, Jiawei, et al.
Publicado: (2024)
por: Ma, Jiawei, et al.
Publicado: (2024)
Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization
por: Lan, Yuqin, et al.
Publicado: (2026)
por: Lan, Yuqin, et al.
Publicado: (2026)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
por: Zhang, Yuyou, et al.
Publicado: (2025)
por: Zhang, Yuyou, et al.
Publicado: (2025)
WildIng: A Wildlife Image Invariant Representation Model for Geographical Domain Shift
por: Santamaria, Julian D., et al.
Publicado: (2026)
por: Santamaria, Julian D., et al.
Publicado: (2026)
Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
por: Zhu, Chunzheng, et al.
Publicado: (2026)
por: Zhu, Chunzheng, et al.
Publicado: (2026)
Stateful Token Reduction for Long-Video Hybrid VLMs
por: Jiang, Jindong, et al.
Publicado: (2026)
por: Jiang, Jindong, et al.
Publicado: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
por: Clark, Christopher, et al.
Publicado: (2026)
por: Clark, Christopher, et al.
Publicado: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
por: Ge, Yuyao, et al.
Publicado: (2025)
por: Ge, Yuyao, et al.
Publicado: (2025)
Towards Lossless Ultimate Vision Token Compression for VLMs
por: Zheng, Dehua, et al.
Publicado: (2025)
por: Zheng, Dehua, et al.
Publicado: (2025)
Treble Counterfactual VLMs: A Causal Approach to Hallucination
por: Li, Shawn, et al.
Publicado: (2025)
por: Li, Shawn, et al.
Publicado: (2025)
Birds of a Different Feather Flock Together: Exploring Opportunities and Challenges in Animal-Human-Machine Teaming
por: Cohen, Myke C., et al.
Publicado: (2025)
por: Cohen, Myke C., et al.
Publicado: (2025)
D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
por: Huang, Yiyang, et al.
Publicado: (2025)
por: Huang, Yiyang, et al.
Publicado: (2025)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
por: Park, Jaehyun, et al.
Publicado: (2026)
por: Park, Jaehyun, et al.
Publicado: (2026)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
por: Hong, Rui, et al.
Publicado: (2026)
por: Hong, Rui, et al.
Publicado: (2026)
Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks
por: Li, Chenjun
Publicado: (2026)
por: Li, Chenjun
Publicado: (2026)
Drive-KD: Multi-Teacher Distillation for VLMs in Autonomous Driving
por: Lian, Weitong, et al.
Publicado: (2026)
por: Lian, Weitong, et al.
Publicado: (2026)
T2T-VICL: Unlocking the Boundaries of Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
por: Xia, Shao-Jun, et al.
Publicado: (2025)
por: Xia, Shao-Jun, et al.
Publicado: (2025)
TAPS : Frustratingly Simple Test Time Active Learning for VLMs
por: Sarkar, Dhruv, et al.
Publicado: (2025)
por: Sarkar, Dhruv, et al.
Publicado: (2025)
DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
por: Wang, Zhenhailong, et al.
Publicado: (2025)
por: Wang, Zhenhailong, et al.
Publicado: (2025)
PolarBEVDet: Exploring Polar Representation for Multi-View 3D Object Detection in Bird's-Eye-View
por: Yu, Zichen, et al.
Publicado: (2024)
por: Yu, Zichen, et al.
Publicado: (2024)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
por: Khan, Ufaq, et al.
Publicado: (2026)
por: Khan, Ufaq, et al.
Publicado: (2026)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
3D Primitives are a Spatial Language for VLMs
por: Liu, Junze, et al.
Publicado: (2026)
por: Liu, Junze, et al.
Publicado: (2026)
Ejemplares similares
-
Reliability Estimation of News Media Sources: Birds of a Feather Flock Together
por: Burdisso, Sergio, et al.
Publicado: (2024) -
Same Answer, Different Representations: Hidden instability in VLMs
por: Wani, Farooq Ahmad, et al.
Publicado: (2026) -
Unified Attention Modeling for Efficient Free-Viewing and Visual Search via Shared Representations
por: Mohammed, Fatma Youssef, et al.
Publicado: (2025) -
What "Not" to Detect: Negation-Aware VLMs via Structured Reasoning and Token Merging
por: Kang, Inha, et al.
Publicado: (2025) -
Rethink MAE with Linear Time-Invariant Dynamics
por: Wang, Zice
Publicado: (2026)