Gespeichert in:
| Hauptverfasser: | Gizdov, Andrey, Procopio, Andrea, Li, Yichen, Harari, Daniel, Ullman, Tomer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.12486 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards aligned body representations in vision models
von: Gizdov, Andrey, et al.
Veröffentlicht: (2025)
von: Gizdov, Andrey, et al.
Veröffentlicht: (2025)
Time-Series at the Edge: Tiny Separable CNNs for Wearable Gait Detection and Optimal Sensor Placement
von: Procopio, Andrea, et al.
Veröffentlicht: (2025)
von: Procopio, Andrea, et al.
Veröffentlicht: (2025)
Chain of Time: In-Context Physical Simulation with Image Generation Models
von: Wang, YingQiao, et al.
Veröffentlicht: (2025)
von: Wang, YingQiao, et al.
Veröffentlicht: (2025)
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
von: Ullman, Tomer
Veröffentlicht: (2024)
von: Ullman, Tomer
Veröffentlicht: (2024)
Category Query Learning for Human-Object Interaction Classification
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
von: Kang, Donggoo, et al.
Veröffentlicht: (2024)
von: Kang, Donggoo, et al.
Veröffentlicht: (2024)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
Like Humans to Few-Shot Learning through Knowledge Permeation of Vision and Text
von: Jia, Yuyu, et al.
Veröffentlicht: (2024)
von: Jia, Yuyu, et al.
Veröffentlicht: (2024)
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
von: Liu, Shanhui, et al.
Veröffentlicht: (2025)
von: Liu, Shanhui, et al.
Veröffentlicht: (2025)
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
von: Aniraj, Ananthu, et al.
Veröffentlicht: (2025)
von: Aniraj, Ananthu, et al.
Veröffentlicht: (2025)
SynHLMA:Synthesizing Hand Language Manipulation for Articulated Object with Discrete Human Object Interaction Representation
von: zhi, Wang, et al.
Veröffentlicht: (2025)
von: zhi, Wang, et al.
Veröffentlicht: (2025)
CoTZero: Annotation-Free Human-Like Vision Reasoning via Hierarchical Synthetic CoT
von: Du, Chengyi, et al.
Veröffentlicht: (2026)
von: Du, Chengyi, et al.
Veröffentlicht: (2026)
Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment
von: Takahashi, Soh, et al.
Veröffentlicht: (2025)
von: Takahashi, Soh, et al.
Veröffentlicht: (2025)
Jack of All Tasks, Master of Many: Designing General-purpose Coarse-to-Fine Vision-Language Model
von: Pramanick, Shraman, et al.
Veröffentlicht: (2023)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2023)
DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
von: Shi, Zhiyi, et al.
Veröffentlicht: (2025)
von: Shi, Zhiyi, et al.
Veröffentlicht: (2025)
When Large Vision-Language Model Meets Large Remote Sensing Imagery: Coarse-to-Fine Text-Guided Token Pruning
von: Luo, Junwei, et al.
Veröffentlicht: (2025)
von: Luo, Junwei, et al.
Veröffentlicht: (2025)
MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models
von: Yang, Garry, et al.
Veröffentlicht: (2025)
von: Yang, Garry, et al.
Veröffentlicht: (2025)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
The Geometry of Representational Failures in Vision Language Models
von: Savietto, Daniele, et al.
Veröffentlicht: (2026)
von: Savietto, Daniele, et al.
Veröffentlicht: (2026)
Human-Object Interaction from Human-Level Instructions
von: Wu, Zhen, et al.
Veröffentlicht: (2024)
von: Wu, Zhen, et al.
Veröffentlicht: (2024)
MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
von: Li, Chenxi, et al.
Veröffentlicht: (2025)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
von: Du, Fan, et al.
Veröffentlicht: (2026)
von: Du, Fan, et al.
Veröffentlicht: (2026)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
von: Cao, Zongsheng, et al.
Veröffentlicht: (2025)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
von: He, Yuting, et al.
Veröffentlicht: (2026)
von: He, Yuting, et al.
Veröffentlicht: (2026)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
Controllable Video Object Insertion via Multiview Priors
von: Qi, Xia, et al.
Veröffentlicht: (2026)
von: Qi, Xia, et al.
Veröffentlicht: (2026)
FMNet: Frequency-Assisted Mamba-Like Linear Attention Network for Camouflaged Object Detection
von: Deng, Ming, et al.
Veröffentlicht: (2025)
von: Deng, Ming, et al.
Veröffentlicht: (2025)
Deep Extrinsic Manifold Representation for Vision Tasks
von: Zhang, Tongtong, et al.
Veröffentlicht: (2024)
von: Zhang, Tongtong, et al.
Veröffentlicht: (2024)
Untrained neural networks can demonstrate memorization-independent abstract reasoning
von: Barak, Tomer, et al.
Veröffentlicht: (2024)
von: Barak, Tomer, et al.
Veröffentlicht: (2024)
Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
von: Ding, Rui, et al.
Veröffentlicht: (2024)
von: Ding, Rui, et al.
Veröffentlicht: (2024)
CoMa: Contextual Massing Generation with Vision-Language Models
von: Maslov, Evgenii, et al.
Veröffentlicht: (2026)
von: Maslov, Evgenii, et al.
Veröffentlicht: (2026)
Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation
von: Zhu, Chunzheng, et al.
Veröffentlicht: (2026)
von: Zhu, Chunzheng, et al.
Veröffentlicht: (2026)
Multi-Object Hallucination in Vision-Language Models
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
von: Zeng, Haoxi, et al.
Veröffentlicht: (2025)
von: Zeng, Haoxi, et al.
Veröffentlicht: (2025)
Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards aligned body representations in vision models
von: Gizdov, Andrey, et al.
Veröffentlicht: (2025) -
Time-Series at the Edge: Tiny Separable CNNs for Wearable Gait Detection and Optimal Sensor Placement
von: Procopio, Andrea, et al.
Veröffentlicht: (2025) -
Chain of Time: In-Context Physical Simulation with Image Generation Models
von: Wang, YingQiao, et al.
Veröffentlicht: (2025) -
The Illusion-Illusion: Vision Language Models See Illusions Where There are None
von: Ullman, Tomer
Veröffentlicht: (2024) -
Category Query Learning for Human-Object Interaction Classification
von: Xie, Chi, et al.
Veröffentlicht: (2023)