Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Zhuowan, Xie, Cihang, Van Durme, Benjamin, Yuille, Alan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
di: Xie, Yunfei, et al.
Pubblicazione: (2024)
di: Xie, Yunfei, et al.
Pubblicazione: (2024)
Rejuvenating image-GPT as Strong Visual Representation Learners
di: Ren, Sucheng, et al.
Pubblicazione: (2023)
di: Ren, Sucheng, et al.
Pubblicazione: (2023)
SPFormer: Enhancing Vision Transformer with Superpixel Representation
di: Mei, Jieru, et al.
Pubblicazione: (2024)
di: Mei, Jieru, et al.
Pubblicazione: (2024)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
di: Xiao, Junfei, et al.
Pubblicazione: (2023)
di: Xiao, Junfei, et al.
Pubblicazione: (2023)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
di: Zhu, Jing, et al.
Pubblicazione: (2025)
di: Zhu, Jing, et al.
Pubblicazione: (2025)
Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
di: Chen, Meiqi, et al.
Pubblicazione: (2024)
Grounding Partially-Defined Events in Multimodal Data
di: Sanders, Kate, et al.
Pubblicazione: (2024)
di: Sanders, Kate, et al.
Pubblicazione: (2024)
M-VAR: Decoupled Scale-wise Autoregressive Modeling for High-Quality Image Generation
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
di: Ren, Sucheng, et al.
Pubblicazione: (2024)
Play to Generalize: Learning to Reason Through Game Play
di: Xie, Yunfei, et al.
Pubblicazione: (2025)
di: Xie, Yunfei, et al.
Pubblicazione: (2025)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
di: Yang, Timing, et al.
Pubblicazione: (2025)
di: Yang, Timing, et al.
Pubblicazione: (2025)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
di: Alper, Morris, et al.
Pubblicazione: (2024)
di: Alper, Morris, et al.
Pubblicazione: (2024)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
di: Sanders, Kate, et al.
Pubblicazione: (2024)
di: Sanders, Kate, et al.
Pubblicazione: (2024)
ViLBench: A Suite for Vision-Language Process Reward Modeling
di: Tu, Haoqin, et al.
Pubblicazione: (2025)
di: Tu, Haoqin, et al.
Pubblicazione: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
di: Martin, Alexander, et al.
Pubblicazione: (2025)
di: Martin, Alexander, et al.
Pubblicazione: (2025)
Semantic Alignment of Unimodal Medical Text and Vision Representations
di: Di Folco, Maxime, et al.
Pubblicazione: (2025)
di: Di Folco, Maxime, et al.
Pubblicazione: (2025)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
ViT-5: Vision Transformers for The Mid-2020s
di: Wang, Feng, et al.
Pubblicazione: (2026)
di: Wang, Feng, et al.
Pubblicazione: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
di: Wang, Feng, et al.
Pubblicazione: (2025)
di: Wang, Feng, et al.
Pubblicazione: (2025)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
di: Yang, Yuwei, et al.
Pubblicazione: (2025)
Visual Representations inside the Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2025)
di: Liu, Benlin, et al.
Pubblicazione: (2025)
Frequency-Modulated Visual Restoration for Matryoshka Large Multimodal Models
di: Pan, Qingtao, et al.
Pubblicazione: (2026)
di: Pan, Qingtao, et al.
Pubblicazione: (2026)
WikiVideo: Article Generation from Multiple Videos
di: Martin, Alexander, et al.
Pubblicazione: (2025)
di: Martin, Alexander, et al.
Pubblicazione: (2025)
Multimodal Representation Learning by Alternating Unimodal Adaptation
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
di: Zhang, Xiaohui, et al.
Pubblicazione: (2023)
Multi-Vector Index Compression in Any Modality
di: Qin, Hanxiang, et al.
Pubblicazione: (2026)
di: Qin, Hanxiang, et al.
Pubblicazione: (2026)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2024)
di: Hu, Yushi, et al.
Pubblicazione: (2024)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
di: Wu, Juncheng, et al.
Pubblicazione: (2026)
di: Wu, Juncheng, et al.
Pubblicazione: (2026)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
di: Lawrence, Logan, et al.
Pubblicazione: (2025)
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
di: Sanders, Kate, et al.
Pubblicazione: (2025)
di: Sanders, Kate, et al.
Pubblicazione: (2025)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
di: Luo, Fuwen, et al.
Pubblicazione: (2024)
di: Luo, Fuwen, et al.
Pubblicazione: (2024)
Captain Safari: A World Engine with Pose-Aligned 3D Memory
di: Chou, Yu-Cheng, et al.
Pubblicazione: (2025)
di: Chou, Yu-Cheng, et al.
Pubblicazione: (2025)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
di: Liu, Haoyu, et al.
Pubblicazione: (2026)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
di: Rostamkhani, Mohammadmostafa, et al.
Pubblicazione: (2024)
di: Rostamkhani, Mohammadmostafa, et al.
Pubblicazione: (2024)
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
di: Wang, Weiyun, et al.
Pubblicazione: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
di: Batra, Hunar, et al.
Pubblicazione: (2025)
di: Batra, Hunar, et al.
Pubblicazione: (2025)
VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation
di: Dao, Alan, et al.
Pubblicazione: (2025)
di: Dao, Alan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
di: Wang, Yuxuan, et al.
Pubblicazione: (2024) -
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
di: Xie, Yunfei, et al.
Pubblicazione: (2024) -
Rejuvenating image-GPT as Strong Visual Representation Learners
di: Ren, Sucheng, et al.
Pubblicazione: (2023) -
SPFormer: Enhancing Vision Transformer with Superpixel Representation
di: Mei, Jieru, et al.
Pubblicazione: (2024) -
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
di: Xiao, Junfei, et al.
Pubblicazione: (2023)