Iterated Learning Improves Compositionality in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Chenhao, Zhang, Jieyu, Kembhavi, Aniruddha, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026)
Task Me Anything
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
One Diffusion to Generate Them All
von: Le, Duong H., et al.
Veröffentlicht: (2024)
von: Le, Duong H., et al.
Veröffentlicht: (2024)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
von: Duggal, Shivam, et al.
Veröffentlicht: (2025)
LATTE: Learning to Think with Vision Specialists
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
Video-Based Reward Modeling for Computer-Use Agents
von: Song, Linxin, et al.
Veröffentlicht: (2026)
von: Song, Linxin, et al.
Veröffentlicht: (2026)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
von: Liu, Shuo, et al.
Veröffentlicht: (2026)
SPARO: Selective Attention for Robust and Compositional Transformer Encodings for Vision
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
von: Vani, Ankit, et al.
Veröffentlicht: (2024)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
MIMIC: Masked Image Modeling with Image Correspondences
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
von: Marathe, Kalyani, et al.
Veröffentlicht: (2023)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
von: Yang, Yue, et al.
Veröffentlicht: (2025)
von: Yang, Yue, et al.
Veröffentlicht: (2025)
Multilingual Diversity Improves Vision-Language Representations
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
von: Nguyen, Thao, et al.
Veröffentlicht: (2024)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
von: Ikezogwo, Wisdom O., et al.
Veröffentlicht: (2025)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
von: Ray, Arijit, et al.
Veröffentlicht: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation
von: Su, Xia, et al.
Veröffentlicht: (2026)
von: Su, Xia, et al.
Veröffentlicht: (2026)
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
von: Yang, Le, et al.
Veröffentlicht: (2024)
von: Yang, Le, et al.
Veröffentlicht: (2024)
Improving Large Vision-Language Models' Understanding for Flow Field Data
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaomei, et al.
Veröffentlicht: (2025)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
Seeing the Unseen: Visual Common Sense for Semantic Placement
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2024)
von: Ramrakhya, Ram, et al.
Veröffentlicht: (2024)
Diversity Covariance-Aware Prompt Learning for Vision-Language Models
von: Dong, Songlin, et al.
Veröffentlicht: (2025)
von: Dong, Songlin, et al.
Veröffentlicht: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
In-Context Learning Improves Compositional Understanding of Vision-Language Models
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
Visual Representations inside the Language Model
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
Natural Language Inference Improves Compositionality in Vision-Language Models
von: Cascante-Bonilla, Paola, et al.
Veröffentlicht: (2024)
von: Cascante-Bonilla, Paola, et al.
Veröffentlicht: (2024)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
von: He, Qijia, et al.
Veröffentlicht: (2026)
von: He, Qijia, et al.
Veröffentlicht: (2026)
Language-Guided Token Compression with Reinforcement Learning in Large Vision-Language Models
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
von: Cao, Sihan, et al.
Veröffentlicht: (2026)
Improving Large Vision and Language Models by Learning from a Panel of Peers
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
von: Gao, Ziqi, et al.
Veröffentlicht: (2024) -
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
von: Huang, Weikai, et al.
Veröffentlicht: (2025) -
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
von: Zhang, Tianyi, et al.
Veröffentlicht: (2026) -
Task Me Anything
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024) -
TrajTok: Learning Trajectory Tokens enables better Video Understanding
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)