EMMA: Efficient Visual Alignment in Multi-Modal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghazanfari, Sara, Araujo, Alexandre, Krishnamurthy, Prashanth, Garg, Siddharth, Khorrami, Farshad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LipSim: A Provably Robust Perceptual Similarity Metric
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2023)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2023)
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2026)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2026)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
RAZER: Robust Accelerated Zero-Shot 3D Open-Vocabulary Panoptic Reconstruction with Spatio-Temporal Aggregation
von: Patel, Naman, et al.
Veröffentlicht: (2025)
von: Patel, Naman, et al.
Veröffentlicht: (2025)
Multi-Modal Hallucination Control by Visual Information Grounding
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
von: Favero, Alessandro, et al.
Veröffentlicht: (2024)
CLIPScope: Enhancing Zero-Shot OOD Detection with Bayesian Scoring
von: Fu, Hao, et al.
Veröffentlicht: (2024)
von: Fu, Hao, et al.
Veröffentlicht: (2024)
Efficient and Distributed Large-Scale 3D Map Registration using Tomographic Features
von: Unlu, Halil Utku, et al.
Veröffentlicht: (2024)
von: Unlu, Halil Utku, et al.
Veröffentlicht: (2024)
Text-centric Alignment for Multi-Modality Learning
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
von: Tsai, Yun-Da, et al.
Veröffentlicht: (2024)
FlashMix: Fast Map-Free LiDAR Localization via Feature Mixing and Contrastive-Constrained Accelerated Training
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
SALSA: Swift Adaptive Lightweight Self-Attention for Enhanced LiDAR Place Recognition
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
SpotEdit: Evaluating Visually-Guided Image Editing Methods
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
EMMA: End-to-End Multimodal Model for Autonomous Driving
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
von: Hwang, Jyh-Jing, et al.
Veröffentlicht: (2024)
RoboPEPP: Vision-Based Robot Pose and Joint Angle Estimation through Embedding Predictive Pre-Training
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
von: Goswami, Raktim Gautam, et al.
Veröffentlicht: (2024)
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
von: Rajabi, Navid, et al.
Veröffentlicht: (2023)
X-VILA: Cross-Modality Alignment for Large Language Model
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
von: Bhat, Vineet, et al.
Veröffentlicht: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
von: Shukor, Mustafa, et al.
Veröffentlicht: (2024)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
Enhancing Visual-Language Modality Alignment in Large Vision Language Models via Self-Improvement
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
von: Han, Songhao, et al.
Veröffentlicht: (2023)
von: Han, Songhao, et al.
Veröffentlicht: (2023)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
von: Ebrahimi, Sayna, et al.
Veröffentlicht: (2024)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
VLLaVO: Mitigating Visual Gap through LLMs
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
von: Chen, Shuhao, et al.
Veröffentlicht: (2024)
A U-Net and Transformer Pipeline for Multilingual Image Translation
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
von: Sahay, Siddharth, et al.
Veröffentlicht: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
von: Madaan, Divyam, et al.
Veröffentlicht: (2024)
von: Madaan, Divyam, et al.
Veröffentlicht: (2024)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models
von: Mishra-Sharma, Siddharth, et al.
Veröffentlicht: (2024)
von: Mishra-Sharma, Siddharth, et al.
Veröffentlicht: (2024)
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
von: Li, Shengzhi, et al.
Veröffentlicht: (2024)
von: Li, Shengzhi, et al.
Veröffentlicht: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
von: Hakim, Zaber Ibn Abdul, et al.
Veröffentlicht: (2023)
von: Hakim, Zaber Ibn Abdul, et al.
Veröffentlicht: (2023)
MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
von: Vasu, Pavan Kumar Anasosalu, et al.
Veröffentlicht: (2023)
Self-Supervised Visual Preference Alignment
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
FRAMES-VQA: Benchmarking Fine-Tuning Robustness across Multi-Modal Shifts in Visual Question Answering
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
von: Huang, Chengyue, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LipSim: A Provably Robust Perceptual Similarity Metric
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2023) -
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024) -
SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2026) -
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025) -
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)