Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hua, Yilun, Artzi, Yoav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoGen: Learning from Feedback with Coupled Comprehension and Generation
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2024)
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2024)
Retrospective Learning from Interactions
von: Chen, Zizhao, et al.
Veröffentlicht: (2024)
von: Chen, Zizhao, et al.
Veröffentlicht: (2024)
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
von: Chen, Zizhao, et al.
Veröffentlicht: (2025)
Post-training for Efficient Communication via Convention Formation
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
von: Hua, Yilun, et al.
Veröffentlicht: (2025)
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
von: Yang, Jianing, et al.
Veröffentlicht: (2024)
von: Yang, Jianing, et al.
Veröffentlicht: (2024)
A Surprising Failure? Multimodal LLMs and the NLVR Challenge
von: Wu, Anne, et al.
Veröffentlicht: (2024)
von: Wu, Anne, et al.
Veröffentlicht: (2024)
Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation
von: Wang, Zixin, et al.
Veröffentlicht: (2024)
von: Wang, Zixin, et al.
Veröffentlicht: (2024)
Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
von: Yang, Yibo, et al.
Veröffentlicht: (2025)
von: Yang, Yibo, et al.
Veröffentlicht: (2025)
Generative Emotion Cause Explanation in Multimodal Conversations
von: Wang, Lin, et al.
Veröffentlicht: (2024)
von: Wang, Lin, et al.
Veröffentlicht: (2024)
Advancing Conversational Diagnostic AI with Multimodal Reasoning
von: Saab, Khaled, et al.
Veröffentlicht: (2025)
von: Saab, Khaled, et al.
Veröffentlicht: (2025)
VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
von: Song, Tingyu, et al.
Veröffentlicht: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
von: Parekh, Jayneel, et al.
Veröffentlicht: (2025)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs
von: Luo, Jinqi, et al.
Veröffentlicht: (2026)
von: Luo, Jinqi, et al.
Veröffentlicht: (2026)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
von: Dongre, Vardhan, et al.
Veröffentlicht: (2025)
von: Dongre, Vardhan, et al.
Veröffentlicht: (2025)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
OpenClaw-RL: Train Any Agent Simply by Talking
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
von: Wang, Yinjie, et al.
Veröffentlicht: (2026)
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
von: Fang, Xianya, et al.
Veröffentlicht: (2026)
MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets
von: Wei, Lai, et al.
Veröffentlicht: (2023)
von: Wei, Lai, et al.
Veröffentlicht: (2023)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
DefAn: Definitive Answer Dataset for LLMs Hallucination Evaluation
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
von: Rahman, A B M Ashikur, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
von: Chang, Kai-Po, et al.
Veröffentlicht: (2025)
von: Chang, Kai-Po, et al.
Veröffentlicht: (2025)
NVLM: Open Frontier-Class Multimodal LLMs
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2026)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Interleaving Reasoning for Better Text-to-Image Generation
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
No Mean Feat: Simple, Strong Baselines for Context Compression
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
von: Feldman, Yair, et al.
Veröffentlicht: (2025)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
von: Li, Yulin, et al.
Veröffentlicht: (2025)
von: Li, Yulin, et al.
Veröffentlicht: (2025)
Is Pre-training Truly Better Than Meta-Learning?
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
von: Miranda, Brando, et al.
Veröffentlicht: (2023)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
von: Chen, Junying, et al.
Veröffentlicht: (2025)
von: Chen, Junying, et al.
Veröffentlicht: (2025)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoGen: Learning from Feedback with Coupled Comprehension and Generation
von: Gul, Mustafa Omer, et al.
Veröffentlicht: (2024) -
Retrospective Learning from Interactions
von: Chen, Zizhao, et al.
Veröffentlicht: (2024) -
Knot So Simple: A Minimalistic Environment for Spatial Reasoning
von: Chen, Zizhao, et al.
Veröffentlicht: (2025) -
Post-training for Efficient Communication via Convention Formation
von: Hua, Yilun, et al.
Veröffentlicht: (2025) -
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
von: Yang, Jianing, et al.
Veröffentlicht: (2024)