Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, James Y., Zhang, Sheng, Liu, Qianchu, Qin, Guanghui, Zhu, Tinghui, Naumann, Tristan, Chen, Muhao, Poon, Hoifung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
by: Ossowski, Timothy, et al.
Published: (2025)
by: Ossowski, Timothy, et al.
Published: (2025)
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
by: Liu, Qianchu, et al.
Published: (2025)
by: Liu, Qianchu, et al.
Published: (2025)
Video Models Can Reason with Verifiable Rewards
by: Zhu, Tinghui, et al.
Published: (2026)
by: Zhu, Tinghui, et al.
Published: (2026)
Is Extending Modality The Right Path Towards Omni-Modality?
by: Zhu, Tinghui, et al.
Published: (2025)
by: Zhu, Tinghui, et al.
Published: (2025)
Offset Unlearning for Large Language Models
by: Huang, James Y., et al.
Published: (2024)
by: Huang, James Y., et al.
Published: (2024)
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
by: Zhou, Wenxuan, et al.
Published: (2023)
by: Zhou, Wenxuan, et al.
Published: (2023)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
by: Zhu, Tinghui, et al.
Published: (2024)
by: Zhu, Tinghui, et al.
Published: (2024)
Exploring Scaling Laws for EHR Foundation Models
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
Scaling medical imaging report generation with multimodal reinforcement learning
by: Liu, Qianchu, et al.
Published: (2026)
by: Liu, Qianchu, et al.
Published: (2026)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
by: Xu, Nan, et al.
Published: (2024)
by: Xu, Nan, et al.
Published: (2024)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
by: Huang, James Y., et al.
Published: (2025)
by: Huang, James Y., et al.
Published: (2025)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
by: Xie, Yiqing, et al.
Published: (2023)
by: Xie, Yiqing, et al.
Published: (2023)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
by: Gero, Zelalem, et al.
Published: (2024)
by: Gero, Zelalem, et al.
Published: (2024)
Pareto Optimal Learning for Estimating Large Language Model Errors
by: Zhao, Theodore, et al.
Published: (2023)
by: Zhao, Theodore, et al.
Published: (2023)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
by: Zhu, Boyu, et al.
Published: (2025)
by: Zhu, Boyu, et al.
Published: (2025)
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
by: Li, Yuankai, et al.
Published: (2026)
by: Li, Yuankai, et al.
Published: (2026)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
by: Zhao, Theodore Zhengde, et al.
Published: (2026)
by: Zhao, Theodore Zhengde, et al.
Published: (2026)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
by: Cai, Rui, et al.
Published: (2025)
by: Cai, Rui, et al.
Published: (2025)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Through My Eyes: Integrative Model of Awareness‐Raising (IMAR)
by: Abdelmajid Kadri, et al.
Published: (2026)
by: Abdelmajid Kadri, et al.
Published: (2026)
Cautionary Tales on Synthetic Controls in Survival Analyses
by: Curth, Alicia, et al.
Published: (2023)
by: Curth, Alicia, et al.
Published: (2023)
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
by: Ossowski, Timothy, et al.
Published: (2026)
by: Ossowski, Timothy, et al.
Published: (2026)
Solar Photovoltaic Assessment with Large Language Model
by: Guo, Muhao, et al.
Published: (2025)
by: Guo, Muhao, et al.
Published: (2025)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
TRIALSCOPE: A Unifying Causal Framework for Scaling Real-World Evidence Generation with Biomedical Language Models
by: González, Javier, et al.
Published: (2023)
by: González, Javier, et al.
Published: (2023)
T-Rex: Text-assisted Retrosynthesis Prediction
by: Liu, Yifeng, et al.
Published: (2024)
by: Liu, Yifeng, et al.
Published: (2024)
Scaling Large Language Model-based Multi-Agent Collaboration
by: Qian, Chen, et al.
Published: (2024)
by: Qian, Chen, et al.
Published: (2024)
Beyond Single-Modal Analytics: A Framework for Integrating Heterogeneous LLM-Based Query Systems for Multi-Modal Data
by: Li, Ruyu, et al.
Published: (2026)
by: Li, Ruyu, et al.
Published: (2026)
Reading With My Eyes Open
by: Quist, Gerdi
Published: (2015)
by: Quist, Gerdi
Published: (2015)
Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
by: Chen, Zhawnen, et al.
Published: (2024)
by: Chen, Zhawnen, et al.
Published: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models
by: Li, Yuanbo, et al.
Published: (2026)
by: Li, Yuanbo, et al.
Published: (2026)
MyGram: Modality-aware Graph Transformer with Global Distribution for Multi-modal Entity Alignment
by: Li, Zhifei, et al.
Published: (2026)
by: Li, Zhifei, et al.
Published: (2026)
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
by: Graf, Victoria, et al.
Published: (2024)
by: Graf, Victoria, et al.
Published: (2024)
Similar Items
-
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
by: Zhang, Sheng, et al.
Published: (2025) -
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
by: Ossowski, Timothy, et al.
Published: (2025) -
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains
by: Liu, Qianchu, et al.
Published: (2025) -
Video Models Can Reason with Verifiable Rewards
by: Zhu, Tinghui, et al.
Published: (2026) -
Is Extending Modality The Right Path Towards Omni-Modality?
by: Zhu, Tinghui, et al.
Published: (2025)