MOAT: Evaluating LMMs for Capability Integration and Instruction Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Zhoutong, Sun, Mingze, Gao, Huan-ang, Wang, Xutong, Wang, Xiangyang, Mei, Yu, Liu, Chang, Li, Qinwei, Zhang, Chengwen, Lan, Qinghuan, Yu, Chun, Shi, Yuanchun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PACEE: Parent-Centered AI Scaffolding for Emotion Education in Early Childhood Conversations
von: Mei, Yu, et al.
Veröffentlicht: (2025)
von: Mei, Yu, et al.
Veröffentlicht: (2025)
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
von: Zhang, Chengwen, et al.
Veröffentlicht: (2026)
von: Zhang, Chengwen, et al.
Veröffentlicht: (2026)
PoseAugment: Generative Human Pose Data Augmentation with Physical Plausibility for IMU-based Motion Capture
von: Li, Zhuojun, et al.
Veröffentlicht: (2024)
von: Li, Zhuojun, et al.
Veröffentlicht: (2024)
Adapting AI to the Moment: Understanding the Dynamics of Parent-AI Collaboration Modes in Real-Time Conversations with Children
von: Mei, Yu, et al.
Veröffentlicht: (2026)
von: Mei, Yu, et al.
Veröffentlicht: (2026)
Monocular Gaussian SLAM with Language Extended Loop Closure
von: Lan, Tian, et al.
Veröffentlicht: (2024)
von: Lan, Tian, et al.
Veröffentlicht: (2024)
A Human-Computer Collaborative Tool for Training a Single Large Language Model Agent into a Network through Few Examples
von: Pan, Lihang, et al.
Veröffentlicht: (2024)
von: Pan, Lihang, et al.
Veröffentlicht: (2024)
Computing with Smart Rings: A Systematic Literature Review
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
von: Wang, Zeyu, et al.
Veröffentlicht: (2025)
CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset
von: Wang, Zhiming, et al.
Veröffentlicht: (2024)
von: Wang, Zhiming, et al.
Veröffentlicht: (2024)
SonarWatch: Field sensing technique for smartwatches based on ultrasound and motion
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
von: Li, Can, et al.
Veröffentlicht: (2025)
von: Li, Can, et al.
Veröffentlicht: (2025)
Falcon-UI: Understanding GUI Before Following User Instructions
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
von: Wang, Jiarui, et al.
Veröffentlicht: (2025)
Say Your Reason: Extract Contextual Rules In Situ for Context-aware Service Recommendation
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
CasualGaze: Towards Modeling and Recognizing Casual Gaze Behavior for Efficient Gaze-based Object Selection
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Generating Mobile Sensing Strategies in Human Behavior Modeling
von: Gao, Nan, et al.
Veröffentlicht: (2023)
von: Gao, Nan, et al.
Veröffentlicht: (2023)
Exploring the Potential of Encoder-free Architectures in 3D LMMs
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
von: Tang, Yiwen, et al.
Veröffentlicht: (2025)
MIBench: Evaluating LMMs on Multimodal Interaction
von: Miao, Yu, et al.
Veröffentlicht: (2026)
von: Miao, Yu, et al.
Veröffentlicht: (2026)
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
von: Yu, Xinquan, et al.
Veröffentlicht: (2025)
von: Yu, Xinquan, et al.
Veröffentlicht: (2025)
PhotoFramer: Multi-modal Image Composition Instruction
von: You, Zhiyuan, et al.
Veröffentlicht: (2025)
von: You, Zhiyuan, et al.
Veröffentlicht: (2025)
Probing the Reliability of Driving VLMs: From Inconsistent Responses to Grounded Temporal Reasoning
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2026)
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2026)
MMSearch-R1: Incentivizing LMMs to Search
von: Wu, Jinming, et al.
Veröffentlicht: (2025)
von: Wu, Jinming, et al.
Veröffentlicht: (2025)
SpeakSoftly: Scaffolding Nonviolent Communication in Intimate Relationships through LLM-Powered Just-In-Time Interventions
von: Chan, Ka I, et al.
Veröffentlicht: (2026)
von: Chan, Ka I, et al.
Veröffentlicht: (2026)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
von: Zeng, Bo, et al.
Veröffentlicht: (2025)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Bridging the gap between natural user expression with complex automation programming in smart homes
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
von: Shi, Yingtian, et al.
Veröffentlicht: (2024)
Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts
von: Huang, Tian, et al.
Veröffentlicht: (2024)
von: Huang, Tian, et al.
Veröffentlicht: (2024)
TaCIE: Enhancing Instruction Comprehension in Large Language Models through Task-Centred Instruction Evolution
von: Yang, Jiuding, et al.
Veröffentlicht: (2024)
von: Yang, Jiuding, et al.
Veröffentlicht: (2024)
CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
von: Shi, Jingzhe, et al.
Veröffentlicht: (2024)
TextOnly: A Unified Function Portal for Text-Related Functions on Smartphones
von: Tu, Minghao, et al.
Veröffentlicht: (2025)
von: Tu, Minghao, et al.
Veröffentlicht: (2025)
Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
von: Gou, Boyu, et al.
Veröffentlicht: (2024)
AngleSizer: Enhancing Spatial Scale Perception for the Visually Impaired with an Interactive Smartphone Assistant
von: Jing, Xiaoqing, et al.
Veröffentlicht: (2024)
von: Jing, Xiaoqing, et al.
Veröffentlicht: (2024)
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
von: Murthy, Rithesh, et al.
Veröffentlicht: (2024)
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs
von: Xu, Xinli, et al.
Veröffentlicht: (2024)
von: Xu, Xinli, et al.
Veröffentlicht: (2024)
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
Enhancing and Assessing Instruction-Following with Fine-Grained Instruction Variants
von: Yang, Jiuding, et al.
Veröffentlicht: (2024)
von: Yang, Jiuding, et al.
Veröffentlicht: (2024)
Enhancing Function-Calling Capabilities in LLMs: Strategies for Prompt Formats, Data Integration, and Multilingual Translation
von: Chen, Yi-Chang, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chang, et al.
Veröffentlicht: (2024)
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
G-VOILA: Gaze-Facilitated Information Querying in Daily Scenarios
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
von: Wang, Zeyu, et al.
Veröffentlicht: (2024)
LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhang, Kaichen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PACEE: Parent-Centered AI Scaffolding for Emotion Education in Early Childhood Conversations
von: Mei, Yu, et al.
Veröffentlicht: (2025) -
HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
von: Zhang, Chengwen, et al.
Veröffentlicht: (2026) -
PoseAugment: Generative Human Pose Data Augmentation with Physical Plausibility for IMU-based Motion Capture
von: Li, Zhuojun, et al.
Veröffentlicht: (2024) -
Adapting AI to the Moment: Understanding the Dynamics of Parent-AI Collaboration Modes in Real-Time Conversations with Children
von: Mei, Yu, et al.
Veröffentlicht: (2026) -
Monocular Gaussian SLAM with Language Extended Loop Closure
von: Lan, Tian, et al.
Veröffentlicht: (2024)