Evaluating Multimodal Large Language Models with Daily Composite Tasks in Home Environments
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Zhenliang, Wang, Yuxi, Xie, Hongzhao, Zhao, Shiyun, Liu, Mingyuan, Lu, Yujie, He, Xinyi, Cheng, Zhenku, Peng, Yujia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents
di: He, Xinyi, et al.
Pubblicazione: (2026)
di: He, Xinyi, et al.
Pubblicazione: (2026)
Learning Long Short-Term Intention within Human Daily Behaviors
di: Sun, Zhe, et al.
Pubblicazione: (2025)
di: Sun, Zhe, et al.
Pubblicazione: (2025)
TongSIM: A General Platform for Simulating Intelligent Machines
di: Sun, Zhe, et al.
Pubblicazione: (2025)
di: Sun, Zhe, et al.
Pubblicazione: (2025)
On Domain-Adaptive Post-Training for Multimodal Large Language Models
di: Cheng, Daixuan, et al.
Pubblicazione: (2024)
di: Cheng, Daixuan, et al.
Pubblicazione: (2024)
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
di: Yin, Yueqin, et al.
Pubblicazione: (2024)
di: Yin, Yueqin, et al.
Pubblicazione: (2024)
An Empirical Analysis on Large Language Models in Debate Evaluation
di: Liu, Xinyi, et al.
Pubblicazione: (2024)
di: Liu, Xinyi, et al.
Pubblicazione: (2024)
Model Composition for Multimodal Large Language Models
di: Chen, Chi, et al.
Pubblicazione: (2024)
di: Chen, Chi, et al.
Pubblicazione: (2024)
SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models
di: Zhao, Xinyi, et al.
Pubblicazione: (2025)
di: Zhao, Xinyi, et al.
Pubblicazione: (2025)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
di: Xu, Xiao, et al.
Pubblicazione: (2024)
di: Xu, Xiao, et al.
Pubblicazione: (2024)
Multimodal Health Risk Prediction System for Chronic Diseases via Vision-Language Fusion and Large Language Models
di: Lu, Dingxin, et al.
Pubblicazione: (2025)
di: Lu, Dingxin, et al.
Pubblicazione: (2025)
MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
di: Liang, Qian, et al.
Pubblicazione: (2025)
di: Liang, Qian, et al.
Pubblicazione: (2025)
XeMap: Contextual Referring in Large-Scale Remote Sensing Environments
di: Li, Yuxi, et al.
Pubblicazione: (2025)
di: Li, Yuxi, et al.
Pubblicazione: (2025)
REALM: RAG-Driven Enhancement of Multimodal Electronic Health Records Analysis via Large Language Models
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
di: Zhu, Yinghao, et al.
Pubblicazione: (2024)
S$^3$IT: A Benchmark for Spatially Situated Social Intelligence Test
di: Sun, Zhe, et al.
Pubblicazione: (2025)
di: Sun, Zhe, et al.
Pubblicazione: (2025)
Evaluating Multimodal Large Language Models on Core Music Perception Tasks
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
di: Carone, Brandon James, et al.
Pubblicazione: (2025)
InstructCoder: Instruction Tuning Large Language Models for Code Editing
di: Li, Kaixin, et al.
Pubblicazione: (2023)
di: Li, Kaixin, et al.
Pubblicazione: (2023)
CompCap: Improving Multimodal Large Language Models with Composite Captions
di: Chen, Xiaohui, et al.
Pubblicazione: (2024)
di: Chen, Xiaohui, et al.
Pubblicazione: (2024)
TangramPuzzle: Evaluating Multimodal Large Language Models with Compositional Spatial Reasoning
di: Liu, Daixian, et al.
Pubblicazione: (2026)
di: Liu, Daixian, et al.
Pubblicazione: (2026)
BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model
di: Li, Yujie, et al.
Pubblicazione: (2025)
di: Li, Yujie, et al.
Pubblicazione: (2025)
Diffusion-RPO: Aligning Diffusion Models through Relative Preference Optimization
di: Gu, Yi, et al.
Pubblicazione: (2024)
di: Gu, Yi, et al.
Pubblicazione: (2024)
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
di: Cheng, Xianfu, et al.
Pubblicazione: (2025)
di: Cheng, Xianfu, et al.
Pubblicazione: (2025)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
di: Yang, Xiaocui, et al.
Pubblicazione: (2024)
di: Yang, Xiaocui, et al.
Pubblicazione: (2024)
Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models
di: Tan, Rui Yang, et al.
Pubblicazione: (2026)
di: Tan, Rui Yang, et al.
Pubblicazione: (2026)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
di: Yan, Ziang, et al.
Pubblicazione: (2024)
di: Yan, Ziang, et al.
Pubblicazione: (2024)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
di: Liu, Hongfu, et al.
Pubblicazione: (2024)
di: Liu, Hongfu, et al.
Pubblicazione: (2024)
LLMs-guided adaptive compensator: Bringing Adaptivity to Automatic Control Systems with Large Language Models
di: Zhou, Zhongchao, et al.
Pubblicazione: (2025)
di: Zhou, Zhongchao, et al.
Pubblicazione: (2025)
Optimized Deployment of HAPS Systems for GNSS Localization Enhancement in Urban Environments
di: Zheng, Hongzhao, et al.
Pubblicazione: (2026)
di: Zheng, Hongzhao, et al.
Pubblicazione: (2026)
Performance Evaluation of Large Language Models in Statistical Programming
di: Song, Xinyi, et al.
Pubblicazione: (2025)
di: Song, Xinyi, et al.
Pubblicazione: (2025)
CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models
di: Wang, Song, et al.
Pubblicazione: (2024)
di: Wang, Song, et al.
Pubblicazione: (2024)
The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
di: Leng, Sicong, et al.
Pubblicazione: (2024)
di: Leng, Sicong, et al.
Pubblicazione: (2024)
Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks
di: Hong, Jindong, et al.
Pubblicazione: (2025)
di: Hong, Jindong, et al.
Pubblicazione: (2025)
Experimental study on the effect of dry water materials on the fire extinguishing efficiency and suppression mechanism of wood crib fire
di: Guoqiang Chai, et al.
Pubblicazione: (2024)
di: Guoqiang Chai, et al.
Pubblicazione: (2024)
Text as Images: Can Multimodal Large Language Models Follow Printed Instructions in Pixels?
di: Li, Xiujun, et al.
Pubblicazione: (2023)
di: Li, Xiujun, et al.
Pubblicazione: (2023)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
di: Cheng, Zebang, et al.
Pubblicazione: (2024)
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
Language-Instructed Reasoning for Group Activity Detection via Multimodal Large Language Model
di: Peng, Jihua, et al.
Pubblicazione: (2025)
di: Peng, Jihua, et al.
Pubblicazione: (2025)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
Evaluating the Quality of Randomness and Entropy in Tasks Supported by Large Language Models
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
di: Karanjai, Rabimba, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Automatic Cognitive Task Generation for In-Situ Evaluation of Embodied Agents
di: He, Xinyi, et al.
Pubblicazione: (2026) -
Learning Long Short-Term Intention within Human Daily Behaviors
di: Sun, Zhe, et al.
Pubblicazione: (2025) -
TongSIM: A General Platform for Simulating Intelligent Machines
di: Sun, Zhe, et al.
Pubblicazione: (2025) -
On Domain-Adaptive Post-Training for Multimodal Large Language Models
di: Cheng, Daixuan, et al.
Pubblicazione: (2024) -
Self-Augmented Preference Optimization: Off-Policy Paradigms for Language Model Alignment
di: Yin, Yueqin, et al.
Pubblicazione: (2024)