PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Liang, Zhang, Yichi, Ren, Shuhuai, Zhao, Haozhe, Cai, Zefan, Wang, Yuchi, Wang, Peiyi, Meng, Xiangdi, Liu, Tianyu, Chang, Baobao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)
di: An, Kaikai, et al.
Pubblicazione: (2024)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice
di: Wang, Rongyang, et al.
Pubblicazione: (2026)
di: Wang, Rongyang, et al.
Pubblicazione: (2026)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web
di: Guo, Hongcheng, et al.
Pubblicazione: (2024)
di: Guo, Hongcheng, et al.
Pubblicazione: (2024)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
di: Ren, Shuhuai, et al.
Pubblicazione: (2023)
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
di: Wang, Yuchi, et al.
Pubblicazione: (2025)
LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
di: Wang, Yuchi, et al.
Pubblicazione: (2024)
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
THINK-Bench: Evaluating Thinking Efficiency and Chain-of-Thought Quality of Large Reasoning Models
di: Li, Zhiyuan, et al.
Pubblicazione: (2025)
di: Li, Zhiyuan, et al.
Pubblicazione: (2025)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
di: Yao, Linli, et al.
Pubblicazione: (2024)
di: Yao, Linli, et al.
Pubblicazione: (2024)
Chain-of-Thought Tokens are Computer Program Variables
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
MER-Bench: A Comprehensive Benchmark for Multimodal Meme Reappraisal
di: Nie, Yiqi, et al.
Pubblicazione: (2026)
di: Nie, Yiqi, et al.
Pubblicazione: (2026)
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis
di: Tu, Zeao, et al.
Pubblicazione: (2024)
di: Tu, Zeao, et al.
Pubblicazione: (2024)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
di: Zhang, Mian, et al.
Pubblicazione: (2024)
di: Zhang, Mian, et al.
Pubblicazione: (2024)
Towards a Unified View of Preference Learning for Large Language Models: A Survey
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
di: Xu, Mengya, et al.
Pubblicazione: (2025)
di: Xu, Mengya, et al.
Pubblicazione: (2025)
Improving Event Definition Following For Zero-Shot Event Detection
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
di: Ying, Kaining, et al.
Pubblicazione: (2024)
di: Ying, Kaining, et al.
Pubblicazione: (2024)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
di: Huang, Yipo, et al.
Pubblicazione: (2024)
di: Huang, Yipo, et al.
Pubblicazione: (2024)
RISEE: A Highly Interactive Naturalistic Driving Trajectories Dataset with Human Subjective Risk Perception and Eye-tracking Information
di: Wu, Xinzheng, et al.
Pubblicazione: (2025)
di: Wu, Xinzheng, et al.
Pubblicazione: (2025)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
di: Li, Lei, et al.
Pubblicazione: (2024)
di: Li, Lei, et al.
Pubblicazione: (2024)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
di: Ma, David, et al.
Pubblicazione: (2025)
di: Ma, David, et al.
Pubblicazione: (2025)
MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
di: Jin, Bohan, et al.
Pubblicazione: (2025)
di: Jin, Bohan, et al.
Pubblicazione: (2025)
VisNumBench: Evaluating Number Sense of Multimodal Large Language Models
di: Weng, Tengjin, et al.
Pubblicazione: (2025)
di: Weng, Tengjin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
di: Chen, Liang, et al.
Pubblicazione: (2024) -
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024) -
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
di: Zhao, Haozhe, et al.
Pubblicazione: (2024) -
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024) -
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
di: Si, Shuzheng, et al.
Pubblicazione: (2023)