VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Duan, Haodong, Fang, Xinyu, Yang, Junming, Zhao, Xiangyu, Qiao, Yuxuan, Li, Mo, Agarwal, Amit, Chen, Zhe, Chen, Lin, Liu, Yuan, Ma, Yubo, Sun, Hailong, Zhang, Yifan, Lu, Shiyin, Wong, Tack Hwa, Wang, Weiyun, Zhou, Peiheng, Li, Xiaozhe, Fu, Chaoyou, Cui, Junbo, Chen, Jixuan, Song, Enxin, Mao, Song, Ding, Shengyuan, Liang, Tianhao, Zhang, Zicheng, Dong, Xiaoyi, Zang, Yuhang, Zhang, Pan, Wang, Jiaqi, Lin, Dahua, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
di: Jia, Qi, et al.
Pubblicazione: (2026)
di: Jia, Qi, et al.
Pubblicazione: (2026)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024)
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
di: Wang, Chonghua, et al.
Pubblicazione: (2024)
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
di: Fang, Xinyu, et al.
Pubblicazione: (2025)
di: Fang, Xinyu, et al.
Pubblicazione: (2025)
Extracellular vesicles—Potential link between periodontal disease and diabetic complications
di: Shengyuan Huang, et al.
Pubblicazione: (2024)
di: Shengyuan Huang, et al.
Pubblicazione: (2024)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
di: Zhuo, Jingming, et al.
Pubblicazione: (2024)
NP-Engine: Empowering Optimization Reasoning in Large Language Models with Verifiable Synthetic NP Problems
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
di: Li, Xiaozhe, et al.
Pubblicazione: (2025)
MM-HELIX: Boosting Multimodal Long-Chain Reflective Reasoning with Holistic Platform and Adaptive Hybrid Policy Optimization
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
SciEvalKit: An Open-source Evaluation Toolkit for Scientific General Intelligence
di: Wang, Yiheng, et al.
Pubblicazione: (2025)
di: Wang, Yiheng, et al.
Pubblicazione: (2025)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2025)
A Rare Case of Disseminated Extrapulmonary Tuberculosis Diagnosed by Endoscopic Ultrasonography‐Guided Fine‐Needle Biopsy: A Case Report
di: Yi‐Lin Lin, et al.
Pubblicazione: (2026)
di: Yi‐Lin Lin, et al.
Pubblicazione: (2026)
Abductive Inference in Retrieval-Augmented Language Models: Generating and Validating Missing Premises
di: Lin, Shiyin
Pubblicazione: (2025)
di: Lin, Shiyin
Pubblicazione: (2025)
LLM-Driven Adaptive Source-Sink Identification and False Positive Mitigation for Static Analysis
di: Lin, Shiyin
Pubblicazione: (2025)
di: Lin, Shiyin
Pubblicazione: (2025)
Hybrid Fuzzing with LLM-Guided Input Mutation and Semantic Feedback
di: Lin, Shiyin
Pubblicazione: (2025)
di: Lin, Shiyin
Pubblicazione: (2025)
Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
di: Li, Xiaozhe, et al.
Pubblicazione: (2026)
Sea–Land Segmentation Dataset Sources
di: Zhang, Jixuan
Pubblicazione: (2025)
di: Zhang, Jixuan
Pubblicazione: (2025)
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model
di: Zang, Yuhang, et al.
Pubblicazione: (2025)
di: Zang, Yuhang, et al.
Pubblicazione: (2025)
Are We on the Right Way for Evaluating Large Vision-Language Models?
di: Chen, Lin, et al.
Pubblicazione: (2024)
di: Chen, Lin, et al.
Pubblicazione: (2024)
SPARK: Synergistic Policy And Reward Co-Evolving Framework
di: Liu, Ziyu, et al.
Pubblicazione: (2025)
di: Liu, Ziyu, et al.
Pubblicazione: (2025)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2024)
di: Zhang, Yi-Kai, et al.
Pubblicazione: (2024)
Decentralized model reference adaptive control for interconnected systems with time‐varying delays and unknown dead‐zone inputs
di: Chen Yang, et al.
Pubblicazione: (2024)
di: Chen Yang, et al.
Pubblicazione: (2024)
Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
di: Ma, Yichuan, et al.
Pubblicazione: (2026)
From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language Models
di: Li, Rongjie, et al.
Pubblicazione: (2024)
di: Li, Rongjie, et al.
Pubblicazione: (2024)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
di: Zhang, Mengchen, et al.
Pubblicazione: (2025)
Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
di: Gu, Yuzhe, et al.
Pubblicazione: (2025)
MM-IFEngine: Towards Multimodal Instruction Following
di: Ding, Shengyuan, et al.
Pubblicazione: (2025)
di: Ding, Shengyuan, et al.
Pubblicazione: (2025)
From Pets to Robots: MojiKit as a Data-Informed Toolkit for Affective HRI Design
di: He, Liwen, et al.
Pubblicazione: (2026)
di: He, Liwen, et al.
Pubblicazione: (2026)
NepTrain and NepTrainKit: Automated Active Learning and Visualization Toolkit for Neuroevolution Potentials
di: Chen, Chengbing, et al.
Pubblicazione: (2025)
di: Chen, Chengbing, et al.
Pubblicazione: (2025)
High‐Throughput Sorting and Single‐Cell Mechanotyping by Hydrodynamic Sorting‐Mechanotyping Cytometry
di: Yao Chen, et al.
Pubblicazione: (2024)
di: Yao Chen, et al.
Pubblicazione: (2024)
Trust Your Critic: Robust Reward Modeling and Reinforcement Learning for Faithful Image Editing and Generation
di: Zhao, Xiangyu, et al.
Pubblicazione: (2026)
di: Zhao, Xiangyu, et al.
Pubblicazione: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance
di: Chen, Wenhao, et al.
Pubblicazione: (2026)
di: Chen, Wenhao, et al.
Pubblicazione: (2026)
Information Density Principle for MLLM Benchmarks
di: Li, Chunyi, et al.
Pubblicazione: (2025)
di: Li, Chunyi, et al.
Pubblicazione: (2025)
On the extreme points of sets of absolulely separable and PPT states
di: Song, Zhiwei, et al.
Pubblicazione: (2024)
di: Song, Zhiwei, et al.
Pubblicazione: (2024)
Spectral characterizations of entanglement witnesses
di: Song, Zhiwei, et al.
Pubblicazione: (2025)
di: Song, Zhiwei, et al.
Pubblicazione: (2025)
A counterexample to the strong spin alignment conjecture
di: Song, Zhiwei, et al.
Pubblicazione: (2026)
di: Song, Zhiwei, et al.
Pubblicazione: (2026)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
di: Li, Xiaozhe, et al.
Pubblicazione: (2025) -
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
di: Li, Xiaozhe, et al.
Pubblicazione: (2026) -
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
di: Jia, Qi, et al.
Pubblicazione: (2026) -
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
di: Qiao, Yuxuan, et al.
Pubblicazione: (2024) -
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
di: Wang, Chonghua, et al.
Pubblicazione: (2024)