Energy-based Automated Model Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Ru, Zou, Heming, Wang, Haobo, Zeng, Yawen, Huang, Zenan, Zhao, Junbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024)
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
von: Zou, Heming, et al.
Veröffentlicht: (2025)
von: Zou, Heming, et al.
Veröffentlicht: (2025)
Mordal: Automated Pretrained Model Selection for Vision Language Models
von: He, Shiqi, et al.
Veröffentlicht: (2025)
von: He, Shiqi, et al.
Veröffentlicht: (2025)
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
von: Cherian, Anoop, et al.
Veröffentlicht: (2024)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)
Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
von: Chen, Kun, et al.
Veröffentlicht: (2025)
von: Chen, Kun, et al.
Veröffentlicht: (2025)
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
Energy-based Models are Zero-Shot Planners for Compositional Scene Rearrangement
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
von: Gkanatsios, Nikolaos, et al.
Veröffentlicht: (2023)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
Auto MC-Reward: Automated Dense Reward Design with Large Language Models for Minecraft
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
von: Huang, Hailang, et al.
Veröffentlicht: (2024)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
HEMM: Holistic Evaluation of Multimodal Foundation Models
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
von: Liang, Paul Pu, et al.
Veröffentlicht: (2024)
Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Models
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
von: Zeng, Yu, et al.
Veröffentlicht: (2026)
Coordinated Robustness Evaluation Framework for Vision-Language Models
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
von: Liu, Dongyang, et al.
Veröffentlicht: (2024)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Improving Automatic VQA Evaluation Using Large Language Models
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
von: Mañas, Oscar, et al.
Veröffentlicht: (2023)
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
von: Eger, Steffen, et al.
Veröffentlicht: (2025)
Unified Lexical Representation for Interpretable Visual-Language Alignment
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
Unsolvable Problem Detection: Robust Understanding Evaluation for Large Multimodal Models
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2024)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
von: Wu, Junjie, et al.
Veröffentlicht: (2024)
von: Wu, Junjie, et al.
Veröffentlicht: (2024)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
von: Lu, Pan, et al.
Veröffentlicht: (2023)
von: Lu, Pan, et al.
Veröffentlicht: (2023)
ALOHa: A New Measure for Hallucination in Captioning Models
von: Petryk, Suzanne, et al.
Veröffentlicht: (2024)
von: Petryk, Suzanne, et al.
Veröffentlicht: (2024)
Automated Black-box Prompt Engineering for Personalized Text-to-Image Generation
von: He, Yutong, et al.
Veröffentlicht: (2024)
von: He, Yutong, et al.
Veröffentlicht: (2024)
SegSub: Evaluating Robustness to Knowledge Conflicts and Hallucinations in Vision-Language Models
von: Carragher, Peter, et al.
Veröffentlicht: (2025)
von: Carragher, Peter, et al.
Veröffentlicht: (2025)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
von: Srivastava, Varun, et al.
Veröffentlicht: (2025)
Grounding and Evaluation for Large Language Models: Practical Challenges and Lessons Learned (Survey)
von: Kenthapadi, Krishnaram, et al.
Veröffentlicht: (2024)
von: Kenthapadi, Krishnaram, et al.
Veröffentlicht: (2024)
Rethinking Genomic Modeling Through Optical Character Recognition
von: Xiang, Hongxin, et al.
Veröffentlicht: (2026)
von: Xiang, Hongxin, et al.
Veröffentlicht: (2026)
VisPlay: Self-Evolving Vision-Language Models from Images
von: He, Yicheng, et al.
Veröffentlicht: (2025)
von: He, Yicheng, et al.
Veröffentlicht: (2025)
Erasing with Precision: Evaluating Specific Concept Erasure from Text-to-Image Generative Models
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
von: Fuchi, Masane, et al.
Veröffentlicht: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
von: Kalibhat, Neha, et al.
Veröffentlicht: (2026)
Energy-Based Transformers are Scalable Learners and Thinkers
von: Gladstone, Alexi, et al.
Veröffentlicht: (2025)
von: Gladstone, Alexi, et al.
Veröffentlicht: (2025)
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2025)
OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
von: Hu, Xueyu, et al.
Veröffentlicht: (2025)
SLaVA-CXR: Small Language and Vision Assistant for Chest X-ray Report Automation
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
von: Wu, Jinge, et al.
Veröffentlicht: (2024)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
von: Zheng, Naishan, et al.
Veröffentlicht: (2025)
Efficient Pruning of Text-to-Image Models: Insights from Pruning Stable Diffusion
von: Ramesh, Samarth N, et al.
Veröffentlicht: (2024)
von: Ramesh, Samarth N, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
von: Mazeika, Mantas, et al.
Veröffentlicht: (2024) -
Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
von: Zou, Heming, et al.
Veröffentlicht: (2025) -
Mordal: Automated Pretrained Model Selection for Vision Language Models
von: He, Shiqi, et al.
Veröffentlicht: (2025) -
Evaluating Large Vision-and-Language Models on Children's Mathematical Olympiads
von: Cherian, Anoop, et al.
Veröffentlicht: (2024) -
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
von: Zhang, Yuhui, et al.
Veröffentlicht: (2025)