Meeseeks: A Feedback-Driven, Iterative Self-Correction Benchmark evaluating LLMs' Instruction Following Capability
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | wang, Jiaming, Zhao, Yunke, Ding, Peng, Kuang, Jun, Shen, Yibin, Tang, Zhe, Jin, Yilin, Wang, ZongYu, Li, Xiaoyu, Cao, Xuezhi, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
par: Zhu, Yaoming, et autres
Publié: (2025)
par: Zhu, Yaoming, et autres
Publié: (2025)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
par: Nie, Dujun, et autres
Publié: (2026)
par: Nie, Dujun, et autres
Publié: (2026)
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
par: Ding, Peng, et autres
Publié: (2025)
par: Ding, Peng, et autres
Publié: (2025)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
par: Ding, Peng, et autres
Publié: (2025)
par: Ding, Peng, et autres
Publié: (2025)
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
par: Wang, Jiaming, et autres
Publié: (2025)
par: Wang, Jiaming, et autres
Publié: (2025)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
par: Fu, Lingyue, et autres
Publié: (2025)
par: Fu, Lingyue, et autres
Publié: (2025)
KG-o1: Enhancing Multi-hop Question Answering in Large Language Models via Knowledge Graph Integration
par: Wang, Nan, et autres
Publié: (2025)
par: Wang, Nan, et autres
Publié: (2025)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
par: Jiayang, Cheng, et autres
Publié: (2026)
par: Jiayang, Cheng, et autres
Publié: (2026)
FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
par: Xia, Congying, et autres
Publié: (2024)
par: Xia, Congying, et autres
Publié: (2024)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
par: Chen, Chen, et autres
Publié: (2025)
par: Chen, Chen, et autres
Publié: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
par: Ding, Peng, et autres
Publié: (2024)
par: Ding, Peng, et autres
Publié: (2024)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
par: Yan, Kaiwen, et autres
Publié: (2025)
par: Yan, Kaiwen, et autres
Publié: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
par: Wen, Bosi, et autres
Publié: (2026)
par: Wen, Bosi, et autres
Publié: (2026)
CodeContests-O: Powering LLMs via Feedback-Driven Iterative Test Case Generation
par: Cai, Jianfeng, et autres
Publié: (2026)
par: Cai, Jianfeng, et autres
Publié: (2026)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
par: He, Yun, et autres
Publié: (2024)
par: He, Yun, et autres
Publié: (2024)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
par: Wang, Ao, et autres
Publié: (2024)
par: Wang, Ao, et autres
Publié: (2024)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
par: Gao, Yiming, et autres
Publié: (2025)
par: Gao, Yiming, et autres
Publié: (2025)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
par: Kolasani, Sai, et autres
Publié: (2025)
par: Kolasani, Sai, et autres
Publié: (2025)
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
par: Qian, Yusu, et autres
Publié: (2024)
par: Qian, Yusu, et autres
Publié: (2024)
DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
par: Che, Lirong, et autres
Publié: (2026)
par: Che, Lirong, et autres
Publié: (2026)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
par: Yang, Zhe, et autres
Publié: (2024)
par: Yang, Zhe, et autres
Publié: (2024)
LERa: Replanning with Visual Feedback in Instruction Following
par: Pchelintsev, Svyatoslav, et autres
Publié: (2025)
par: Pchelintsev, Svyatoslav, et autres
Publié: (2025)
A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback
par: Duan, Guoliang, et autres
Publié: (2025)
par: Duan, Guoliang, et autres
Publié: (2025)
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
par: Ji, Jiaming, et autres
Publié: (2024)
par: Ji, Jiaming, et autres
Publié: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
par: Wang, Xihuai, et autres
Publié: (2025)
par: Wang, Xihuai, et autres
Publié: (2025)
When Instructions Multiply: Measuring and Estimating LLM Capabilities of Multiple Instructions Following
par: Harada, Keno, et autres
Publié: (2025)
par: Harada, Keno, et autres
Publié: (2025)
The Instruction Gap: LLMs get lost in Following Instruction
par: Tripathi, Vishesh, et autres
Publié: (2025)
par: Tripathi, Vishesh, et autres
Publié: (2025)
Unraveling the Mystery of Scaling Laws: Part I
par: Su, Hui, et autres
Publié: (2024)
par: Su, Hui, et autres
Publié: (2024)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
par: Liu, Junlin, et autres
Publié: (2026)
par: Liu, Junlin, et autres
Publié: (2026)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
par: Sun, Chung-En, et autres
Publié: (2024)
par: Sun, Chung-En, et autres
Publié: (2024)
MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
par: Yan, Siyu, et autres
Publié: (2025)
par: Yan, Siyu, et autres
Publié: (2025)
HalluScan: A Systematic Benchmark for Detecting and Mitigating Hallucinations in Instruction-Following LLMs
par: Cherif, Ahmed
Publié: (2026)
par: Cherif, Ahmed
Publié: (2026)
Self-Review Framework for Enhancing Instruction Following Capability of LLM
par: Park, Sihyun
Publié: (2025)
par: Park, Sihyun
Publié: (2025)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
par: Zhang, Junyan, et autres
Publié: (2025)
par: Zhang, Junyan, et autres
Publié: (2025)
CoBA-RL: Capability-Oriented Budget Allocation for Reinforcement Learning in LLMs
par: Yao, Zhiyuan, et autres
Publié: (2026)
par: Yao, Zhiyuan, et autres
Publié: (2026)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
par: Xu, Wanghan, et autres
Publié: (2025)
par: Xu, Wanghan, et autres
Publié: (2025)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
par: Jin, Rihui, et autres
Publié: (2026)
par: Jin, Rihui, et autres
Publié: (2026)
Iterative methods of linearized moment equations for rarefied gases
par: Dong, Xiaoyu, et autres
Publié: (2023)
par: Dong, Xiaoyu, et autres
Publié: (2023)
sprofing the DLSCA
par: wang
Publié: (2025)
par: wang
Publié: (2025)
Documents similaires
-
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
par: Zhu, Yaoming, et autres
Publié: (2025) -
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
par: Nie, Dujun, et autres
Publié: (2026) -
Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement
par: Ding, Peng, et autres
Publié: (2025) -
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
par: Fu, Lingyue, et autres
Publié: (2025) -
Friend or Foe: How LLMs' Safety Mind Gets Fooled by Intent Shift Attack
par: Ding, Peng, et autres
Publié: (2025)