AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Qi, Yunjia, Peng, Hao, Wang, Xiaozhi, Xin, Amy, Liu, Youfeng, Xu, Bin, Hou, Lei, Li, Juanzi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Constraint Back-translation Improves Complex Instruction Following of Large Language Models
par: Qi, Yunjia, et autres
Publié: (2024)
par: Qi, Yunjia, et autres
Publié: (2024)
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
par: Peng, Hao, et autres
Publié: (2025)
par: Peng, Hao, et autres
Publié: (2025)
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
par: Peng, Hao, et autres
Publié: (2025)
par: Peng, Hao, et autres
Publié: (2025)
On the Paradoxical Interference between Instruction-Following and Task Solving
par: Qi, Yunjia, et autres
Publié: (2026)
par: Qi, Yunjia, et autres
Publié: (2026)
ADELIE: Aligning Large Language Models on Information Extraction
par: Qi, Yunjia, et autres
Publié: (2024)
par: Qi, Yunjia, et autres
Publié: (2024)
StoryAlign: Evaluating and Training Reward Models for Story Generation
par: Xia, Haotian, et autres
Publié: (2026)
par: Xia, Haotian, et autres
Publié: (2026)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
par: Peng, Hao, et autres
Publié: (2026)
par: Peng, Hao, et autres
Publié: (2026)
StoryWriter: A Multi-Agent Framework for Long Story Generation
par: Xia, Haotian, et autres
Publié: (2025)
par: Xia, Haotian, et autres
Publié: (2025)
MAVEN-Fact: A Large-scale Event Factuality Detection Dataset
par: Li, Chunyang, et autres
Publié: (2024)
par: Li, Chunyang, et autres
Publié: (2024)
LLMAEL: Large Language Models are Good Context Augmenters for Entity Linking
par: Xin, Amy, et autres
Publié: (2024)
par: Xin, Amy, et autres
Publié: (2024)
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
par: Ma, Shengkun, et autres
Publié: (2025)
par: Ma, Shengkun, et autres
Publié: (2025)
Pre-training Distillation for Large Language Models: A Design Space Exploration
par: Peng, Hao, et autres
Publié: (2024)
par: Peng, Hao, et autres
Publié: (2024)
R-Eval: A Unified Toolkit for Evaluating Domain Knowledge of Retrieval Augmented Large Language Models
par: Tu, Shangqing, et autres
Publié: (2024)
par: Tu, Shangqing, et autres
Publié: (2024)
Boundary-Guided Policy Optimization for Memory-efficient RL of Diffusion Large Language Models
par: Lin, Nianyi, et autres
Publié: (2025)
par: Lin, Nianyi, et autres
Publié: (2025)
Event-level Knowledge Editing
par: Peng, Hao, et autres
Publié: (2024)
par: Peng, Hao, et autres
Publié: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
par: LI, Yizhi, et autres
Publié: (2024)
par: LI, Yizhi, et autres
Publié: (2024)
ChatLog: Carefully Evaluating the Evolution of ChatGPT Across Time
par: Tu, Shangqing, et autres
Publié: (2023)
par: Tu, Shangqing, et autres
Publié: (2023)
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
par: Ding, Deming, et autres
Publié: (2026)
par: Ding, Deming, et autres
Publié: (2026)
OpenEP: Open-Ended Future Event Prediction
par: Guan, Yong, et autres
Publié: (2024)
par: Guan, Yong, et autres
Publié: (2024)
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models
par: Tu, Shangqing, et autres
Publié: (2023)
par: Tu, Shangqing, et autres
Publié: (2023)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
par: Sun, Wangtao, et autres
Publié: (2024)
par: Sun, Wangtao, et autres
Publié: (2024)
Towards Understanding Safety Alignment: A Mechanistic Perspective from Safety Neurons
par: Chen, Jianhui, et autres
Publié: (2024)
par: Chen, Jianhui, et autres
Publié: (2024)
ABC-Eval: Benchmarking Large Language Models on Symbolic Music Understanding and Instruction Following
par: Zhao, Jiahao, et autres
Publié: (2025)
par: Zhao, Jiahao, et autres
Publié: (2025)
Preserving Knowledge Invariance: Rethinking Robustness Evaluation of Open Information Extraction
par: Qi, Ji, et autres
Publié: (2023)
par: Qi, Ji, et autres
Publié: (2023)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
par: Jing, Yi, et autres
Publié: (2026)
par: Jing, Yi, et autres
Publié: (2026)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
par: Kim, Dongjun, et autres
Publié: (2025)
par: Kim, Dongjun, et autres
Publié: (2025)
TacoERE: Cluster-aware Compression for Event Relation Extraction
par: Guan, Yong, et autres
Publié: (2024)
par: Guan, Yong, et autres
Publié: (2024)
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
par: Bai, Yushi, et autres
Publié: (2024)
par: Bai, Yushi, et autres
Publié: (2024)
KoLA: Carefully Benchmarking World Knowledge of Large Language Models
par: Yu, Jifan, et autres
Publié: (2023)
par: Yu, Jifan, et autres
Publié: (2023)
LIFEBench: Evaluating Length Instruction Following in Large Language Models
par: Zhang, Wei, et autres
Publié: (2025)
par: Zhang, Wei, et autres
Publié: (2025)
LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder
par: Jing, Yi, et autres
Publié: (2025)
par: Jing, Yi, et autres
Publié: (2025)
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning
par: Xin, Amy, et autres
Publié: (2024)
par: Xin, Amy, et autres
Publié: (2024)
Benchmarking Complex Instruction-Following with Multiple Constraints Composition
par: Wen, Bosi, et autres
Publié: (2024)
par: Wen, Bosi, et autres
Publié: (2024)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
par: Qin, Yiwei, et autres
Publié: (2024)
par: Qin, Yiwei, et autres
Publié: (2024)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
par: Yan, Jianhao, et autres
Publié: (2024)
par: Yan, Jianhao, et autres
Publié: (2024)
Countering Catastrophic Forgetting of Large Language Models for Better Instruction Following via Weight-Space Model Merging
par: Lyu, Mengxian, et autres
Publié: (2026)
par: Lyu, Mengxian, et autres
Publié: (2026)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
par: Si, Shuzheng, et autres
Publié: (2025)
par: Si, Shuzheng, et autres
Publié: (2025)
Can Language Models Follow Multiple Turns of Entangled Instructions?
par: Han, Chi, et autres
Publié: (2025)
par: Han, Chi, et autres
Publié: (2025)
MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
par: Ye, Junjie, et autres
Publié: (2025)
par: Ye, Junjie, et autres
Publié: (2025)
Instruction Following by Principled Boosting Attention of Large Language Models
par: Guardieiro, Vitoria, et autres
Publié: (2025)
par: Guardieiro, Vitoria, et autres
Publié: (2025)
Documents similaires
-
Constraint Back-translation Improves Complex Instruction Following of Large Language Models
par: Qi, Yunjia, et autres
Publié: (2024) -
VerIF: Verification Engineering for Reinforcement Learning in Instruction Following
par: Peng, Hao, et autres
Publié: (2025) -
Agentic Reward Modeling: Integrating Human Preferences with Verifiable Correctness Signals for Reliable Reward Systems
par: Peng, Hao, et autres
Publié: (2025) -
On the Paradoxical Interference between Instruction-Following and Task Solving
par: Qi, Yunjia, et autres
Publié: (2026) -
ADELIE: Aligning Large Language Models on Information Extraction
par: Qi, Yunjia, et autres
Publié: (2024)