LIFBench: Evaluating the Instruction Following Performance and Stability of Large Language Models in Long-Context Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiaodong, Wang, Minhao, Liu, Yichen, Shi, Xiaoming, Yan, He, Lu, Xiangju, Zhu, Junmin, Zhang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels
by: Wei, Lingxiao, et al.
Published: (2024)
by: Wei, Lingxiao, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
by: Qi, Yunjia, et al.
Published: (2025)
by: Qi, Yunjia, et al.
Published: (2025)
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning
by: Lu, Yuheng, et al.
Published: (2025)
by: Lu, Yuheng, et al.
Published: (2025)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
by: Sun, Wangtao, et al.
Published: (2024)
by: Sun, Wangtao, et al.
Published: (2024)
Evaluating Large Language Models at Evaluating Instruction Following
by: Zeng, Zhiyuan, et al.
Published: (2023)
by: Zeng, Zhiyuan, et al.
Published: (2023)
LIFEBench: Evaluating Length Instruction Following in Large Language Models
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Instruction-Following Evaluation of Large Vision-Language Models
by: Shiono, Daiki, et al.
Published: (2025)
by: Shiono, Daiki, et al.
Published: (2025)
LongCodeZip: Compress Long Context for Code Language Models
by: Shi, Yuling, et al.
Published: (2025)
by: Shi, Yuling, et al.
Published: (2025)
LongSafety: Evaluating Long-Context Safety of Large Language Models
by: Lu, Yida, et al.
Published: (2025)
by: Lu, Yida, et al.
Published: (2025)
InFoBench: Evaluating Instruction Following Ability in Large Language Models
by: Qin, Yiwei, et al.
Published: (2024)
by: Qin, Yiwei, et al.
Published: (2024)
WildLong: Synthesizing Realistic Long-Context Instruction Data at Scale
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
by: Zeng, Bo, et al.
Published: (2025)
by: Zeng, Bo, et al.
Published: (2025)
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Instruction-Following Pruning for Large Language Models
by: Hou, Bairu, et al.
Published: (2025)
by: Hou, Bairu, et al.
Published: (2025)
Long Context Alignment with Short Instructions and Synthesized Positions
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models
by: Lian, Haoran, et al.
Published: (2024)
by: Lian, Haoran, et al.
Published: (2024)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
IHEval: Evaluating Language Models on Following the Instruction Hierarchy
by: Zhang, Zhihan, et al.
Published: (2025)
by: Zhang, Zhihan, et al.
Published: (2025)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
MulDimIF: A Multi-Dimensional Constraint Framework for Evaluating and Improving Instruction Following in Large Language Models
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
by: Zhang, Kaiyan, et al.
Published: (2024)
by: Zhang, Kaiyan, et al.
Published: (2024)
Dynamic Chunking for Diffusion Language Models
by: Zhu, Yichen, et al.
Published: (2026)
by: Zhu, Yichen, et al.
Published: (2026)
From Complex to Simple: Enhancing Multi-Constraint Complex Instruction Following Ability of Large Language Models
by: He, Qianyu, et al.
Published: (2024)
by: He, Qianyu, et al.
Published: (2024)
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models
by: Hu, Zhiyuan, et al.
Published: (2024)
by: Hu, Zhiyuan, et al.
Published: (2024)
Thus Spake Long-Context Large Language Model
by: Liu, Xiaoran, et al.
Published: (2025)
by: Liu, Xiaoran, et al.
Published: (2025)
Revisiting the Reliability of Language Models in Instruction-Following
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Can LLMs Simulate Personas with Reversed Performance? A Systematic Investigation for Counterfactual Instruction Following in Math Reasoning Context
by: Kumar, Sai Adith Senthil, et al.
Published: (2025)
by: Kumar, Sai Adith Senthil, et al.
Published: (2025)
UniMem: Towards a Unified View of Long-Context Large Language Models
by: Fang, Junjie, et al.
Published: (2024)
by: Fang, Junjie, et al.
Published: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Span-level Emotion-Cause-Category Triplet Extraction with Instruction Tuning LLMs and Data Augmentation
by: Li, Xiangju, et al.
Published: (2025)
by: Li, Xiangju, et al.
Published: (2025)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
Retrieval meets Long Context Large Language Models
by: Xu, Peng, et al.
Published: (2023)
by: Xu, Peng, et al.
Published: (2023)
Evaluating the Translation Performance of Large Language Models Based on Euas-20
by: Huang, Yan, et al.
Published: (2024)
by: Huang, Yan, et al.
Published: (2024)
Long Context is Not Long at All: A Prospector of Long-Dependency Data for Large Language Models
by: Chen, Longze, et al.
Published: (2024)
by: Chen, Longze, et al.
Published: (2024)
On Many-Shot In-Context Learning for Long-Context Evaluation
by: Zou, Kaijian, et al.
Published: (2024)
by: Zou, Kaijian, et al.
Published: (2024)
ReIFE: Re-evaluating Instruction-Following Evaluation
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Similar Items
-
CNNSum: Exploring Long-Context Summarization with Large Language Models in Chinese Novels
by: Wei, Lingxiao, et al.
Published: (2024) -
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025) -
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
by: Qi, Yunjia, et al.
Published: (2025) -
XIFBench: Evaluating Large Language Models on Multilingual Instruction Following
by: Li, Zhenyu, et al.
Published: (2025) -
Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning
by: Lu, Yuheng, et al.
Published: (2025)