InFoBench: Evaluating Instruction Following Ability in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Qin, Yiwei, Song, Kaiqiang, Hu, Yebowen, Yao, Wenlin, Cho, Sangwoo, Wang, Xiaoyang, Wu, Xuansheng, Liu, Fei, Liu, Pengfei, Yu, Dong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Large Language Models do Analytical Reasoning?
by: Hu, Yebowen, et al.
Published: (2024)
by: Hu, Yebowen, et al.
Published: (2024)
When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
by: Hu, Yebowen, et al.
Published: (2024)
by: Hu, Yebowen, et al.
Published: (2024)
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
by: Hu, Yebowen, et al.
Published: (2024)
by: Hu, Yebowen, et al.
Published: (2024)
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
by: Liu, Fuxiao, et al.
Published: (2023)
by: Liu, Fuxiao, et al.
Published: (2023)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
SPECTRUM: Speaker-Enhanced Pre-Training for Long Dialogue Summarization
by: Cho, Sangwoo, et al.
Published: (2024)
by: Cho, Sangwoo, et al.
Published: (2024)
A Versatile Multimodal Agent for Multimedia Content Generation
by: Zhang, Daoan, et al.
Published: (2026)
by: Zhang, Daoan, et al.
Published: (2026)
DeFine: Decision-Making with Analogical Reasoning over Factor Profiles
by: Hu, Yebowen, et al.
Published: (2024)
by: Hu, Yebowen, et al.
Published: (2024)
Polarity Calibration for Opinion Summarization
by: Lei, Yuanyuan, et al.
Published: (2024)
by: Lei, Yuanyuan, et al.
Published: (2024)
TCIA: A Task-Centric Instruction Augmentation Method for Instruction Finetuning
by: Ma, Simin, et al.
Published: (2025)
by: Ma, Simin, et al.
Published: (2025)
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models
by: Sun, Haoran, et al.
Published: (2024)
by: Sun, Haoran, et al.
Published: (2024)
Could Small Language Models Serve as Recommenders? Towards Data-centric Cold-start Recommendations
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
Complex Logical Instruction Generation
by: Zhang, Mian, et al.
Published: (2025)
by: Zhang, Mian, et al.
Published: (2025)
Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
Evaluation of Instruction-Following Ability for Large Language Models on Story-Ending Generation
by: Hida, Rem, et al.
Published: (2024)
by: Hida, Rem, et al.
Published: (2024)
Soundness-Aware Level: A Microscopic Signature that Predicts LLM Reasoning Potential
by: Wu, Xuansheng, et al.
Published: (2025)
by: Wu, Xuansheng, et al.
Published: (2025)
RefuteBench: Evaluating Refuting Instruction-Following for Large Language Models
by: Yan, Jianhao, et al.
Published: (2024)
by: Yan, Jianhao, et al.
Published: (2024)
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
by: Chu, Zheng, et al.
Published: (2023)
by: Chu, Zheng, et al.
Published: (2023)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
by: Wang, Peiding, et al.
Published: (2025)
by: Wang, Peiding, et al.
Published: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
by: Kim, Dongjun, et al.
Published: (2025)
by: Kim, Dongjun, et al.
Published: (2025)
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions
by: Cassano, Federico, et al.
Published: (2023)
by: Cassano, Federico, et al.
Published: (2023)
Deconstructing Instruction-Following: A New Benchmark for Granular Evaluation of Large Language Model Instruction Compliance Abilities
by: Purpura, Alberto, et al.
Published: (2026)
by: Purpura, Alberto, et al.
Published: (2026)
Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models
by: Zeng, Bo, et al.
Published: (2025)
by: Zeng, Bo, et al.
Published: (2025)
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
by: Chen, Xinyi, et al.
Published: (2024)
by: Chen, Xinyi, et al.
Published: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
by: LI, Yizhi, et al.
Published: (2024)
by: LI, Yizhi, et al.
Published: (2024)
Skills-in-Context Prompting: Unlocking Compositionality in Large Language Models
by: Chen, Jiaao, et al.
Published: (2023)
by: Chen, Jiaao, et al.
Published: (2023)
Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
by: Sun, Wangtao, et al.
Published: (2024)
by: Sun, Wangtao, et al.
Published: (2024)
Revisiting Compositional Generalization Capability of Large Language Models Considering Instruction Following Ability
by: Sakai, Yusuke, et al.
Published: (2025)
by: Sakai, Yusuke, et al.
Published: (2025)
StakeBench: Evaluating Language Understanding Grounded in Market Commitment
by: Pei, Yunhua, et al.
Published: (2026)
by: Pei, Yunhua, et al.
Published: (2026)
Evaluating Large Language Models at Evaluating Instruction Following
by: Zeng, Zhiyuan, et al.
Published: (2023)
by: Zeng, Zhiyuan, et al.
Published: (2023)
MathChat: Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
by: Liang, Zhenwen, et al.
Published: (2024)
by: Liang, Zhenwen, et al.
Published: (2024)
A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
TeachBench: A Syllabus-Grounded Framework for Evaluating Teaching Ability in Large Language Models
by: Li, Zheng, et al.
Published: (2026)
by: Li, Zheng, et al.
Published: (2026)
STRUX: An LLM for Decision-Making with Structured Explanations
by: Lu, Yiming, et al.
Published: (2024)
by: Lu, Yiming, et al.
Published: (2024)
LIFEBench: Evaluating Length Instruction Following in Large Language Models
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging
by: He, Jinlong, et al.
Published: (2024)
by: He, Jinlong, et al.
Published: (2024)
DIVE: Diversified Iterative Self-Improvement
by: Qin, Yiwei, et al.
Published: (2025)
by: Qin, Yiwei, et al.
Published: (2025)
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
by: Wang, Xiaoxuan, et al.
Published: (2023)
by: Wang, Xiaoxuan, et al.
Published: (2023)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
by: Wu, Yunze, et al.
Published: (2025)
by: Wu, Yunze, et al.
Published: (2025)
Similar Items
-
Can Large Language Models do Analytical Reasoning?
by: Hu, Yebowen, et al.
Published: (2024) -
When Reasoning Meets Information Aggregation: A Case Study with Sports Narratives
by: Hu, Yebowen, et al.
Published: (2024) -
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
by: Hu, Yebowen, et al.
Published: (2024) -
MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
by: Liu, Fuxiao, et al.
Published: (2023) -
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)