INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Hanseok, Lee, Hyunji, Ye, Seonghyeon, Shin, Haebin, Jang, Hansol, Jun, Changwook, Seo, Minjoon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KTRL+F: Knowledge-Augmented In-Document Search
by: Oh, Hanseok, et al.
Published: (2023)
by: Oh, Hanseok, et al.
Published: (2023)
Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis
by: Yang, Sohee, et al.
Published: (2023)
by: Yang, Sohee, et al.
Published: (2023)
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
by: Won, Yunjae, et al.
Published: (2025)
by: Won, Yunjae, et al.
Published: (2025)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)
by: Hwang, Hyeonbin, et al.
Published: (2024)
Generative Prompt Internalization
by: Shin, Haebin, et al.
Published: (2024)
by: Shin, Haebin, et al.
Published: (2024)
Semiparametric Token-Sequence Co-Supervision
by: Lee, Hyunji, et al.
Published: (2024)
by: Lee, Hyunji, et al.
Published: (2024)
How Well Do Large Language Models Truly Ground?
by: Lee, Hyunji, et al.
Published: (2023)
by: Lee, Hyunji, et al.
Published: (2023)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
KoDialogBench: Evaluating Conversational Understanding of Language Models with Korean Dialogue Benchmark
by: Jang, Seongbo, et al.
Published: (2024)
by: Jang, Seongbo, et al.
Published: (2024)
On the Effectiveness of Integration Methods for Multimodal Dialogue Response Retrieval
by: Jang, Seongbo, et al.
Published: (2025)
by: Jang, Seongbo, et al.
Published: (2025)
Understanding and Enhancing Mamba-Transformer Hybrids for Memory Recall and Language Modeling
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
by: Chang, Hoyeon, et al.
Published: (2024)
by: Chang, Hoyeon, et al.
Published: (2024)
Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
by: Kim, Jiyeon, et al.
Published: (2026)
by: Kim, Jiyeon, et al.
Published: (2026)
Knowledge Entropy Decay during Language Model Pretraining Hinders New Knowledge Acquisition
by: Kim, Jiyeon, et al.
Published: (2024)
by: Kim, Jiyeon, et al.
Published: (2024)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
by: Jin, Rihui, et al.
Published: (2026)
by: Jin, Rihui, et al.
Published: (2026)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
LGAI-EMBEDDING-Preview Technical Report
by: Choi, Jooyoung, et al.
Published: (2025)
by: Choi, Jooyoung, et al.
Published: (2025)
Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
by: Moon, Hyeonseok, et al.
Published: (2025)
by: Moon, Hyeonseok, et al.
Published: (2025)
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)
by: Lee, Changho, et al.
Published: (2024)
BiasJailbreak:Analyzing Ethical Biases and Jailbreak Vulnerabilities in Large Language Models
by: Lee, Isack, et al.
Published: (2024)
by: Lee, Isack, et al.
Published: (2024)
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
by: Lee, Seongyun, et al.
Published: (2026)
by: Lee, Seongyun, et al.
Published: (2026)
Rethinking the Role of Proxy Rewards in Language Model Alignment
by: Kim, Sungdong, et al.
Published: (2024)
by: Kim, Sungdong, et al.
Published: (2024)
Exploring Adversarial Robustness in Classification tasks using DNA Language Models
by: Yoo, Hyunwoo, et al.
Published: (2024)
by: Yoo, Hyunwoo, et al.
Published: (2024)
TSLM: Tree-Structured Language Modeling for Divergent Thinking
by: Kim, Doyoung, et al.
Published: (2026)
by: Kim, Doyoung, et al.
Published: (2026)
FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets
by: Ye, Seonghyeon, et al.
Published: (2023)
by: Ye, Seonghyeon, et al.
Published: (2023)
Eliciting Instruction-tuned Code Language Models' Capabilities to Utilize Auxiliary Function for Code Generation
by: Lee, Seonghyeon, et al.
Published: (2024)
by: Lee, Seonghyeon, et al.
Published: (2024)
IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
by: Song, Tingyu, et al.
Published: (2025)
by: Song, Tingyu, et al.
Published: (2025)
On Efficient Language and Vision Assistants for Visually-Situated Natural Language Understanding: What Matters in Reading and Reasoning
by: Kim, Geewook, et al.
Published: (2024)
by: Kim, Geewook, et al.
Published: (2024)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
by: BehnamGhader, Parishad, et al.
Published: (2025)
by: BehnamGhader, Parishad, et al.
Published: (2025)
How language models extrapolate outside the training data: A case study in Textualized Gridworld
by: Kim, Doyoung, et al.
Published: (2024)
by: Kim, Doyoung, et al.
Published: (2024)
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
by: Shin, Haebin, et al.
Published: (2025)
by: Shin, Haebin, et al.
Published: (2025)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
Exploring Language Model's Code Generation Ability with Auxiliary Functions
by: Lee, Seonghyeon, et al.
Published: (2024)
by: Lee, Seonghyeon, et al.
Published: (2024)
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
by: Weller, Orion, et al.
Published: (2025)
by: Weller, Orion, et al.
Published: (2025)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
by: Shin, Haebin, et al.
Published: (2025)
by: Shin, Haebin, et al.
Published: (2025)
EHRSQL: A Practical Text-to-SQL Benchmark for Electronic Health Records
by: Lee, Gyubok, et al.
Published: (2023)
by: Lee, Gyubok, et al.
Published: (2023)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023)
by: Kim, Seungone, et al.
Published: (2023)
DAHL: Domain-specific Automated Hallucination Evaluation of Long-Form Text through a Benchmark Dataset in Biomedicine
by: Seo, Jean, et al.
Published: (2024)
by: Seo, Jean, et al.
Published: (2024)
Towards Better Instruction Following Retrieval Models
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
Similar Items
-
KTRL+F: Knowledge-Augmented In-Document Search
by: Oh, Hanseok, et al.
Published: (2023) -
Improving Probability-based Prompt Selection Through Unified Evaluation and Analysis
by: Yang, Sohee, et al.
Published: (2023) -
Instruction Tuning with and without Context: Behavioral Shifts and Downstream Impact
by: Lee, Hyunji, et al.
Published: (2025) -
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
by: Won, Yunjae, et al.
Published: (2025) -
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
by: Hwang, Hyeonbin, et al.
Published: (2024)