Metric Calculating Benchmark: Code-Verifiable Complicate Instruction Following Benchmark for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moon, Hyeonseok, Hong, Seongtae, Seo, Jaehyung, Lim, Heuiseok |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Impact of Negated Text on Hallucination with Large Language Models
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025)
Call for Rigor in Reporting Quality of Instruction Tuning Data
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
Cross-Lingual Optimization for Language Transfer in Large Language Models
von: Lee, Jungseob, et al.
Veröffentlicht: (2025)
von: Lee, Jungseob, et al.
Veröffentlicht: (2025)
Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer
von: Lee, Seungyoon, et al.
Veröffentlicht: (2025)
von: Lee, Seungyoon, et al.
Veröffentlicht: (2025)
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
von: Park, Chanhee, et al.
Veröffentlicht: (2025)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
von: Koo, Seonmin, et al.
Veröffentlicht: (2024)
Find the Intention of Instruction: Comprehensive Evaluation of Instruction Understanding for Large Language Models
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2024)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocal
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
Beyond Hard Negatives: The Importance of Score Distribution in Knowledge Distillation for Dense Retrieval
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
CoME: An Unlearning-based Approach to Conflict-free Model Editing
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
von: Jung, Jimin, et al.
Veröffentlicht: (2026)
von: Jung, Jimin, et al.
Veröffentlicht: (2026)
SemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic Bridges
von: Hong, Seongtae, et al.
Veröffentlicht: (2026)
von: Hong, Seongtae, et al.
Veröffentlicht: (2026)
LegalMidm: Use-Case-Driven Legal Domain Specialization for Korean Large Language Model
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2026)
Improving Semantic Proximity in Information Retrieval through Cross-Lingual Alignment
von: Hong, Seongtae, et al.
Veröffentlicht: (2026)
von: Hong, Seongtae, et al.
Veröffentlicht: (2026)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training
von: Lee, Seungyoon, et al.
Veröffentlicht: (2026)
von: Lee, Seungyoon, et al.
Veröffentlicht: (2026)
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
von: Eo, Sugyeong, et al.
Veröffentlicht: (2026)
von: Eo, Sugyeong, et al.
Veröffentlicht: (2026)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
von: Jayakumar, Thanmay, et al.
Veröffentlicht: (2026)
EIFBENCH: Extremely Complex Instruction Following Benchmark for Large Language Models
von: Zou, Tao, et al.
Veröffentlicht: (2025)
von: Zou, Tao, et al.
Veröffentlicht: (2025)
INSTRUCTIR: A Benchmark for Instruction Following of Information Retrieval Models
von: Oh, Hanseok, et al.
Veröffentlicht: (2024)
von: Oh, Hanseok, et al.
Veröffentlicht: (2024)
The SIFo Benchmark: Investigating the Sequential Instruction Following Ability of Large Language Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2024)
von: Chen, Xinyi, et al.
Veröffentlicht: (2024)
AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
von: Qi, Yunjia, et al.
Veröffentlicht: (2025)
Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
von: Kim, Dongjun, et al.
Veröffentlicht: (2025)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
von: Gao, Yiming, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Calculus Problem-Solving: A Comparative Analysis
von: Moon, In Hak
Veröffentlicht: (2025)
von: Moon, In Hak
Veröffentlicht: (2025)
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
von: Oh, Jio, et al.
Veröffentlicht: (2024)
von: Oh, Jio, et al.
Veröffentlicht: (2024)
Post-hoc Utterance Refining Method by Entity Mining for Faithful Knowledge Grounded Conversations
von: Jang, Yoonna, et al.
Veröffentlicht: (2024)
von: Jang, Yoonna, et al.
Veröffentlicht: (2024)
CIF-Bench: A Chinese Instruction-Following Benchmark for Evaluating the Generalizability of Large Language Models
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
von: LI, Yizhi, et al.
Veröffentlicht: (2024)
Generalizing Verifiable Instruction Following
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
von: Pyatkin, Valentina, et al.
Veröffentlicht: (2025)
Improving Korean-English Cross-Lingual Retrieval: A Data-Centric Study of Language Composition and Model Merging
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
von: Jang, Youngjoon, et al.
Veröffentlicht: (2025)
Llamion Technical Report
von: Yang, Kisu, et al.
Veröffentlicht: (2026)
von: Yang, Kisu, et al.
Veröffentlicht: (2026)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
M3DocDep: Multi-modal, Multi-page, Multi-document Dependency Chunking with Large Vision-Language Models
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
von: Shin, Joongmin, et al.
Veröffentlicht: (2026)
Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy
von: Shim, Gyuho, et al.
Veröffentlicht: (2026)
von: Shim, Gyuho, et al.
Veröffentlicht: (2026)
Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
von: Chun, Yongchan, et al.
Veröffentlicht: (2025)
von: Chun, Yongchan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Impact of Negated Text on Hallucination with Large Language Models
von: Seo, Jaehyung, et al.
Veröffentlicht: (2025) -
Call for Rigor in Reporting Quality of Instruction Tuning Data
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025) -
Cross-Lingual Optimization for Language Transfer in Large Language Models
von: Lee, Jungseob, et al.
Veröffentlicht: (2025) -
Semantic Aware Linear Transfer by Recycling Pre-trained Language Models for Cross-lingual Transfer
von: Lee, Seungyoon, et al.
Veröffentlicht: (2025) -
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models
von: Moon, Hyeonseok, et al.
Veröffentlicht: (2025)