Benchmark Profiling: Mechanistic Diagnosis of LLM Benchmarks
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Dongjun, Shim, Gyuho, Chun, Yongchan, Kim, Minhyuk, Park, Chanjun, Lim, Heuiseok |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
di: Chun, Yongchan, et al.
Pubblicazione: (2025)
di: Chun, Yongchan, et al.
Pubblicazione: (2025)
Exploring Coding Spot: Understanding Parametric Contributions to LLM Coding Performance
di: Kim, Dongjun, et al.
Pubblicazione: (2024)
di: Kim, Dongjun, et al.
Pubblicazione: (2024)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
di: Park, Chanhee, et al.
Pubblicazione: (2025)
di: Park, Chanhee, et al.
Pubblicazione: (2025)
ChatLang-8: An LLM-Based Synthetic Data Generation Framework for Grammatical Error Correction
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
CharacterGPT: A Persona Reconstruction Framework for Role-Playing Agents
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
di: Jung, Dahyun, et al.
Pubblicazione: (2025)
di: Jung, Dahyun, et al.
Pubblicazione: (2025)
LANGSAE EDITING: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal
di: Kim, Dongjun, et al.
Pubblicazione: (2026)
di: Kim, Dongjun, et al.
Pubblicazione: (2026)
Understanding LLM Development Through Longitudinal Study: Insights from the Open Ko-LLM Leaderboard
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents
di: Shin, Joongmin, et al.
Pubblicazione: (2026)
di: Shin, Joongmin, et al.
Pubblicazione: (2026)
CoME: An Unlearning-based Approach to Conflict-free Model Editing
di: Jung, Dahyun, et al.
Pubblicazione: (2025)
di: Jung, Dahyun, et al.
Pubblicazione: (2025)
Alternative Speech: Complementary Method to Counter-Narrative for Better Discourse
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
di: Lee, Seungyoon, et al.
Pubblicazione: (2024)
Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy
di: Shim, Gyuho, et al.
Pubblicazione: (2026)
di: Shim, Gyuho, et al.
Pubblicazione: (2026)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
From Ambiguity to Accuracy: The Transformative Effect of Coreference Resolution on Retrieval-Augmented Generation systems
di: Jang, Youngjoon, et al.
Pubblicazione: (2025)
di: Jang, Youngjoon, et al.
Pubblicazione: (2025)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
Translation of Multifaceted Data without Re-Training of Machine Translation Systems
di: Moon, Hyeonseok, et al.
Pubblicazione: (2024)
di: Moon, Hyeonseok, et al.
Pubblicazione: (2024)
InstaTrans: An Instruction-Aware Translation Framework for Non-English Instruction Datasets
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
Model-Based Data-Centric AI: Bridging the Divide Between Academic Ideals and Industrial Pragmatism
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
di: Shin, Joongmin, et al.
Pubblicazione: (2026)
di: Shin, Joongmin, et al.
Pubblicazione: (2026)
TORSO: Template-Oriented Reasoning Towards General Tasks
di: Kim, Minhyuk, et al.
Pubblicazione: (2025)
di: Kim, Minhyuk, et al.
Pubblicazione: (2025)
NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models
di: Moon, Hyeonseok, et al.
Pubblicazione: (2025)
di: Moon, Hyeonseok, et al.
Pubblicazione: (2025)
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses
di: Eo, Sugyeong, et al.
Pubblicazione: (2026)
di: Eo, Sugyeong, et al.
Pubblicazione: (2026)
Evalverse: Unified and Accessible Library for Large Language Model Evaluation
di: Kim, Jihoo, et al.
Pubblicazione: (2024)
di: Kim, Jihoo, et al.
Pubblicazione: (2024)
No Reader Left Behind: Multi-Agent Summaries Everyone Can Understand
di: Jung, Jimin, et al.
Pubblicazione: (2026)
di: Jung, Jimin, et al.
Pubblicazione: (2026)
Toward Practical Automatic Speech Recognition and Post-Processing: a Call for Explainable Error Benchmark Guideline
di: Koo, Seonmin, et al.
Pubblicazione: (2024)
di: Koo, Seonmin, et al.
Pubblicazione: (2024)
Dataverse: Open-Source ETL (Extract, Transform, Load) Pipeline for Large Language Models
di: Park, Hyunbyung, et al.
Pubblicazione: (2024)
di: Park, Hyunbyung, et al.
Pubblicazione: (2024)
Analysis of Utterance Embeddings and Clustering Methods Related to Intent Induction for Task-Oriented Dialogue
di: Park, Jeiyoon, et al.
Pubblicazione: (2022)
di: Park, Jeiyoon, et al.
Pubblicazione: (2022)
sDPO: Don't Use Your Data All at Once
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
LP Data Pipeline: Lightweight, Purpose-driven Data Pipeline for Large Language Models
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
di: Kim, Yungi, et al.
Pubblicazione: (2024)
di: Kim, Yungi, et al.
Pubblicazione: (2024)
The Impact of Negated Text on Hallucination with Large Language Models
di: Seo, Jaehyung, et al.
Pubblicazione: (2025)
di: Seo, Jaehyung, et al.
Pubblicazione: (2025)
Call for Rigor in Reporting Quality of Instruction Tuning Data
di: Moon, Hyeonseok, et al.
Pubblicazione: (2025)
di: Moon, Hyeonseok, et al.
Pubblicazione: (2025)
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
KFinEval-Pilot: A Comprehensive Benchmark Suite for Korean Financial Language Understanding
di: Hwang, Bokwang, et al.
Pubblicazione: (2025)
di: Hwang, Bokwang, et al.
Pubblicazione: (2025)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation
di: Yoon, Jeongho, et al.
Pubblicazione: (2026)
di: Yoon, Jeongho, et al.
Pubblicazione: (2026)
Evidential Transformation Network: Turning Pretrained Models into Evidential Models for Post-hoc Uncertainty Estimation
di: Chun, Yongchan, et al.
Pubblicazione: (2026)
di: Chun, Yongchan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
di: Kim, Dongjun, et al.
Pubblicazione: (2025) -
Enhancing Automatic Term Extraction with Large Language Models via Syntactic Retrieval
di: Chun, Yongchan, et al.
Pubblicazione: (2025) -
Exploring Coding Spot: Understanding Parametric Contributions to LLM Coding Performance
di: Kim, Dongjun, et al.
Pubblicazione: (2024) -
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
di: Park, Chanhee, et al.
Pubblicazione: (2025) -
ChatLang-8: An LLM-Based Synthetic Data Generation Framework for Grammatical Error Correction
di: Park, Jeiyoon, et al.
Pubblicazione: (2024)