Benchmarks Are Not That Out of Distribution: Word Overlap Predicts Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Chung, Woojin, Kim, Jeonghoon |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
by: Chung, Woojin, et al.
Published: (2025)
by: Chung, Woojin, et al.
Published: (2025)
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025)
by: Kim, Nayeon, et al.
Published: (2025)
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
by: Kim, Seunghee, et al.
Published: (2026)
by: Kim, Seunghee, et al.
Published: (2026)
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
by: Choi, Yunho, et al.
Published: (2026)
by: Choi, Yunho, et al.
Published: (2026)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Benchmarking LLMs on the Semantic Overlap Summarization Task
by: Salvador, John, et al.
Published: (2024)
by: Salvador, John, et al.
Published: (2024)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
by: Wu, Siyang, et al.
Published: (2025)
by: Wu, Siyang, et al.
Published: (2025)
Constructions are Revealed in Word Distributions
by: Rozner, Joshua, et al.
Published: (2025)
by: Rozner, Joshua, et al.
Published: (2025)
Handling Ambiguity in Emotion: From Out-of-Domain Detection to Distribution Estimation
by: Wu, Wen, et al.
Published: (2024)
by: Wu, Wen, et al.
Published: (2024)
Is Cross-Lingual Transfer in Bilingual Models Human-Like? A Study with Overlapping Word Forms in Dutch and English
by: Škrjanec, Iza, et al.
Published: (2026)
by: Škrjanec, Iza, et al.
Published: (2026)
Improving Multi-hop Logical Reasoning in Knowledge Graphs with Context-Aware Query Representation Learning
by: Kim, Jeonghoon, et al.
Published: (2024)
by: Kim, Jeonghoon, et al.
Published: (2024)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
by: Wu, Zhaofeng, et al.
Published: (2025)
by: Wu, Zhaofeng, et al.
Published: (2025)
Words that Matter: The Impact of Negative Words on News Sentiment and Stock Market Index
by: Kim, Wonseong
Published: (2023)
by: Kim, Wonseong
Published: (2023)
Stable Language Model Pre-training by Reducing Embedding Variability
by: Chung, Woojin, et al.
Published: (2024)
by: Chung, Woojin, et al.
Published: (2024)
Can You Learn Semantics Through Next-Word Prediction? The Case of Entailment
by: Merrill, William, et al.
Published: (2024)
by: Merrill, William, et al.
Published: (2024)
On Support Samples of Next Word Prediction
by: Li, Yuqian, et al.
Published: (2025)
by: Li, Yuqian, et al.
Published: (2025)
Word Synchronization Challenge: A Benchmark for Word Association Responses for Large Language Models
by: Cazalets, Tanguy, et al.
Published: (2025)
by: Cazalets, Tanguy, et al.
Published: (2025)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
by: Verhoeven, Ivo, et al.
Published: (2024)
by: Verhoeven, Ivo, et al.
Published: (2024)
Broken Words, Broken Performance: Effect of Tokenization on Performance of LLMs
by: Pawar, Sachin, et al.
Published: (2025)
by: Pawar, Sachin, et al.
Published: (2025)
Empirical Study of Named Entity Recognition Performance Using Distribution-aware Word Embedding
by: Chen, Xin, et al.
Published: (2021)
by: Chen, Xin, et al.
Published: (2021)
The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks
by: Schlechtweg, Dominik, et al.
Published: (2024)
by: Schlechtweg, Dominik, et al.
Published: (2024)
ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language Models
by: Shim, Jeonghoon, et al.
Published: (2025)
by: Shim, Jeonghoon, et al.
Published: (2025)
Using Correspondence Patterns to Identify Irregular Words in Cognate sets Through Leave-One-Out Validation
by: Blum, Frederic, et al.
Published: (2026)
by: Blum, Frederic, et al.
Published: (2026)
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
UniCoM: A Universal Code-Switching Speech Generator
by: Lee, Sangmin, et al.
Published: (2025)
by: Lee, Sangmin, et al.
Published: (2025)
How Good Are LLMs at Out-of-Distribution Detection?
by: Liu, Bo, et al.
Published: (2023)
by: Liu, Bo, et al.
Published: (2023)
CO2: Efficient Distributed Training with Full Communication-Computation Overlap
by: Sun, Weigao, et al.
Published: (2024)
by: Sun, Weigao, et al.
Published: (2024)
VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models
by: Kim, Woojin, et al.
Published: (2026)
by: Kim, Woojin, et al.
Published: (2026)
BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
by: Shen, Ming, et al.
Published: (2025)
by: Shen, Ming, et al.
Published: (2025)
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
by: Sun, Yuhong, et al.
Published: (2024)
by: Sun, Yuhong, et al.
Published: (2024)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
From Words to Worth: Newborn Article Impact Prediction with LLM
by: Zhao, Penghai, et al.
Published: (2024)
by: Zhao, Penghai, et al.
Published: (2024)
No Encore: Unlearning as Opt-Out in Music Generation
by: Kim, Jinju, et al.
Published: (2025)
by: Kim, Jinju, et al.
Published: (2025)
Internal Causal Mechanisms Robustly Predict Language Model Out-of-Distribution Behaviors
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
A Survey on Out-of-Distribution Evaluation of Neural NLP Models
by: Li, Xinzhe, et al.
Published: (2023)
by: Li, Xinzhe, et al.
Published: (2023)
TAIA: Large Language Models are Out-of-Distribution Data Learners
by: Jiang, Shuyang, et al.
Published: (2024)
by: Jiang, Shuyang, et al.
Published: (2024)
BED: Bi-Encoder-Based Detectors for Out-of-Distribution Detection
by: Owen, Louis, et al.
Published: (2023)
by: Owen, Louis, et al.
Published: (2023)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning
by: Jiang, Juyong, et al.
Published: (2026)
by: Jiang, Juyong, et al.
Published: (2026)
Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
by: Anantheswaran, Ujjwala, et al.
Published: (2024)
Similar Items
-
Exploiting Vocabulary Frequency Imbalance in Language Model Pre-training
by: Chung, Woojin, et al.
Published: (2025) -
Handling Korean Out-of-Vocabulary Words with Phoneme Representation Learning
by: Kim, Nayeon, et al.
Published: (2025) -
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
by: Kim, Seunghee, et al.
Published: (2026) -
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
by: Choi, Yunho, et al.
Published: (2026) -
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)