Models Know Models Best: Evaluation via Model-Preferred Formats
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Joonhak, Jung, Sungmok, Park, Jongyeon, Lee, Jaejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
von: Jung, Sungmok, et al.
Veröffentlicht: (2026)
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025)
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
von: Kim, Seorin, et al.
Veröffentlicht: (2025)
von: Kim, Seorin, et al.
Veröffentlicht: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
von: Park, Chanwoo, et al.
Veröffentlicht: (2025)
Language Models Prefer What They Know: Relative Confidence Estimation via Confidence Preferences
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
von: Shrivastava, Vaishnavi, et al.
Veröffentlicht: (2025)
Thunder-DeID: Accurate and Efficient De-identification Framework for Korean Court Judgments
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
von: Hahm, Sungeun, et al.
Veröffentlicht: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
Pragmatic Competence Evaluation of Large Language Models for the Korean Language
von: Park, Dojun, et al.
Veröffentlicht: (2024)
von: Park, Dojun, et al.
Veröffentlicht: (2024)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
von: Jung, Jeesu, et al.
Veröffentlicht: (2025)
von: Jung, Jeesu, et al.
Veröffentlicht: (2025)
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment
von: Lee, Janghwan, et al.
Veröffentlicht: (2024)
von: Lee, Janghwan, et al.
Veröffentlicht: (2024)
Unsupervised Extractive Dialogue Summarization in Hyperdimensional Space
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
von: Park, Seongmin, et al.
Veröffentlicht: (2024)
Diverging Preferences: When do Annotators Disagree and do Models Know?
von: Zhang, Michael JQ, et al.
Veröffentlicht: (2024)
von: Zhang, Michael JQ, et al.
Veröffentlicht: (2024)
Reasoning by Commented Code for Table Question Answering
von: Pyo, Seho, et al.
Veröffentlicht: (2026)
von: Pyo, Seho, et al.
Veröffentlicht: (2026)
SEDD: Scalable and Efficient Dataset Deduplication with GPUs
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
von: Son, Youngjun, et al.
Veröffentlicht: (2025)
FLEX: A Benchmark for Evaluating Robustness of Fairness in Large Language Models
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
von: Jung, Dahyun, et al.
Veröffentlicht: (2025)
MultiPragEval: Multilingual Pragmatic Evaluation of Large Language Models
von: Park, Dojun, et al.
Veröffentlicht: (2024)
von: Park, Dojun, et al.
Veröffentlicht: (2024)
Fine-Tuning Language Models to Know What They Know
von: Park, Sangjun, et al.
Veröffentlicht: (2026)
von: Park, Sangjun, et al.
Veröffentlicht: (2026)
HAE-RAE Bench: Evaluation of Korean Knowledge in Language Models
von: Son, Guijin, et al.
Veröffentlicht: (2023)
von: Son, Guijin, et al.
Veröffentlicht: (2023)
Investigating Language Preference of Multilingual RAG Systems
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
von: Park, Jeonghyun, et al.
Veröffentlicht: (2025)
Choices Speak Louder than Questions
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
von: Han, Hojae, et al.
Veröffentlicht: (2024)
von: Han, Hojae, et al.
Veröffentlicht: (2024)
Aligning to Thousands of Preferences via System Message Generalization
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
von: Jung, Sunghee, et al.
Veröffentlicht: (2025)
von: Jung, Sunghee, et al.
Veröffentlicht: (2025)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
Thunder-LLM: Efficiently Adapting LLMs to Korean with Minimal Resources
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
von: Kim, Jinpyo, et al.
Veröffentlicht: (2025)
TeXBLEU: Automatic Metric for Evaluate LaTeX Format
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
von: Jung, Kyudan, et al.
Veröffentlicht: (2024)
Return of EM: Entity-driven Answer Set Expansion for QA Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2024)
Structured Language Generation Model: Loss Calibration and Formatted Decoding for Robust Structure Prediction and Knowledge Retrieval
von: Lee, Minho, et al.
Veröffentlicht: (2024)
von: Lee, Minho, et al.
Veröffentlicht: (2024)
ORPO: Monolithic Preference Optimization without Reference Model
von: Hong, Jiwoo, et al.
Veröffentlicht: (2024)
von: Hong, Jiwoo, et al.
Veröffentlicht: (2024)
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
von: Kim, Seunghee, et al.
Veröffentlicht: (2026)
von: Kim, Seunghee, et al.
Veröffentlicht: (2026)
SCALE: Upscaled Continual Learning of Large Language Models
von: Lee, Jin-woo, et al.
Veröffentlicht: (2025)
von: Lee, Jin-woo, et al.
Veröffentlicht: (2025)
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
von: Jo, Hwiyeol, et al.
Veröffentlicht: (2025)
Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
von: Balepur, Nishant, et al.
Veröffentlicht: (2026)
The Comparative Trap: Pairwise Comparisons Amplifies Biased Preferences of LLM Evaluators
von: Jeong, Hawon, et al.
Veröffentlicht: (2024)
von: Jeong, Hawon, et al.
Veröffentlicht: (2024)
Models That Know How Evaluations Are Designed Score Safer
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
Evaluation of Best-of-N Sampling Strategies for Language Model Alignment
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
von: Ichihara, Yuki, et al.
Veröffentlicht: (2025)
KnowRL: Teaching Language Models to Know What They Know
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams
von: Lee, Yukyung, et al.
Veröffentlicht: (2026)
von: Lee, Yukyung, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
von: Jung, Sungmok, et al.
Veröffentlicht: (2026) -
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
von: So, Yeonkyoung, et al.
Veröffentlicht: (2025) -
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models
von: Cho, Gyeongje, et al.
Veröffentlicht: (2025) -
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
von: Park, Chanwoo, et al.
Veröffentlicht: (2025) -
KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
von: Kim, Seorin, et al.
Veröffentlicht: (2025)