A Single Character can Make or Break Your LLM Evals
Fuente:
arXiv
Saved in:
| Main Authors: | Su, Jingtong, Zhang, Jianyu, Ullrich, Karen, Bottou, Léon, Ibrahim, Mark |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024)
by: Su, Jingtong, et al.
Published: (2024)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025)
by: Su, Jingtong, et al.
Published: (2025)
Memory Mosaics at scale
by: Zhang, Jianyu, et al.
Published: (2025)
by: Zhang, Jianyu, et al.
Published: (2025)
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024)
by: An, Shengnan, et al.
Published: (2024)
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026)
by: Wang, Xiangqi, et al.
Published: (2026)
These Are Not All the Features You Are Looking For: A Fundamental Bottleneck in Supervised Pretraining
by: Yang, Xingyu Alice, et al.
Published: (2025)
by: Yang, Xingyu Alice, et al.
Published: (2025)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
by: Lee, Jaebok, et al.
Published: (2025)
by: Lee, Jaebok, et al.
Published: (2025)
Memory Mosaics
by: Zhang, Jianyu, et al.
Published: (2024)
by: Zhang, Jianyu, et al.
Published: (2024)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
by: Fish, Sara, et al.
Published: (2025)
by: Fish, Sara, et al.
Published: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
by: Khatun, Aisha, et al.
Published: (2024)
by: Khatun, Aisha, et al.
Published: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
by: Zhao, Jiahao, et al.
Published: (2025)
by: Zhao, Jiahao, et al.
Published: (2025)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
by: McCammon, James
Published: (2025)
by: McCammon, James
Published: (2025)
Is Your LLM Outdated? A Deep Look at Temporal Generalization
by: Zhu, Chenghao, et al.
Published: (2024)
by: Zhu, Chenghao, et al.
Published: (2024)
Single Character Perturbations Break LLM Alignment
by: Lin, Leon, et al.
Published: (2024)
by: Lin, Leon, et al.
Published: (2024)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
by: Xiong, Zhen, et al.
Published: (2025)
by: Xiong, Zhen, et al.
Published: (2025)
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
by: Chun, Ye Eun, et al.
Published: (2025)
by: Chun, Ye Eun, et al.
Published: (2025)
AcademicEval: Live Long-Context LLM Benchmark
by: Zhang, Haozhen, et al.
Published: (2025)
by: Zhang, Haozhen, et al.
Published: (2025)
Measuring all the noises of LLM Evals
by: Wang, Sida
Published: (2025)
by: Wang, Sida
Published: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
by: D'Souza, Jennifer, et al.
Published: (2025)
by: D'Souza, Jennifer, et al.
Published: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
by: Yang, Langqi, et al.
Published: (2025)
by: Yang, Langqi, et al.
Published: (2025)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
by: Khan, Haidar, et al.
Published: (2025)
by: Khan, Haidar, et al.
Published: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
by: Zhou, Cai, et al.
Published: (2025)
by: Zhou, Cai, et al.
Published: (2025)
DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer
by: Hong, Junyuan, et al.
Published: (2023)
by: Hong, Junyuan, et al.
Published: (2023)
Collaborative Quest Completion with LLM-driven Non-Player Characters in Minecraft
by: Rao, Sudha, et al.
Published: (2024)
by: Rao, Sudha, et al.
Published: (2024)
Experiences Build Characters: The Linguistic Origins and Functional Impact of LLM Personality
by: Wang, Xi, et al.
Published: (2026)
by: Wang, Xi, et al.
Published: (2026)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
by: Liu, Ryan, et al.
Published: (2024)
by: Liu, Ryan, et al.
Published: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
by: Ren, Huimin, et al.
Published: (2025)
by: Ren, Huimin, et al.
Published: (2025)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
by: Ullrich, Herbert, et al.
Published: (2025)
by: Ullrich, Herbert, et al.
Published: (2025)
Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
by: Silver, Sam, et al.
Published: (2025)
by: Silver, Sam, et al.
Published: (2025)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
by: Alyahya, Hisham A., et al.
Published: (2025)
by: Alyahya, Hisham A., et al.
Published: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence
by: Wu, Kevin, et al.
Published: (2024)
by: Wu, Kevin, et al.
Published: (2024)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
by: Zhao, Jingqian, et al.
Published: (2025)
by: Zhao, Jingqian, et al.
Published: (2025)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
by: Xu, Shuhang, et al.
Published: (2025)
by: Xu, Shuhang, et al.
Published: (2025)
Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
by: Buakhaw, Pasin, et al.
Published: (2025)
by: Buakhaw, Pasin, et al.
Published: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
by: Phan, Buu, et al.
Published: (2024)
by: Phan, Buu, et al.
Published: (2024)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
by: Chen, Xingwu, et al.
Published: (2026)
by: Chen, Xingwu, et al.
Published: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
by: Luo, Sijia, et al.
Published: (2026)
by: Luo, Sijia, et al.
Published: (2026)
Who can we trust? LLM-as-a-jury for Comparative Assessment
by: Qian, Mengjie, et al.
Published: (2026)
by: Qian, Mengjie, et al.
Published: (2026)
Similar Items
-
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
by: Su, Jingtong, et al.
Published: (2024) -
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
by: Su, Jingtong, et al.
Published: (2025) -
Memory Mosaics at scale
by: Zhang, Jianyu, et al.
Published: (2025) -
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024) -
Dual Optimal: Make Your LLM Peer-like with Dignity
by: Wang, Xiangqi, et al.
Published: (2026)