A Single Character can Make or Break Your LLM Evals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Jingtong, Zhang, Jianyu, Ullrich, Karen, Bottou, Léon, Ibrahim, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
Memory Mosaics at scale
von: Zhang, Jianyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jianyu, et al.
Veröffentlicht: (2025)
Make Your LLM Fully Utilize the Context
von: An, Shengnan, et al.
Veröffentlicht: (2024)
von: An, Shengnan, et al.
Veröffentlicht: (2024)
Dual Optimal: Make Your LLM Peer-like with Dignity
von: Wang, Xiangqi, et al.
Veröffentlicht: (2026)
von: Wang, Xiangqi, et al.
Veröffentlicht: (2026)
These Are Not All the Features You Are Looking For: A Fundamental Bottleneck in Supervised Pretraining
von: Yang, Xingyu Alice, et al.
Veröffentlicht: (2025)
von: Yang, Xingyu Alice, et al.
Veröffentlicht: (2025)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
Memory Mosaics
von: Zhang, Jianyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyu, et al.
Veröffentlicht: (2024)
EconEvals: Benchmarks and Litmus Tests for Economic Decision-Making by LLM Agents
von: Fish, Sara, et al.
Veröffentlicht: (2025)
von: Fish, Sara, et al.
Veröffentlicht: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahao, et al.
Veröffentlicht: (2025)
TimeStampEval: A Simple LLM Eval and a Little Fuzzy Matching Trick to Improve Search Accuracy
von: McCammon, James
Veröffentlicht: (2025)
von: McCammon, James
Veröffentlicht: (2025)
Is Your LLM Outdated? A Deep Look at Temporal Generalization
von: Zhu, Chenghao, et al.
Veröffentlicht: (2024)
von: Zhu, Chenghao, et al.
Veröffentlicht: (2024)
Single Character Perturbations Break LLM Alignment
von: Lin, Leon, et al.
Veröffentlicht: (2024)
von: Lin, Leon, et al.
Veröffentlicht: (2024)
Enhancing LLM Character-Level Manipulation via Divide and Conquer
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
von: Chun, Ye Eun, et al.
Veröffentlicht: (2025)
von: Chun, Ye Eun, et al.
Veröffentlicht: (2025)
AcademicEval: Live Long-Context LLM Benchmark
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
Measuring all the noises of LLM Evals
von: Wang, Sida
Veröffentlicht: (2025)
von: Wang, Sida
Veröffentlicht: (2025)
YESciEval: Robust LLM-as-a-Judge for Scientific Question Answering
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
von: Khan, Haidar, et al.
Veröffentlicht: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
von: Nguyen, Trong-Hieu, et al.
Veröffentlicht: (2024)
Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner
von: Zhou, Cai, et al.
Veröffentlicht: (2025)
von: Zhou, Cai, et al.
Veröffentlicht: (2025)
DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer
von: Hong, Junyuan, et al.
Veröffentlicht: (2023)
von: Hong, Junyuan, et al.
Veröffentlicht: (2023)
Collaborative Quest Completion with LLM-driven Non-Player Characters in Minecraft
von: Rao, Sudha, et al.
Veröffentlicht: (2024)
von: Rao, Sudha, et al.
Veröffentlicht: (2024)
Experiences Build Characters: The Linguistic Origins and Functional Impact of LLM Personality
von: Wang, Xi, et al.
Veröffentlicht: (2026)
von: Wang, Xi, et al.
Veröffentlicht: (2026)
Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
von: Ullrich, Herbert, et al.
Veröffentlicht: (2025)
von: Ullrich, Herbert, et al.
Veröffentlicht: (2025)
Language Models can perform Single-Utterance Self-Correction of Perturbed Reasoning
von: Silver, Sam, et al.
Veröffentlicht: (2025)
von: Silver, Sam, et al.
Veröffentlicht: (2025)
ZeroSumEval: An Extensible Framework For Scaling LLM Evaluation with Inter-Model Competition
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
von: Alyahya, Hisham A., et al.
Veröffentlicht: (2025)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
von: Wu, JiaRu, et al.
Veröffentlicht: (2025)
ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence
von: Wu, Kevin, et al.
Veröffentlicht: (2024)
von: Wu, Kevin, et al.
Veröffentlicht: (2024)
CoreEval: Automatically Building Contamination-Resilient Datasets with Real-World Knowledge toward Reliable LLM Evaluation
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
von: Zhao, Jingqian, et al.
Veröffentlicht: (2025)
Is Your LLM Really Mastering the Concept? A Multi-Agent Benchmark
von: Xu, Shuhang, et al.
Veröffentlicht: (2025)
von: Xu, Shuhang, et al.
Veröffentlicht: (2025)
Deflanderization for Game Dialogue: Balancing Character Authenticity with Task Execution in LLM-based NPCs
von: Buakhaw, Pasin, et al.
Veröffentlicht: (2025)
von: Buakhaw, Pasin, et al.
Veröffentlicht: (2025)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
Who can we trust? LLM-as-a-jury for Comparative Assessment
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
von: Qian, Mengjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024) -
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025) -
Memory Mosaics at scale
von: Zhang, Jianyu, et al.
Veröffentlicht: (2025) -
Make Your LLM Fully Utilize the Context
von: An, Shengnan, et al.
Veröffentlicht: (2024) -
Dual Optimal: Make Your LLM Peer-like with Dignity
von: Wang, Xiangqi, et al.
Veröffentlicht: (2026)