The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Song, Yifan, Wang, Guoyin, Li, Sujian, Lin, Bill Yuchen |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
par: Song, Yifan, et autres
Publié: (2024)
par: Song, Yifan, et autres
Publié: (2024)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
par: Yi, Zihao, et autres
Publié: (2025)
par: Yi, Zihao, et autres
Publié: (2025)
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
par: Li, Wenjie, et autres
Publié: (2026)
par: Li, Wenjie, et autres
Publié: (2026)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
par: Liu, Junpeng, et autres
Publié: (2024)
par: Liu, Junpeng, et autres
Publié: (2024)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
par: Lo, Leo Yu-Ho, et autres
Publié: (2024)
par: Lo, Leo Yu-Ho, et autres
Publié: (2024)
The Good, The Bad, and Why: Unveiling Emotions in Generative AI
par: Li, Cheng, et autres
Publié: (2023)
par: Li, Cheng, et autres
Publié: (2023)
SCALM: Detecting Bad Practices in Smart Contracts Through LLMs
par: Li, Zongwei, et autres
Publié: (2025)
par: Li, Zongwei, et autres
Publié: (2025)
When Bad Data Leads to Good Models
par: Li, Kenneth, et autres
Publié: (2025)
par: Li, Kenneth, et autres
Publié: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
par: Wang, Yicheng, et autres
Publié: (2024)
par: Wang, Yicheng, et autres
Publié: (2024)
Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?
par: Chen, Pinzhen, et autres
Publié: (2024)
par: Chen, Pinzhen, et autres
Publié: (2024)
Automated Evaluation can Distinguish the Good and Bad AI Responses to Patient Questions about Hospitalization
par: Soni, Sarvesh, et autres
Publié: (2025)
par: Soni, Sarvesh, et autres
Publié: (2025)
MPO: Boosting LLM Agents with Meta Plan Optimization
par: Xiong, Weimin, et autres
Publié: (2025)
par: Xiong, Weimin, et autres
Publié: (2025)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
par: Leng, Jixuan, et autres
Publié: (2025)
par: Leng, Jixuan, et autres
Publié: (2025)
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories
par: Song, Yifan, et autres
Publié: (2024)
par: Song, Yifan, et autres
Publié: (2024)
Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
par: Hu, Beizhe, et autres
Publié: (2023)
par: Hu, Beizhe, et autres
Publié: (2023)
FADE: Why Bad Descriptions Happen to Good Features
par: Puri, Bruno, et autres
Publié: (2025)
par: Puri, Bruno, et autres
Publié: (2025)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
par: Sadallah, Abdelrahman, et autres
Publié: (2025)
par: Sadallah, Abdelrahman, et autres
Publié: (2025)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
par: Xu, Zhangchen, et autres
Publié: (2024)
par: Xu, Zhangchen, et autres
Publié: (2024)
Can LLMs Ask Good Questions?
par: Zhang, Yueheng, et autres
Publié: (2025)
par: Zhang, Yueheng, et autres
Publié: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
par: Berezin, Sergei, et autres
Publié: (2025)
par: Berezin, Sergei, et autres
Publié: (2025)
Template-assisted Contrastive Learning of Task-oriented Dialogue Sentence Embeddings
par: Oh, Minsik, et autres
Publié: (2023)
par: Oh, Minsik, et autres
Publié: (2023)
P-Aligner: Enabling Pre-Alignment of Language Models via Principled Instruction Synthesis
par: Song, Feifan, et autres
Publié: (2025)
par: Song, Feifan, et autres
Publié: (2025)
A Training-Free Length Extrapolation Approach for LLMs: Greedy Attention Logit Interpolation (GALI)
par: Li, Yan, et autres
Publié: (2025)
par: Li, Yan, et autres
Publié: (2025)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
par: Zhang, Ran, et autres
Publié: (2024)
par: Zhang, Ran, et autres
Publié: (2024)
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG)
par: Zeng, Shenglai, et autres
Publié: (2024)
par: Zeng, Shenglai, et autres
Publié: (2024)
Small Edits, Big Consequences: Telling Good from Bad Robustness in Large Language Models
par: Ismailov, Altynbek, et autres
Publié: (2025)
par: Ismailov, Altynbek, et autres
Publié: (2025)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
par: Alghisi, Simone, et autres
Publié: (2024)
par: Alghisi, Simone, et autres
Publié: (2024)
Hierarchical Memory Organization for Wikipedia Generation
par: Yu, Eugene J., et autres
Publié: (2025)
par: Yu, Eugene J., et autres
Publié: (2025)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
par: Zhang, Xuemiao, et autres
Publié: (2025)
par: Zhang, Xuemiao, et autres
Publié: (2025)
Simulation, Modelling and Classification of Wiki Contributors: Spotting The Good, The Bad, and The Ugly
par: Méndez, Silvia García, et autres
Publié: (2024)
par: Méndez, Silvia García, et autres
Publié: (2024)
Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement
par: Xiong, Weimin, et autres
Publié: (2024)
par: Xiong, Weimin, et autres
Publié: (2024)
Similarity-based Neighbor Selection for Graph LLMs
par: Li, Rui, et autres
Publié: (2024)
par: Li, Rui, et autres
Publié: (2024)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
par: Kim, Yungi, et autres
Publié: (2024)
par: Kim, Yungi, et autres
Publié: (2024)
EERPD: Leveraging Emotion and Emotion Regulation for Improving Personality Detection
par: Li, Zheng, et autres
Publié: (2024)
par: Li, Zheng, et autres
Publié: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
par: Lin, Bill Yuchen, et autres
Publié: (2024)
par: Lin, Bill Yuchen, et autres
Publié: (2024)
LLMs Should Incorporate Explicit Mechanisms for Human Empathy
par: You, Xiaoxing, et autres
Publié: (2026)
par: You, Xiaoxing, et autres
Publié: (2026)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
par: Mazumder, Aritra, et autres
Publié: (2026)
par: Mazumder, Aritra, et autres
Publié: (2026)
Shapley Value-based Contrastive Alignment for Multimodal Information Extraction
par: Luo, Wen, et autres
Publié: (2024)
par: Luo, Wen, et autres
Publié: (2024)
Can GNN be Good Adapter for LLMs?
par: Huang, Xuanwen, et autres
Publié: (2024)
par: Huang, Xuanwen, et autres
Publié: (2024)
Position: LLMs Can be Good Tutors in English Education
par: Ye, Jingheng, et autres
Publié: (2025)
par: Ye, Jingheng, et autres
Publié: (2025)
Documents similaires
-
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
par: Song, Yifan, et autres
Publié: (2024) -
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
par: Yi, Zihao, et autres
Publié: (2025) -
Identifying Good and Bad Neurons for Task-Level Controllable LLMs
par: Li, Wenjie, et autres
Publié: (2026) -
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
par: Liu, Junpeng, et autres
Publié: (2024) -
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
par: Lo, Leo Yu-Ho, et autres
Publié: (2024)