The Generative Energy Arena (GEA): Incorporating Energy Awareness in Large Language Model (LLM) Human Evaluations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arriaga, Carlos, Martínez, Gonzalo, Sendin, Eneko, Conde, Javier, Reviriego, Pedro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
von: Ferrando, Raquel, et al.
Veröffentlicht: (2025)
von: Ferrando, Raquel, et al.
Veröffentlicht: (2025)
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
von: Fu, Tairan, et al.
Veröffentlicht: (2026)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
von: Awad, Samer, et al.
Veröffentlicht: (2026)
von: Awad, Samer, et al.
Veröffentlicht: (2026)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
To Words and Beyond: Probing Large Language Models for Sentence-Level Psycholinguistic Norms of Memorability and Reading Times
von: Clark, Thomas Hikaru, et al.
Veröffentlicht: (2026)
von: Clark, Thomas Hikaru, et al.
Veröffentlicht: (2026)
Can ChatGPT Learn to Count Letters?
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
von: Reviriego, Pedro, et al.
Veröffentlicht: (2023)
von: Reviriego, Pedro, et al.
Veröffentlicht: (2023)
Concurrent Linguistic Error Detection (CLED): a New Methodology for Error Detection in Large Language Models
von: Zhu, Jinhua, et al.
Veröffentlicht: (2024)
von: Zhu, Jinhua, et al.
Veröffentlicht: (2024)
How does fine-tuning improve sensorimotor representations in large language models?
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
Why Do Large Language Models (LLMs) Struggle to Count Letters?
von: Fu, Tairan, et al.
Veröffentlicht: (2024)
von: Fu, Tairan, et al.
Veröffentlicht: (2024)
Spanish and LLM Benchmarks: is MMLU Lost in Translation?
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
von: Plaza, Irene, et al.
Veröffentlicht: (2024)
Elsevier Arena: Human Evaluation of Chemistry/Biology/Health Foundational Large Language Models
von: Thorne, Camilo, et al.
Veröffentlicht: (2024)
von: Thorne, Camilo, et al.
Veröffentlicht: (2024)
MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
von: Tang, Jianheng, et al.
Veröffentlicht: (2024)
von: Tang, Jianheng, et al.
Veröffentlicht: (2024)
Spilled Energy in Large Language Models
von: Minut, Adrian Robert, et al.
Veröffentlicht: (2026)
von: Minut, Adrian Robert, et al.
Veröffentlicht: (2026)
CultureLLM: Incorporating Cultural Differences into Large Language Models
von: Li, Cheng, et al.
Veröffentlicht: (2024)
von: Li, Cheng, et al.
Veröffentlicht: (2024)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
Evaluating Creative Short Story Generation in Humans and Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Arena-Lite: Efficient and Reliable Large Language Model Evaluation via Tournament-Based Direct Comparisons
von: Son, Seonil, et al.
Veröffentlicht: (2024)
von: Son, Seonil, et al.
Veröffentlicht: (2024)
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation
von: Ke, Pei, et al.
Veröffentlicht: (2023)
von: Ke, Pei, et al.
Veröffentlicht: (2023)
Stochastic Streets: A Walk Through Random LLM Address Generation in four European Cities
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
von: Fu, Tairan, et al.
Veröffentlicht: (2025)
Establishing Vocabulary Tests as a Benchmark for Evaluating Large Language Models
von: Martínez, Gonzalo, et al.
Veröffentlicht: (2023)
von: Martínez, Gonzalo, et al.
Veröffentlicht: (2023)
Understanding the Impact of Artificial Intelligence in Academic Writing: Metadata to the Rescue
von: Conde, Javier, et al.
Veröffentlicht: (2025)
von: Conde, Javier, et al.
Veröffentlicht: (2025)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
von: Chiang, Wei-Lin, et al.
Veröffentlicht: (2024)
Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use
von: Arriaga, Carlos, et al.
Veröffentlicht: (2024)
von: Arriaga, Carlos, et al.
Veröffentlicht: (2024)
LongReasonArena: A Long Reasoning Benchmark for Large Language Models
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
von: Ding, Jiayu, et al.
Veröffentlicht: (2025)
GameArena: Evaluating LLM Reasoning through Live Computer Games
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
von: Hu, Lanxiang, et al.
Veröffentlicht: (2024)
Werewolf Arena: A Case Study in LLM Evaluation via Social Deduction
von: Bailis, Suma, et al.
Veröffentlicht: (2024)
von: Bailis, Suma, et al.
Veröffentlicht: (2024)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
von: Han, Hojae, et al.
Veröffentlicht: (2024)
von: Han, Hojae, et al.
Veröffentlicht: (2024)
Beware of Words: Evaluating the Lexical Diversity of Conversational LLMs using ChatGPT as Case Study
von: Martínez, Gonzalo, et al.
Veröffentlicht: (2024)
von: Martínez, Gonzalo, et al.
Veröffentlicht: (2024)
Energy-Efficient Stochastic Computing (SC) Neural Networks for Internet of Things Devices With Layer-Wise Adjustable Sequence Length (ASL)
von: Wang, Ziheng, et al.
Veröffentlicht: (2025)
von: Wang, Ziheng, et al.
Veröffentlicht: (2025)
DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
von: Hu, Wenhao, et al.
Veröffentlicht: (2025)
Evaluating Morphological Compositional Generalization in Large Language Models
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
von: Ismayilzada, Mete, et al.
Veröffentlicht: (2024)
Energy Landscapes Enable Reliable Abstention in Retrieval-Augmented Large Language Models for Healthcare
von: Shankar, Ravi, et al.
Veröffentlicht: (2025)
von: Shankar, Ravi, et al.
Veröffentlicht: (2025)
Evaluating Neural Language Models as Cognitive Models of Language Acquisition
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
von: Martínez, Héctor Javier Vázquez, et al.
Veröffentlicht: (2023)
Automatic Question & Answer Generation Using Generative Large Language Model (LLM)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
von: Ehsan, Md. Alvee, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Is There a Case for Conversation Optimized Tokenizers in Large Language Models?
von: Ferrando, Raquel, et al.
Veröffentlicht: (2025) -
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident, Especially When They are Wrong
von: Fu, Tairan, et al.
Veröffentlicht: (2025) -
Psycholinguistic Word Features: a New Approach for the Evaluation of LLMs Alignment with Humans
von: Conde, Javier, et al.
Veröffentlicht: (2025) -
Beyond Reproducibility: Token Probabilities Expose Large Language Model Nondeterminism
von: Fu, Tairan, et al.
Veröffentlicht: (2026) -
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
von: Awad, Samer, et al.
Veröffentlicht: (2026)