Evaluating LLM Metrics Through Real-World Capabilities
Fuente:
arXiv
Guardado en:
| Autores principales: | Miller, Justin K, Tang, Wenjia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
por: Wang, Liang, et al.
Publicado: (2026)
por: Wang, Liang, et al.
Publicado: (2026)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
por: Ivanov, Igor
Publicado: (2025)
por: Ivanov, Igor
Publicado: (2025)
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
por: Davide, Fabrizio, et al.
Publicado: (2024)
por: Davide, Fabrizio, et al.
Publicado: (2024)
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
por: Wu, Dekun, et al.
Publicado: (2023)
por: Wu, Dekun, et al.
Publicado: (2023)
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
por: Redkar, Neel
Publicado: (2024)
por: Redkar, Neel
Publicado: (2024)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
por: Chang, Edward Y., et al.
Publicado: (2025)
por: Chang, Edward Y., et al.
Publicado: (2025)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
por: Danilevsky, Marina, et al.
Publicado: (2025)
por: Danilevsky, Marina, et al.
Publicado: (2025)
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
por: Chang, Edward Y.
Publicado: (2026)
por: Chang, Edward Y.
Publicado: (2026)
Toward Architecture-Aware Evaluation Metrics for LLM Agents
por: Souza, Débora, et al.
Publicado: (2026)
por: Souza, Débora, et al.
Publicado: (2026)
A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
por: Christop, Iwona, et al.
Publicado: (2026)
por: Christop, Iwona, et al.
Publicado: (2026)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
Evaluating Relational Reasoning in LLMs with REL
por: Fesser, Lukas, et al.
Publicado: (2026)
por: Fesser, Lukas, et al.
Publicado: (2026)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
por: Chang, Edward Y., et al.
Publicado: (2025)
por: Chang, Edward Y., et al.
Publicado: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
por: Lai, Junyu, et al.
Publicado: (2025)
por: Lai, Junyu, et al.
Publicado: (2025)
EVINCE: Optimizing Multi-LLM Dialogues Using Conditional Statistics and Information Theory
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Active Context Compression: Autonomous Memory Management in LLM Agents
por: Verma, Nikhil
Publicado: (2026)
por: Verma, Nikhil
Publicado: (2026)
Evaluating Steering Techniques using Human Similarity Judgments
por: Studdiford, Zach, et al.
Publicado: (2025)
por: Studdiford, Zach, et al.
Publicado: (2025)
Intrinsic Evaluation of RAG Systems for Deep-Logic Questions
por: Hu, Junyi, et al.
Publicado: (2024)
por: Hu, Junyi, et al.
Publicado: (2024)
KNOW: A Real-World Ontology for Knowledge Capture with Large Language Models
por: Bendiken, Arto
Publicado: (2024)
por: Bendiken, Arto
Publicado: (2024)
Applying Cognitive Design Patterns to General LLM Agents
por: Wray, Robert E., et al.
Publicado: (2025)
por: Wray, Robert E., et al.
Publicado: (2025)
Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
por: Paulsen, Norman
Publicado: (2025)
por: Paulsen, Norman
Publicado: (2025)
DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
por: Hasan, Md Hasebul, et al.
Publicado: (2026)
por: Hasan, Md Hasebul, et al.
Publicado: (2026)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
From Fake Focus to Real Precision: Confusion-Driven Adversarial Attention Learning in Transformers
por: Liu, Yawei
Publicado: (2025)
por: Liu, Yawei
Publicado: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
por: Preuveneers, Jack, et al.
Publicado: (2025)
por: Preuveneers, Jack, et al.
Publicado: (2025)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
por: Sun, Kangkang, et al.
Publicado: (2026)
por: Sun, Kangkang, et al.
Publicado: (2026)
Efficient LLM Safety Evaluation through Multi-Agent Debate
por: Lin, Dachuan, et al.
Publicado: (2025)
por: Lin, Dachuan, et al.
Publicado: (2025)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Evaluating Large Language Models on Historical Health Crisis Knowledge in Resource-Limited Settings: A Hybrid Multi-Metric Study
por: Hasan, Mohammed Rakibul
Publicado: (2026)
por: Hasan, Mohammed Rakibul
Publicado: (2026)
Generative Active Testing: Efficient LLM Evaluation via Proxy Task Adaptation
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2026)
por: Ramakrishnan, Aashish Anantha, et al.
Publicado: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs, et al.
Publicado: (2023)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
por: Abramov, Roman, et al.
Publicado: (2025)
por: Abramov, Roman, et al.
Publicado: (2025)
Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons
por: Sandan, Isik Baran, et al.
Publicado: (2025)
por: Sandan, Isik Baran, et al.
Publicado: (2025)
LLM-Based SQL Generation: Prompting, Self-Refinement, and Adaptive Weighted Majority Voting
por: Yang, Yu-Jie, et al.
Publicado: (2026)
por: Yang, Yu-Jie, et al.
Publicado: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
por: Vera, Patricio
Publicado: (2026)
por: Vera, Patricio
Publicado: (2026)
Ejemplares similares
-
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
por: Wang, Liang, et al.
Publicado: (2026) -
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
por: Ivanov, Igor
Publicado: (2025) -
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
por: Davide, Fabrizio, et al.
Publicado: (2024) -
Deciphering Digital Detectives: Understanding LLM Behaviors and Capabilities in Multi-Agent Mystery Games
por: Wu, Dekun, et al.
Publicado: (2023) -
Planning vs Reasoning: Ablations to Test Capabilities of LoRA layers
por: Redkar, Neel
Publicado: (2024)