CORE: Full-Path Evaluation of LLM Agents Beyond Final State
Fuente:
arXiv
Guardado en:
| Autores principales: | Michelakis, Panagiotis, Hadjiyiannis, Yiannis, Stamoulis, Dimitrios |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows
por: Stamoulis, Dimitrios, et al.
Publicado: (2025)
por: Stamoulis, Dimitrios, et al.
Publicado: (2025)
GeckOpt: LLM System Efficiency via Intent-Based Tool Selection
por: Fore, Michael, et al.
Publicado: (2024)
por: Fore, Michael, et al.
Publicado: (2024)
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
Automated Multi-Agent Workflows for RTL Design
por: Bhattaram, Amulya, et al.
Publicado: (2025)
por: Bhattaram, Amulya, et al.
Publicado: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
por: Kim, Wonjoong, et al.
Publicado: (2025)
por: Kim, Wonjoong, et al.
Publicado: (2025)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
por: Badmus, Emmanuel O., et al.
Publicado: (2025)
por: Badmus, Emmanuel O., et al.
Publicado: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
por: Pandey, Punya Syon, et al.
Publicado: (2025)
por: Pandey, Punya Syon, et al.
Publicado: (2025)
Mining Path Association Rules in Large Property Graphs (with Appendix)
por: Sasaki, Yuya, et al.
Publicado: (2024)
por: Sasaki, Yuya, et al.
Publicado: (2024)
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
por: Cao, Hongliu, et al.
Publicado: (2026)
por: Cao, Hongliu, et al.
Publicado: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
por: Shi, Zhichao, et al.
Publicado: (2025)
por: Shi, Zhichao, et al.
Publicado: (2025)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
por: Li, Yifei, et al.
Publicado: (2026)
por: Li, Yifei, et al.
Publicado: (2026)
LLM Agents Beyond Utility: An Open-Ended Perspective
por: Nachkov, Asen, et al.
Publicado: (2025)
por: Nachkov, Asen, et al.
Publicado: (2025)
CORE: Collaborative Reasoning via Cross Teaching
por: Mishra, Kshitij, et al.
Publicado: (2026)
por: Mishra, Kshitij, et al.
Publicado: (2026)
Evaluating LLM Reasoning Beyond Correctness and CoT
por: Abbasloo, Soheil
Publicado: (2025)
por: Abbasloo, Soheil
Publicado: (2025)
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
por: Wang, Jun, et al.
Publicado: (2025)
por: Wang, Jun, et al.
Publicado: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
por: Chopra, Harshita, et al.
Publicado: (2026)
por: Chopra, Harshita, et al.
Publicado: (2026)
CORE: Comprehensive Ontological Relation Evaluation for Large Language Models
por: Dwivedi, Satyam, et al.
Publicado: (2026)
por: Dwivedi, Satyam, et al.
Publicado: (2026)
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
por: Pasternak, Gil, et al.
Publicado: (2025)
por: Pasternak, Gil, et al.
Publicado: (2025)
Beyond Final Answers: Evaluating Large Language Models for Math Tutoring
por: Gupta, Adit, et al.
Publicado: (2025)
por: Gupta, Adit, et al.
Publicado: (2025)
CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning
por: Nasvytis, Linas, et al.
Publicado: (2026)
por: Nasvytis, Linas, et al.
Publicado: (2026)
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
por: Shaw, Seiji, et al.
Publicado: (2026)
por: Shaw, Seiji, et al.
Publicado: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
por: Wang, Yiding, et al.
Publicado: (2025)
por: Wang, Yiding, et al.
Publicado: (2025)
Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
por: Su, Miao, et al.
Publicado: (2026)
por: Su, Miao, et al.
Publicado: (2026)
Evaluating and Understanding Scheming Propensity in LLM Agents
por: Hopman, Mia, et al.
Publicado: (2026)
por: Hopman, Mia, et al.
Publicado: (2026)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
por: Jain, Suryaansh, et al.
Publicado: (2025)
por: Jain, Suryaansh, et al.
Publicado: (2025)
CORE:Toward Ubiquitous 6G Intelligence Through Collaborative Orchestration of Large Language Model Agents Over Hierarchical Edge
por: Yu, Zitong, et al.
Publicado: (2026)
por: Yu, Zitong, et al.
Publicado: (2026)
Sample-Efficient Reinforcement Learning with Temporal Logic Objectives: Leveraging the Task Specification to Guide Exploration
por: Kantaros, Yiannis, et al.
Publicado: (2024)
por: Kantaros, Yiannis, et al.
Publicado: (2024)
AgentAuditor: Human-Level Safety and Security Evaluation for LLM Agents
por: Luo, Hanjun, et al.
Publicado: (2025)
por: Luo, Hanjun, et al.
Publicado: (2025)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
por: Luo, Mingyu, et al.
Publicado: (2026)
por: Luo, Mingyu, et al.
Publicado: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
por: Kim, Doyoung, et al.
Publicado: (2026)
por: Kim, Doyoung, et al.
Publicado: (2026)
Ego-Foresight: Self-supervised Learning of Agent-Aware Representations for Improved RL
por: Nunes, Manuel Serra, et al.
Publicado: (2024)
por: Nunes, Manuel Serra, et al.
Publicado: (2024)
CORE: Contrastive Masked Feature Reconstruction on Graphs
por: Bo, Jianyuan, et al.
Publicado: (2025)
por: Bo, Jianyuan, et al.
Publicado: (2025)
CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks
por: Meher, Dipak, et al.
Publicado: (2025)
por: Meher, Dipak, et al.
Publicado: (2025)
The Evaluation Game: Beyond Static LLM Benchmarking
por: Wang, Paul, et al.
Publicado: (2026)
por: Wang, Paul, et al.
Publicado: (2026)
Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation
por: Tian, Hanlin, et al.
Publicado: (2024)
por: Tian, Hanlin, et al.
Publicado: (2024)
Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agents
por: Chuang, Yun-Shiuan, et al.
Publicado: (2026)
por: Chuang, Yun-Shiuan, et al.
Publicado: (2026)
Ejemplares similares
-
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
por: Singh, Simranjit, et al.
Publicado: (2024) -
Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows
por: Stamoulis, Dimitrios, et al.
Publicado: (2025) -
GeckOpt: LLM System Efficiency via Intent-Based Tool Selection
por: Fore, Michael, et al.
Publicado: (2024) -
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
por: Singh, Simranjit, et al.
Publicado: (2024) -
Automated Multi-Agent Workflows for RTL Design
por: Bhattaram, Amulya, et al.
Publicado: (2025)