Saved in:
| Main Authors: | Michelakis, Panagiotis, Hadjiyiannis, Yiannis, Stamoulis, Dimitrios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.20998 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows
by: Stamoulis, Dimitrios, et al.
Published: (2025)
by: Stamoulis, Dimitrios, et al.
Published: (2025)
GeckOpt: LLM System Efficiency via Intent-Based Tool Selection
by: Fore, Michael, et al.
Published: (2024)
by: Fore, Michael, et al.
Published: (2024)
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
Automated Multi-Agent Workflows for RTL Design
by: Bhattaram, Amulya, et al.
Published: (2025)
by: Bhattaram, Amulya, et al.
Published: (2025)
An LLM-Tool Compiler for Fused Parallel Function Calling
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses
by: Badmus, Emmanuel O., et al.
Published: (2025)
by: Badmus, Emmanuel O., et al.
Published: (2025)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Mining Path Association Rules in Large Property Graphs (with Appendix)
by: Sasaki, Yuya, et al.
Published: (2024)
by: Sasaki, Yuya, et al.
Published: (2024)
CoFineLLM: Conformal Finetuning of LLMs for Language-Instructed Robot Planning
by: Wang, Jun, et al.
Published: (2025)
by: Wang, Jun, et al.
Published: (2025)
CORE: Comprehensive Ontological Relation Evaluation for Large Language Models
by: Dwivedi, Satyam, et al.
Published: (2026)
by: Dwivedi, Satyam, et al.
Published: (2026)
CORE: Collaborative Reasoning via Cross Teaching
by: Mishra, Kshitij, et al.
Published: (2026)
by: Mishra, Kshitij, et al.
Published: (2026)
JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
by: Shi, Zhichao, et al.
Published: (2025)
by: Shi, Zhichao, et al.
Published: (2025)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
by: Cao, Hongliu, et al.
Published: (2026)
by: Cao, Hongliu, et al.
Published: (2026)
Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation
by: Tian, Hanlin, et al.
Published: (2024)
by: Tian, Hanlin, et al.
Published: (2024)
GeoFlow: Agentic Workflow Automation for Geospatial Tasks
by: Bhattaram, Amulya, et al.
Published: (2025)
by: Bhattaram, Amulya, et al.
Published: (2025)
Beyond Final Answers: Evaluating Large Language Models for Math Tutoring
by: Gupta, Adit, et al.
Published: (2025)
by: Gupta, Adit, et al.
Published: (2025)
LLM Agents Beyond Utility: An Open-Ended Perspective
by: Nachkov, Asen, et al.
Published: (2025)
by: Nachkov, Asen, et al.
Published: (2025)
CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning
by: Nasvytis, Linas, et al.
Published: (2026)
by: Nasvytis, Linas, et al.
Published: (2026)
Multimodal and Multiview Deep Fusion for Autonomous Marine Navigation
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
by: Dagdilelis, Dimitrios, et al.
Published: (2025)
Evaluating LLM Reasoning Beyond Correctness and CoT
by: Abbasloo, Soheil
Published: (2025)
by: Abbasloo, Soheil
Published: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
Sample-Efficient Reinforcement Learning with Temporal Logic Objectives: Leveraging the Task Specification to Guide Exploration
by: Kantaros, Yiannis, et al.
Published: (2024)
by: Kantaros, Yiannis, et al.
Published: (2024)
Ego-Foresight: Self-supervised Learning of Agent-Aware Representations for Improved RL
by: Nunes, Manuel Serra, et al.
Published: (2024)
by: Nunes, Manuel Serra, et al.
Published: (2024)
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
by: Shaw, Seiji, et al.
Published: (2026)
by: Shaw, Seiji, et al.
Published: (2026)
CORE: Contrastive Masked Feature Reconstruction on Graphs
by: Bo, Jianyuan, et al.
Published: (2025)
by: Bo, Jianyuan, et al.
Published: (2025)
CORE-KG: An LLM-Driven Knowledge Graph Construction Framework for Human Smuggling Networks
by: Meher, Dipak, et al.
Published: (2025)
by: Meher, Dipak, et al.
Published: (2025)
CORE -- A Cell-Level Coarse-to-Fine Image Registration Engine for Multi-stain Image Alignment
by: Nasir, Esha Sadia, et al.
Published: (2025)
by: Nasir, Esha Sadia, et al.
Published: (2025)
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
by: Pasternak, Gil, et al.
Published: (2025)
by: Pasternak, Gil, et al.
Published: (2025)
CORE:Toward Ubiquitous 6G Intelligence Through Collaborative Orchestration of Large Language Model Agents Over Hierarchical Edge
by: Yu, Zitong, et al.
Published: (2026)
by: Yu, Zitong, et al.
Published: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
by: Zheng, Junhao, et al.
Published: (2025)
by: Zheng, Junhao, et al.
Published: (2025)
State Representations as Incentives for Reinforcement Learning Agents: A Sim2Real Analysis on Robotic Grasping
by: Petropoulakis, Panagiotis, et al.
Published: (2023)
by: Petropoulakis, Panagiotis, et al.
Published: (2023)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
by: Siegel, Zachary S., et al.
Published: (2024)
by: Siegel, Zachary S., et al.
Published: (2024)
When Alignment Isn't Enough: Response-Path Attacks on LLM Agents
by: Luo, Mingyu, et al.
Published: (2026)
by: Luo, Mingyu, et al.
Published: (2026)
Beyond Outcome Reward: Decoupling Search and Answering Improves LLM Agents
by: Wang, Yiding, et al.
Published: (2025)
by: Wang, Yiding, et al.
Published: (2025)
Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents
by: Su, Miao, et al.
Published: (2026)
by: Su, Miao, et al.
Published: (2026)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
by: Kim, Doyoung, et al.
Published: (2026)
by: Kim, Doyoung, et al.
Published: (2026)
Similar Items
-
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
by: Singh, Simranjit, et al.
Published: (2024) -
Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows
by: Stamoulis, Dimitrios, et al.
Published: (2025) -
GeckOpt: LLM System Efficiency via Intent-Based Tool Selection
by: Fore, Michael, et al.
Published: (2024) -
GeoLLM-Engine: A Realistic Environment for Building Geospatial Copilots
by: Singh, Simranjit, et al.
Published: (2024) -
Automated Multi-Agent Workflows for RTL Design
by: Bhattaram, Amulya, et al.
Published: (2025)