Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Wonjoong, Park, Sangwu, In, Yeonjun, Kim, Sein, Lee, Dongha, Park, Chanyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
di: Kim, Wonjoong, et al.
Pubblicazione: (2026)
di: Kim, Wonjoong, et al.
Pubblicazione: (2026)
Reasoning Structure Matters for Safety Alignment of Reasoning Models
di: In, Yeonjun, et al.
Pubblicazione: (2026)
di: In, Yeonjun, et al.
Pubblicazione: (2026)
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
di: Kim, Wonjoong, et al.
Pubblicazione: (2024)
di: Kim, Wonjoong, et al.
Pubblicazione: (2024)
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
di: In, Yeonjun, et al.
Pubblicazione: (2026)
di: In, Yeonjun, et al.
Pubblicazione: (2026)
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
di: Kim, Sein, et al.
Pubblicazione: (2026)
di: Kim, Sein, et al.
Pubblicazione: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
di: Park, Sangwu, et al.
Pubblicazione: (2026)
di: Park, Sangwu, et al.
Pubblicazione: (2026)
DSLR: Diversity Enhancement and Structure Learning for Rehearsal-based Graph Continual Learning
di: Choi, Seungyoon, et al.
Pubblicazione: (2024)
di: Choi, Seungyoon, et al.
Pubblicazione: (2024)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Rethinking Failure Attribution in Multi-Agent Systems: A Multi-Perspective Benchmark and Evaluation
di: In, Yeonjun, et al.
Pubblicazione: (2026)
di: In, Yeonjun, et al.
Pubblicazione: (2026)
Revisiting Fake News Detection: Towards Temporality-aware Evaluation by Leveraging Engagement Earliness
di: Kim, Junghoon, et al.
Pubblicazione: (2024)
di: Kim, Junghoon, et al.
Pubblicazione: (2024)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
di: Hwang, Yeonjun, et al.
Pubblicazione: (2026)
di: Hwang, Yeonjun, et al.
Pubblicazione: (2026)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
di: Kim, Serin, et al.
Pubblicazione: (2026)
di: Kim, Serin, et al.
Pubblicazione: (2026)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
di: Kim, Jaehoon, et al.
Pubblicazione: (2026)
di: Kim, Jaehoon, et al.
Pubblicazione: (2026)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
di: Yoon, Kanghoon, et al.
Pubblicazione: (2025)
Beyond Ontology in Dialogue State Tracking for Goal-Oriented Chatbot
di: Lee, Sejin, et al.
Pubblicazione: (2024)
di: Lee, Sejin, et al.
Pubblicazione: (2024)
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interactions
di: Lee, Jeongeun, et al.
Pubblicazione: (2026)
di: Lee, Jeongeun, et al.
Pubblicazione: (2026)
Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning
di: Jo, Hwiyeol, et al.
Pubblicazione: (2025)
di: Jo, Hwiyeol, et al.
Pubblicazione: (2025)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
di: Lee, Gisang, et al.
Pubblicazione: (2024)
di: Lee, Gisang, et al.
Pubblicazione: (2024)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
Adaptive Self-training Framework for Fine-grained Scene Graph Generation
di: Kim, Kibum, et al.
Pubblicazione: (2024)
di: Kim, Kibum, et al.
Pubblicazione: (2024)
Zero-shot Commonsense Reasoning over Machine Imagination
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
di: Park, Hyuntae, et al.
Pubblicazione: (2024)
Region4Web: Rethinking Observation Space Granularity for Web Agents
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
Beyond Task-Oriented and Chitchat Dialogues: Proactive and Transition-Aware Conversational Agents
di: Yoon, Yejin, et al.
Pubblicazione: (2025)
di: Yoon, Yejin, et al.
Pubblicazione: (2025)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
di: Kim, Haechan, et al.
Pubblicazione: (2026)
di: Kim, Haechan, et al.
Pubblicazione: (2026)
ProgRAG: Hallucination-Resistant Progressive Retrieval and Reasoning over Knowledge Graphs
di: Park, Minbae, et al.
Pubblicazione: (2025)
di: Park, Minbae, et al.
Pubblicazione: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
InvThink: Premortem Reasoning for Safer Language Models
di: Kim, Yubin, et al.
Pubblicazione: (2025)
di: Kim, Yubin, et al.
Pubblicazione: (2025)
When Is Enough Not Enough? Illusory Completion in Search Agents
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
Open Ko-LLM Leaderboard2: Bridging Foundational and Practical Evaluation for Korean LLMs
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
SAAS: Solving Ability Amplification Strategy for Enhanced Mathematical Reasoning in Large Language Models
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
di: Kim, Hyeonwoo, et al.
Pubblicazione: (2024)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
Disentangling and Generating Modalities for Recommendation in Missing Modality Scenarios
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
COCOA: CBT-based Conversational Counseling Agent using Memory Specialized in Cognitive Distortions and Dynamic Prompt
di: Lee, Suyeon, et al.
Pubblicazione: (2024)
di: Lee, Suyeon, et al.
Pubblicazione: (2024)
QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs
di: Kim, Minsang, et al.
Pubblicazione: (2024)
di: Kim, Minsang, et al.
Pubblicazione: (2024)
PARAN: Persona-Augmented Review ANswering system on Food Delivery Review Dataset
di: Park, Moonsoo, et al.
Pubblicazione: (2025)
di: Park, Moonsoo, et al.
Pubblicazione: (2025)
Ko-MuSR: A Multistep Soft Reasoning Benchmark for LLMs Capable of Understanding Korean
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
di: Kim, Dongjun, et al.
Pubblicazione: (2025)
Open Ko-LLM Leaderboard: Evaluating Large Language Models in Korean with Ko-H5 Benchmark
di: Park, Chanjun, et al.
Pubblicazione: (2024)
di: Park, Chanjun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
di: In, Yeonjun, et al.
Pubblicazione: (2025) -
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
di: Kim, Wonjoong, et al.
Pubblicazione: (2026) -
Reasoning Structure Matters for Safety Alignment of Reasoning Models
di: In, Yeonjun, et al.
Pubblicazione: (2026) -
SIMPLOT: Enhancing Chart Question Answering by Distilling Essentials
di: Kim, Wonjoong, et al.
Pubblicazione: (2024) -
Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agents
di: In, Yeonjun, et al.
Pubblicazione: (2026)