ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Kwak, Beong-woo, Kim, Minju, Lim, Dongha, Chae, Hyungjoo, Kang, Dongjin, Kim, Sunghwan, Yang, Dongil, Yeo, Jinyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
by: Kim, Sunghwan, et al.
Published: (2025)
by: Kim, Sunghwan, et al.
Published: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
by: Yang, Dongil, et al.
Published: (2025)
by: Yang, Dongil, et al.
Published: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
by: Lee, Seungbeen, et al.
Published: (2024)
by: Lee, Seungbeen, et al.
Published: (2024)
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
by: Chae, Hyungjoo, et al.
Published: (2025)
by: Chae, Hyungjoo, et al.
Published: (2025)
Pearl: A Review-driven Persona-Knowledge Grounded Conversational Recommendation Dataset
by: Kim, Minjin, et al.
Published: (2024)
by: Kim, Minjin, et al.
Published: (2024)
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
by: Kim, Minju, et al.
Published: (2025)
by: Kim, Minju, et al.
Published: (2025)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
by: Ko, Sungho, et al.
Published: (2024)
by: Ko, Sungho, et al.
Published: (2024)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
LEGO-Eval: Towards Fine-Grained Evaluation on Synthesizing 3D Embodied Environments with Tool Augmentation
by: Hwangbo, Gyeom, et al.
Published: (2025)
by: Hwangbo, Gyeom, et al.
Published: (2025)
Embodied Agents Meet Personalization: Investigating Challenges and Solutions Through the Lens of Memory Utilization
by: Kwon, Taeyoon, et al.
Published: (2025)
by: Kwon, Taeyoon, et al.
Published: (2025)
Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language Models
by: Chae, Hyungjoo, et al.
Published: (2024)
by: Chae, Hyungjoo, et al.
Published: (2024)
Cactus: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
by: Lee, Suyeon, et al.
Published: (2024)
by: Lee, Suyeon, et al.
Published: (2024)
Stop Playing the Guessing Game! Target-free User Simulation for Evaluating Conversational Recommender Systems
by: Kim, Sunghwan, et al.
Published: (2024)
by: Kim, Sunghwan, et al.
Published: (2024)
Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support Conversation
by: Kang, Dongjin, et al.
Published: (2024)
by: Kang, Dongjin, et al.
Published: (2024)
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
by: Hwang, Yeonjun, et al.
Published: (2026)
by: Hwang, Yeonjun, et al.
Published: (2026)
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
YA-TA: Towards Personalized Question-Answering Teaching Assistants using Instructor-Student Dual Retrieval-augmented Knowledge Fusion
by: Yang, Dongil, et al.
Published: (2024)
by: Yang, Dongil, et al.
Published: (2024)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
by: Choi, Dongwook, et al.
Published: (2025)
by: Choi, Dongwook, et al.
Published: (2025)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
by: Kim, Hana, et al.
Published: (2024)
by: Kim, Hana, et al.
Published: (2024)
Coffee: Boost Your Code LLMs by Fixing Bugs with Feedback
by: Moon, Seungjun, et al.
Published: (2023)
by: Moon, Seungjun, et al.
Published: (2023)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
by: Kim, Seoyeon, et al.
Published: (2024)
by: Kim, Seoyeon, et al.
Published: (2024)
MVIGER: Multi-View Variational Integration of Complementary Knowledge for Generative Recommender
by: Kim, Tongyoung, et al.
Published: (2024)
by: Kim, Tongyoung, et al.
Published: (2024)
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
by: Kim, Wonjoong, et al.
Published: (2025)
by: Kim, Wonjoong, et al.
Published: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
by: Seo, Yeongbin, et al.
Published: (2024)
by: Seo, Yeongbin, et al.
Published: (2024)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
by: Seo, Yeongbin, et al.
Published: (2025)
by: Seo, Yeongbin, et al.
Published: (2025)
Unveiling Implicit Table Knowledge with Question-Then-Pinpoint Reasoner for Insightful Table Summarization
by: Seo, Kwangwook, et al.
Published: (2024)
by: Seo, Kwangwook, et al.
Published: (2024)
Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
by: Yoon, Soojin, et al.
Published: (2024)
by: Yoon, Soojin, et al.
Published: (2024)
COCOA: CBT-based Conversational Counseling Agent using Memory Specialized in Cognitive Distortions and Dynamic Prompt
by: Lee, Suyeon, et al.
Published: (2024)
by: Lee, Suyeon, et al.
Published: (2024)
Self-Consistent Reasoning-based Aspect-Sentiment Quad Prediction with Extract-Then-Assign Strategy
by: Kim, Jieyong, et al.
Published: (2024)
by: Kim, Jieyong, et al.
Published: (2024)
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
by: Kim, Jieyong, et al.
Published: (2024)
by: Kim, Jieyong, et al.
Published: (2024)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
by: Park, Minjun, et al.
Published: (2026)
by: Park, Minjun, et al.
Published: (2026)
PaP-NF: Probabilistic Long-Term Time Series Forecasting via Prefix-as-Prompt Reprogramming and Normalizing Flows
by: Kim, Minju, et al.
Published: (2026)
by: Kim, Minju, et al.
Published: (2026)
Similar Items
-
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
by: Kim, Sunghwan, et al.
Published: (2025) -
Evaluating Robustness of Reward Models for Mathematical Reasoning
by: Kim, Sunghwan, et al.
Published: (2024) -
One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL
by: Chae, Hyungjoo, et al.
Published: (2025) -
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study
by: Yang, Dongil, et al.
Published: (2025) -
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
by: Lee, Seungbeen, et al.
Published: (2024)