CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hwang, Yeonjun, Park, Sungyong, Kim, Minju, Lee, Dongha, Yeo, Jinyoung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
von: Kim, Wonjoong, et al.
Veröffentlicht: (2025)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2025)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
Ever-Evolving Memory by Blending and Refining the Past
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2024)
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2024)
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
von: Kim, Hana, et al.
Veröffentlicht: (2024)
von: Kim, Hana, et al.
Veröffentlicht: (2024)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
von: Chon, Heejae, et al.
Veröffentlicht: (2024)
Large Language Models are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
von: Kwon, Taeyoon, et al.
Veröffentlicht: (2023)
von: Kwon, Taeyoon, et al.
Veröffentlicht: (2023)
COCOA: CBT-based Conversational Counseling Agent using Memory Specialized in Cognitive Distortions and Dynamic Prompt
von: Lee, Suyeon, et al.
Veröffentlicht: (2024)
von: Lee, Suyeon, et al.
Veröffentlicht: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2025)
Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
von: Yoon, Soojin, et al.
Veröffentlicht: (2024)
von: Yoon, Soojin, et al.
Veröffentlicht: (2024)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
von: Ko, Sungho, et al.
Veröffentlicht: (2024)
von: Ko, Sungho, et al.
Veröffentlicht: (2024)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
von: Jang, Kyochul, et al.
Veröffentlicht: (2025)
von: Jang, Kyochul, et al.
Veröffentlicht: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
von: Lee, Seungbeen, et al.
Veröffentlicht: (2024)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
von: In, Yeonjun, et al.
Veröffentlicht: (2025)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2024)
von: Ong, Kai Tzu-iunn, et al.
Veröffentlicht: (2024)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
von: Park, Minjun, et al.
Veröffentlicht: (2026)
von: Park, Minjun, et al.
Veröffentlicht: (2026)
Leveraging Large Language Models for Active Merchant Non-player Characters
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
von: Kim, Byungjun, et al.
Veröffentlicht: (2024)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
von: Prahlad, Deeksha, et al.
Veröffentlicht: (2025)
von: Prahlad, Deeksha, et al.
Veröffentlicht: (2025)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
von: Kim, Jaehoon, et al.
Veröffentlicht: (2026)
von: Kim, Jaehoon, et al.
Veröffentlicht: (2026)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
von: Kim, Dohyun, et al.
Veröffentlicht: (2024)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
von: Lee, Sangyub, et al.
Veröffentlicht: (2026)
Region4Web: Rethinking Observation Space Granularity for Web Agents
von: Kwon, Donguk, et al.
Veröffentlicht: (2026)
von: Kwon, Donguk, et al.
Veröffentlicht: (2026)
A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
von: Choi, Junhyuk, et al.
Veröffentlicht: (2025)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
von: Kim, Wonjoong, et al.
Veröffentlicht: (2026)
von: Kim, Wonjoong, et al.
Veröffentlicht: (2026)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
von: Kim, Serin, et al.
Veröffentlicht: (2026)
von: Kim, Serin, et al.
Veröffentlicht: (2026)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
von: Kim, Minju, et al.
Veröffentlicht: (2025)
von: Kim, Minju, et al.
Veröffentlicht: (2025)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
von: Park, Jinyoung, et al.
Veröffentlicht: (2023)
BIPED: Pedagogically Informed Tutoring System for ESL Education
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2024)
von: Kwon, Soonwoo, et al.
Veröffentlicht: (2024)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2024)
Unveiling Implicit Table Knowledge with Question-Then-Pinpoint Reasoner for Insightful Table Summarization
von: Seo, Kwangwook, et al.
Veröffentlicht: (2024)
von: Seo, Kwangwook, et al.
Veröffentlicht: (2024)
Enhancing Decision-Making of Large Language Models via Actor-Critic
von: Dong, Heng, et al.
Veröffentlicht: (2025)
von: Dong, Heng, et al.
Veröffentlicht: (2025)
On the Decision-Making Abilities in Role-Playing using Large Language Models
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
von: Shen, Chenglei, et al.
Veröffentlicht: (2024)
Beyond Ontology in Dialogue State Tracking for Goal-Oriented Chatbot
von: Lee, Sejin, et al.
Veröffentlicht: (2024)
von: Lee, Sejin, et al.
Veröffentlicht: (2024)
Revisiting Fake News Detection: Towards Temporality-aware Evaluation by Leveraging Engagement Earliness
von: Kim, Junghoon, et al.
Veröffentlicht: (2024)
von: Kim, Junghoon, et al.
Veröffentlicht: (2024)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
von: Choi, Dongwook, et al.
Veröffentlicht: (2025)
von: Choi, Dongwook, et al.
Veröffentlicht: (2025)
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
von: Kim, Namyoung, et al.
Veröffentlicht: (2025)
von: Kim, Namyoung, et al.
Veröffentlicht: (2025)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
von: Lee, Kyungro, et al.
Veröffentlicht: (2025)
von: Lee, Kyungro, et al.
Veröffentlicht: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
von: Kim, Kibum, et al.
Veröffentlicht: (2023)
von: Kim, Kibum, et al.
Veröffentlicht: (2023)
LLM+AL: Bridging Large Language Models and Action Languages for Complex Reasoning about Actions
von: Ishay, Adam, et al.
Veröffentlicht: (2025)
von: Ishay, Adam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
von: Kim, Wonjoong, et al.
Veröffentlicht: (2025) -
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025) -
Ever-Evolving Memory by Blending and Refining the Past
von: Kim, Seo Hyun, et al.
Veröffentlicht: (2024) -
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
von: Kim, Hana, et al.
Veröffentlicht: (2024) -
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
von: Chon, Heejae, et al.
Veröffentlicht: (2024)