CONDESION-BENCH: Conditional Decision-Making of Large Language Models in Compositional Action Space
Fuente:
arXiv
Salvato in:
| Autori principali: | Hwang, Yeonjun, Park, Sungyong, Kim, Minju, Lee, Dongha, Yeo, Jinyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
di: Kim, Wonjoong, et al.
Pubblicazione: (2025)
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
di: Seo, Yeongbin, et al.
Pubblicazione: (2025)
Ever-Evolving Memory by Blending and Refining the Past
di: Kim, Seo Hyun, et al.
Pubblicazione: (2024)
di: Kim, Seo Hyun, et al.
Pubblicazione: (2024)
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
di: Kim, Hana, et al.
Pubblicazione: (2024)
di: Kim, Hana, et al.
Pubblicazione: (2024)
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
di: Chon, Heejae, et al.
Pubblicazione: (2024)
di: Chon, Heejae, et al.
Pubblicazione: (2024)
Large Language Models are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales
di: Kwon, Taeyoon, et al.
Pubblicazione: (2023)
di: Kwon, Taeyoon, et al.
Pubblicazione: (2023)
COCOA: CBT-based Conversational Counseling Agent using Memory Specialized in Cognitive Distortions and Dynamic Prompt
di: Lee, Suyeon, et al.
Pubblicazione: (2024)
di: Lee, Suyeon, et al.
Pubblicazione: (2024)
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
di: Kim, Sunghwan, et al.
Pubblicazione: (2025)
di: Kim, Sunghwan, et al.
Pubblicazione: (2025)
Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
di: Yoon, Soojin, et al.
Pubblicazione: (2024)
di: Yoon, Soojin, et al.
Pubblicazione: (2024)
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question Answering
di: Ko, Sungho, et al.
Pubblicazione: (2024)
di: Ko, Sungho, et al.
Pubblicazione: (2024)
Evaluating Robustness of Reward Models for Mathematical Reasoning
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
di: Kim, Sunghwan, et al.
Pubblicazione: (2024)
DICE-BENCH: Evaluating the Tool-Use Capabilities of Large Language Models in Multi-Round, Multi-Party Dialogues
di: Jang, Kyochul, et al.
Pubblicazione: (2025)
di: Jang, Kyochul, et al.
Pubblicazione: (2025)
Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics
di: Lee, Seungbeen, et al.
Pubblicazione: (2024)
di: Lee, Seungbeen, et al.
Pubblicazione: (2024)
R1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety Knowledge
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
di: Ong, Kai Tzu-iunn, et al.
Pubblicazione: (2024)
di: Ong, Kai Tzu-iunn, et al.
Pubblicazione: (2024)
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
di: Park, Minjun, et al.
Pubblicazione: (2026)
di: Park, Minjun, et al.
Pubblicazione: (2026)
Leveraging Large Language Models for Active Merchant Non-player Characters
di: Kim, Byungjun, et al.
Pubblicazione: (2024)
di: Kim, Byungjun, et al.
Pubblicazione: (2024)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
di: Prahlad, Deeksha, et al.
Pubblicazione: (2025)
OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Models
di: Kim, Jaehoon, et al.
Pubblicazione: (2026)
di: Kim, Jaehoon, et al.
Pubblicazione: (2026)
LINGO-Space: Language-Conditioned Incremental Grounding for Space
di: Kim, Dohyun, et al.
Pubblicazione: (2024)
di: Kim, Dohyun, et al.
Pubblicazione: (2024)
Evaluating LLMs for Police Decision-Making: A Framework Based on Police Action Scenarios
di: Lee, Sangyub, et al.
Pubblicazione: (2026)
di: Lee, Sangyub, et al.
Pubblicazione: (2026)
Region4Web: Rethinking Observation Space Granularity for Web Agents
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
di: Kwon, Donguk, et al.
Pubblicazione: (2026)
A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
di: Choi, Junhyuk, et al.
Pubblicazione: (2025)
di: Choi, Junhyuk, et al.
Pubblicazione: (2025)
VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
di: Kim, Wonjoong, et al.
Pubblicazione: (2026)
di: Kim, Wonjoong, et al.
Pubblicazione: (2026)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
di: Kim, Serin, et al.
Pubblicazione: (2026)
di: Kim, Serin, et al.
Pubblicazione: (2026)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
di: Kim, Minju, et al.
Pubblicazione: (2025)
di: Kim, Minju, et al.
Pubblicazione: (2025)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
di: Park, Jinyoung, et al.
Pubblicazione: (2023)
di: Park, Jinyoung, et al.
Pubblicazione: (2023)
BIPED: Pedagogically Informed Tutoring System for ESL Education
di: Kwon, Soonwoo, et al.
Pubblicazione: (2024)
di: Kwon, Soonwoo, et al.
Pubblicazione: (2024)
Can Code-Switched Texts Activate a Knowledge Switch in LLMs? A Case Study on English-Korean Code-Switching
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
di: Kim, Seoyeon, et al.
Pubblicazione: (2024)
Unveiling Implicit Table Knowledge with Question-Then-Pinpoint Reasoner for Insightful Table Summarization
di: Seo, Kwangwook, et al.
Pubblicazione: (2024)
di: Seo, Kwangwook, et al.
Pubblicazione: (2024)
Enhancing Decision-Making of Large Language Models via Actor-Critic
di: Dong, Heng, et al.
Pubblicazione: (2025)
di: Dong, Heng, et al.
Pubblicazione: (2025)
On the Decision-Making Abilities in Role-Playing using Large Language Models
di: Shen, Chenglei, et al.
Pubblicazione: (2024)
di: Shen, Chenglei, et al.
Pubblicazione: (2024)
Beyond Ontology in Dialogue State Tracking for Goal-Oriented Chatbot
di: Lee, Sejin, et al.
Pubblicazione: (2024)
di: Lee, Sejin, et al.
Pubblicazione: (2024)
Revisiting Fake News Detection: Towards Temporality-aware Evaluation by Leveraging Engagement Earliness
di: Kim, Junghoon, et al.
Pubblicazione: (2024)
di: Kim, Junghoon, et al.
Pubblicazione: (2024)
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
di: Choi, Dongwook, et al.
Pubblicazione: (2025)
di: Choi, Dongwook, et al.
Pubblicazione: (2025)
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
di: Kim, Namyoung, et al.
Pubblicazione: (2025)
di: Kim, Namyoung, et al.
Pubblicazione: (2025)
MoCoRP: Modeling Consistent Relations between Persona and Response for Persona-based Dialogue
di: Lee, Kyungro, et al.
Pubblicazione: (2025)
di: Lee, Kyungro, et al.
Pubblicazione: (2025)
LLM4SGG: Large Language Models for Weakly Supervised Scene Graph Generation
di: Kim, Kibum, et al.
Pubblicazione: (2023)
di: Kim, Kibum, et al.
Pubblicazione: (2023)
LLM+AL: Bridging Large Language Models and Action Languages for Complex Reasoning about Actions
di: Ishay, Adam, et al.
Pubblicazione: (2025)
di: Ishay, Adam, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
di: Kim, Wonjoong, et al.
Pubblicazione: (2025) -
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
di: Seo, Yeongbin, et al.
Pubblicazione: (2025) -
Ever-Evolving Memory by Blending and Refining the Past
di: Kim, Seo Hyun, et al.
Pubblicazione: (2024) -
Commonsense-augmented Memory Construction and Management in Long-term Conversations via Context-aware Persona Refinement
di: Kim, Hana, et al.
Pubblicazione: (2024) -
Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes
di: Chon, Heejae, et al.
Pubblicazione: (2024)