You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Jaewoo, Zhu, Andrew, Callison-Burch, Chris |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
von: Zhu, Andrew, et al.
Veröffentlicht: (2025)
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
von: Liu, Muxin, et al.
Veröffentlicht: (2026)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
von: Horvitz, Zachary, et al.
Veröffentlicht: (2023)
von: Horvitz, Zachary, et al.
Veröffentlicht: (2023)
CALYPSO: LLMs as Dungeon Masters' Assistants
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
von: Zhu, Andrew, et al.
Veröffentlicht: (2023)
Domain Gating Ensemble Networks for AI-Generated Text Detection
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
von: Tripathi, Arihant, et al.
Veröffentlicht: (2025)
Large Language Models Can Self-Improve At Web Agent Tasks
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
von: Patel, Ajay, et al.
Veröffentlicht: (2024)
Concept Lancet: Image Editing with Compositional Representation Transplant
von: Luo, Jinqi, et al.
Veröffentlicht: (2025)
von: Luo, Jinqi, et al.
Veröffentlicht: (2025)
Does Time Have Its Place? Temporal Heads: Where Language Models Recall Time-specific Information
von: Park, Yein, et al.
Veröffentlicht: (2025)
von: Park, Yein, et al.
Veröffentlicht: (2025)
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
von: Ahn, Jaewoo, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models on Bistable Images
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
The Labyrinth and the Thread: Rethinking Regularizations in Sequential Knowledge Editing for Large Language Models
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
von: Wang, Zheng, et al.
Veröffentlicht: (2026)
Asynchronous LLM Function Calling
von: Gim, In, et al.
Veröffentlicht: (2024)
von: Gim, In, et al.
Veröffentlicht: (2024)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Navigate through Enigmatic Labyrinth A Survey of Chain of Thought Reasoning: Advances, Frontiers and Future
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
von: Chu, Zheng, et al.
Veröffentlicht: (2023)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
von: Zhu, Andrew, et al.
Veröffentlicht: (2024)
PaCE: Parsimonious Concept Engineering for Large Language Models
von: Luo, Jinqi, et al.
Veröffentlicht: (2024)
von: Luo, Jinqi, et al.
Veröffentlicht: (2024)
The Labyrinth of Links: Navigating the Associative Maze of Multi-modal LLMs
von: Li, Hong, et al.
Veröffentlicht: (2024)
von: Li, Hong, et al.
Veröffentlicht: (2024)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
von: Kitouni, Ouail, et al.
Veröffentlicht: (2024)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
Linguistic and Argument Diversity in Synthetic Data for Function-Calling Agents
von: Greenstein, Dan, et al.
Veröffentlicht: (2026)
von: Greenstein, Dan, et al.
Veröffentlicht: (2026)
Alopex: A Computational Framework for Enabling On-Device Function Calls with LLMs
von: Ran, Yide, et al.
Veröffentlicht: (2024)
von: Ran, Yide, et al.
Veröffentlicht: (2024)
Enhance Reasoning for Large Language Models in the Game Werewolf
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
von: Wu, Shuang, et al.
Veröffentlicht: (2024)
Green AI: Which Programming Language Consumes the Most?
von: Marini, Niccolò, et al.
Veröffentlicht: (2024)
von: Marini, Niccolò, et al.
Veröffentlicht: (2024)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
von: Yang, Yue, et al.
Veröffentlicht: (2023)
von: Yang, Yue, et al.
Veröffentlicht: (2023)
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
von: Fu, Xingyu, et al.
Veröffentlicht: (2025)
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
von: Abdelaziz, Ibrahim, et al.
Veröffentlicht: (2024)
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
von: Li, Bryan, et al.
Veröffentlicht: (2024)
von: Li, Bryan, et al.
Veröffentlicht: (2024)
Towards Faithful Model Explanation in NLP: A Survey
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
von: Lyu, Qing, et al.
Veröffentlicht: (2022)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
von: Li, Bryan, et al.
Veröffentlicht: (2023)
von: Li, Bryan, et al.
Veröffentlicht: (2023)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
REAMS: Reasoning Enhanced Algorithm for Maths Solving
von: Singh, Eishkaran, et al.
Veröffentlicht: (2025)
von: Singh, Eishkaran, et al.
Veröffentlicht: (2025)
GenesisFunc: Multi-Agent Data Generation for Accurate and Generalizable Function-Calling
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
von: Xu, Hao-Xiang, et al.
Veröffentlicht: (2026)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
von: Chai, Huacan, et al.
Veröffentlicht: (2025)
von: Chai, Huacan, et al.
Veröffentlicht: (2025)
You need to MIMIC to get FAME: Solving Meeting Transcript Scarcity with a Multi-Agent Conversations
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
von: Kirstein, Frederic, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
von: Zhu, Andrew, et al.
Veröffentlicht: (2025) -
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
von: Zhu, Andrew, et al.
Veröffentlicht: (2025) -
Autorubric: Unifying Rubric-based LLM Evaluation
von: Rao, Delip, et al.
Veröffentlicht: (2026) -
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
von: Zhu, Andrew, et al.
Veröffentlicht: (2024) -
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)