An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vyas, Kaustubh, Graux, Damien, Montella, Sébastien, Vougiouklis, Pavlos, Lai, Ruofei, Li, Keshuang, Ren, Yang, Pan, Jeff Z. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle
von: Vyas, Kaustubh, et al.
Veröffentlicht: (2024)
von: Vyas, Kaustubh, et al.
Veröffentlicht: (2024)
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
von: Shen, Zhili, et al.
Veröffentlicht: (2024)
von: Shen, Zhili, et al.
Veröffentlicht: (2024)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation
von: Shen, Zhili, et al.
Veröffentlicht: (2024)
von: Shen, Zhili, et al.
Veröffentlicht: (2024)
A Usage-centric Take on Intent Understanding in E-Commerce
von: Zhou, Wendi, et al.
Veröffentlicht: (2024)
von: Zhou, Wendi, et al.
Veröffentlicht: (2024)
Millions of $\text{GeAR}$-s: Extending GraphRAG to Millions of Documents
von: Shen, Zhili, et al.
Veröffentlicht: (2025)
von: Shen, Zhili, et al.
Veröffentlicht: (2025)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
von: Huang, Wenyu, et al.
Veröffentlicht: (2025)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
von: Huang, Wenyu, et al.
Veröffentlicht: (2024)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
OpenSIR: Open-Ended Self-Improving Reasoner
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
von: Kwan, Wai-Chung, et al.
Veröffentlicht: (2025)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
von: Zheng, Danna, et al.
Veröffentlicht: (2024)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
von: Zheng, Danna, et al.
Veröffentlicht: (2025)
von: Zheng, Danna, et al.
Veröffentlicht: (2025)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
von: Stein, Katharina, et al.
Veröffentlicht: (2023)
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
von: Labroo, Arya, et al.
Veröffentlicht: (2026)
von: Labroo, Arya, et al.
Veröffentlicht: (2026)
Can Language Models Analyze Data? Evaluating Large Language Models for Question Answering over Datasets
von: Xenofontos, Andreas, et al.
Veröffentlicht: (2026)
von: Xenofontos, Andreas, et al.
Veröffentlicht: (2026)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
von: Kwon, Deuksin, et al.
Veröffentlicht: (2024)
von: Kwon, Deuksin, et al.
Veröffentlicht: (2024)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
von: Ruan, Kai, et al.
Veröffentlicht: (2024)
von: Ruan, Kai, et al.
Veröffentlicht: (2024)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
von: Dhole, Kaustubh
Veröffentlicht: (2025)
von: Dhole, Kaustubh
Veröffentlicht: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
von: Allaham, Mowafak, et al.
Veröffentlicht: (2024)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
von: Wei, Hui, et al.
Veröffentlicht: (2025)
von: Wei, Hui, et al.
Veröffentlicht: (2025)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
von: Du, Yiming, et al.
Veröffentlicht: (2025)
von: Du, Yiming, et al.
Veröffentlicht: (2025)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
von: Xu, Wanghan, et al.
Veröffentlicht: (2025)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
von: Zhang, Shimao, et al.
Veröffentlicht: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
von: Wang, Lei, et al.
Veröffentlicht: (2024)
von: Wang, Lei, et al.
Veröffentlicht: (2024)
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
von: Sakabe, Ritsu, et al.
Veröffentlicht: (2025)
von: Sakabe, Ritsu, et al.
Veröffentlicht: (2025)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhihuan, et al.
Veröffentlicht: (2024)
Spectral Attention Steering for Prompt Highlighting
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
von: Li, Weixian Waylon, et al.
Veröffentlicht: (2026)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
von: Turk, Matt
Veröffentlicht: (2026)
von: Turk, Matt
Veröffentlicht: (2026)
Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms
von: Macko, Dominik
Veröffentlicht: (2026)
von: Macko, Dominik
Veröffentlicht: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Automated Capability Discovery via Foundation Model Self-Exploration
von: Lu, Cong, et al.
Veröffentlicht: (2025)
von: Lu, Cong, et al.
Veröffentlicht: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
von: Wang, Shu, et al.
Veröffentlicht: (2024)
von: Wang, Shu, et al.
Veröffentlicht: (2024)
Assessing the Capability of LLMs in Solving POSCOMP Questions
von: Viegas, Cayo, et al.
Veröffentlicht: (2025)
von: Viegas, Cayo, et al.
Veröffentlicht: (2025)
DeepInnovator: Triggering the Innovative Capabilities of LLMs
von: Fan, Tianyu, et al.
Veröffentlicht: (2026)
von: Fan, Tianyu, et al.
Veröffentlicht: (2026)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
von: Wang, Hanbin, et al.
Veröffentlicht: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
von: Rystrøm, Jonathan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle
von: Vyas, Kaustubh, et al.
Veröffentlicht: (2024) -
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
von: Shen, Zhili, et al.
Veröffentlicht: (2024) -
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
von: Huang, Wenyu, et al.
Veröffentlicht: (2024) -
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation
von: Shen, Zhili, et al.
Veröffentlicht: (2024) -
A Usage-centric Take on Intent Understanding in E-Commerce
von: Zhou, Wendi, et al.
Veröffentlicht: (2024)