An Extensive Evaluation of PDDL Capabilities in off-the-shelf LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Vyas, Kaustubh, Graux, Damien, Montella, Sébastien, Vougiouklis, Pavlos, Lai, Ruofei, Li, Keshuang, Ren, Yang, Pan, Jeff Z. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle
por: Vyas, Kaustubh, et al.
Publicado: (2024)
por: Vyas, Kaustubh, et al.
Publicado: (2024)
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
por: Shen, Zhili, et al.
Publicado: (2024)
por: Shen, Zhili, et al.
Publicado: (2024)
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
por: Huang, Wenyu, et al.
Publicado: (2024)
por: Huang, Wenyu, et al.
Publicado: (2024)
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation
por: Shen, Zhili, et al.
Publicado: (2024)
por: Shen, Zhili, et al.
Publicado: (2024)
A Usage-centric Take on Intent Understanding in E-Commerce
por: Zhou, Wendi, et al.
Publicado: (2024)
por: Zhou, Wendi, et al.
Publicado: (2024)
Millions of $\text{GeAR}$-s: Extending GraphRAG to Millions of Documents
por: Shen, Zhili, et al.
Publicado: (2025)
por: Shen, Zhili, et al.
Publicado: (2025)
Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation
por: Huang, Wenyu, et al.
Publicado: (2025)
por: Huang, Wenyu, et al.
Publicado: (2025)
Less is More: Making Smaller Language Models Competent Subgraph Retrievers for Multi-hop KGQA
por: Huang, Wenyu, et al.
Publicado: (2024)
por: Huang, Wenyu, et al.
Publicado: (2024)
PDDL-Mind: Large Language Models are Capable on Belief Reasoning with Reliable State Tracking
por: Zhu, Wang Bill, et al.
Publicado: (2026)
por: Zhu, Wang Bill, et al.
Publicado: (2026)
OpenSIR: Open-Ended Self-Improving Reasoner
por: Kwan, Wai-Chung, et al.
Publicado: (2025)
por: Kwan, Wai-Chung, et al.
Publicado: (2025)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
por: Zheng, Danna, et al.
Publicado: (2024)
por: Zheng, Danna, et al.
Publicado: (2024)
Long-Form Information Alignment Evaluation Beyond Atomic Facts
por: Zheng, Danna, et al.
Publicado: (2025)
por: Zheng, Danna, et al.
Publicado: (2025)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
por: Yu, Simon, et al.
Publicado: (2024)
por: Yu, Simon, et al.
Publicado: (2024)
Automating the Generation of Prompts for LLM-based Action Choice in PDDL Planning
por: Stein, Katharina, et al.
Publicado: (2023)
por: Stein, Katharina, et al.
Publicado: (2023)
Funny or Persuasive, but Not Both: Evaluating Fine-Grained Multi-Concept Control in LLMs
por: Labroo, Arya, et al.
Publicado: (2026)
por: Labroo, Arya, et al.
Publicado: (2026)
Can Language Models Analyze Data? Evaluating Large Language Models for Question Answering over Datasets
por: Xenofontos, Andreas, et al.
Publicado: (2026)
por: Xenofontos, Andreas, et al.
Publicado: (2026)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
por: Kwon, Deuksin, et al.
Publicado: (2024)
por: Kwon, Deuksin, et al.
Publicado: (2024)
Evaluating LLMs' Divergent Thinking Capabilities for Scientific Idea Generation with Minimal Context
por: Ruan, Kai, et al.
Publicado: (2024)
por: Ruan, Kai, et al.
Publicado: (2024)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
por: Dhole, Kaustubh
Publicado: (2025)
por: Dhole, Kaustubh
Publicado: (2025)
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment
por: Allaham, Mowafak, et al.
Publicado: (2024)
por: Allaham, Mowafak, et al.
Publicado: (2024)
PlanGenLLMs: A Modern Survey of LLM Planning Capabilities
por: Wei, Hui, et al.
Publicado: (2025)
por: Wei, Hui, et al.
Publicado: (2025)
Rethinking Memory in LLM based Agents: Representations, Operations, and Emerging Topics
por: Du, Yiming, et al.
Publicado: (2025)
por: Du, Yiming, et al.
Publicado: (2025)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
por: Xu, Wanghan, et al.
Publicado: (2025)
por: Xu, Wanghan, et al.
Publicado: (2025)
How Does Alignment Enhance LLMs' Multilingual Capabilities? A Language Neurons Perspective
por: Zhang, Shimao, et al.
Publicado: (2025)
por: Zhang, Shimao, et al.
Publicado: (2025)
CharacterBox: Evaluating the Role-Playing Capabilities of LLMs in Text-Based Virtual Worlds
por: Wang, Lei, et al.
Publicado: (2024)
por: Wang, Lei, et al.
Publicado: (2024)
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
por: Sakabe, Ritsu, et al.
Publicado: (2025)
por: Sakabe, Ritsu, et al.
Publicado: (2025)
Are Your LLMs Capable of Stable Reasoning?
por: Liu, Junnan, et al.
Publicado: (2024)
por: Liu, Junnan, et al.
Publicado: (2024)
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning
por: Jiang, Zhihuan, et al.
Publicado: (2024)
por: Jiang, Zhihuan, et al.
Publicado: (2024)
Spectral Attention Steering for Prompt Highlighting
por: Li, Weixian Waylon, et al.
Publicado: (2026)
por: Li, Weixian Waylon, et al.
Publicado: (2026)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
por: Turk, Matt
Publicado: (2026)
por: Turk, Matt
Publicado: (2026)
Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms
por: Macko, Dominik
Publicado: (2026)
por: Macko, Dominik
Publicado: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
por: Sirdeshmukh, Ved, et al.
Publicado: (2025)
por: Sirdeshmukh, Ved, et al.
Publicado: (2025)
Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data
por: Liu, Xiao, et al.
Publicado: (2024)
por: Liu, Xiao, et al.
Publicado: (2024)
Automated Capability Discovery via Foundation Model Self-Exploration
por: Lu, Cong, et al.
Publicado: (2025)
por: Lu, Cong, et al.
Publicado: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
por: Wang, Shu, et al.
Publicado: (2024)
por: Wang, Shu, et al.
Publicado: (2024)
Assessing the Capability of LLMs in Solving POSCOMP Questions
por: Viegas, Cayo, et al.
Publicado: (2025)
por: Viegas, Cayo, et al.
Publicado: (2025)
DeepInnovator: Triggering the Innovative Capabilities of LLMs
por: Fan, Tianyu, et al.
Publicado: (2026)
por: Fan, Tianyu, et al.
Publicado: (2026)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
por: Hu, Haiquan, et al.
Publicado: (2025)
por: Hu, Haiquan, et al.
Publicado: (2025)
Code-Vision: Evaluating Multimodal LLMs Logic Understanding and Code Generation Capabilities
por: Wang, Hanbin, et al.
Publicado: (2025)
por: Wang, Hanbin, et al.
Publicado: (2025)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
por: Rystrøm, Jonathan, et al.
Publicado: (2025)
Ejemplares similares
-
From An LLM Swarm To A PDDL-Empowered HIVE: Planning Self-Executed Instructions In A Multi-Modal Jungle
por: Vyas, Kaustubh, et al.
Publicado: (2024) -
Improving Retrieval-augmented Text-to-SQL with AST-based Ranking and Schema Pruning
por: Shen, Zhili, et al.
Publicado: (2024) -
Prompting Large Language Models with Knowledge Graphs for Question Answering Involving Long-tail Facts
por: Huang, Wenyu, et al.
Publicado: (2024) -
GeAR: Graph-enhanced Agent for Retrieval-augmented Generation
por: Shen, Zhili, et al.
Publicado: (2024) -
A Usage-centric Take on Intent Understanding in E-Commerce
por: Zhou, Wendi, et al.
Publicado: (2024)