FABLE: A Novel Data-Flow Analysis Benchmark on Procedural Text for Large Language Model Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Pallagani, Vishal, Gupta, Nitin, Aydin, John, Srivastava, Biplav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
On Sample-Efficient Generalized Planning via Learned Transition Models
by: Gupta, Nitin, et al.
Published: (2026)
by: Gupta, Nitin, et al.
Published: (2026)
The Case for Developing a Foundation Model for Planning-like Tasks from Scratch
by: Srivastava, Biplav, et al.
Published: (2024)
by: Srivastava, Biplav, et al.
Published: (2024)
A Neurosymbolic Fast and Slow Architecture for Graph Coloring
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
A Planning Ontology to Represent and Exploit Planning Knowledge for Performance Efficiency
by: Muppasani, Bharath, et al.
Published: (2023)
by: Muppasani, Bharath, et al.
Published: (2023)
On the Prospects of Incorporating Large Language Models (LLMs) in Automated Planning and Scheduling (APS)
by: Pallagani, Vishal, et al.
Published: (2024)
by: Pallagani, Vishal, et al.
Published: (2024)
On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
by: Widener, Michael, et al.
Published: (2025)
by: Widener, Michael, et al.
Published: (2025)
A Novel Approach to Balance Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes and its Implementation in BEACON
by: Nagpal, Vansh, et al.
Published: (2024)
by: Nagpal, Vansh, et al.
Published: (2024)
Neurosymbolic AI for Enhancing Instructability in Generative AI
by: Sheth, Amit, et al.
Published: (2024)
by: Sheth, Amit, et al.
Published: (2024)
Fundamental Limits of Black-Box Safety Evaluation: Information-Theoretic and Computational Barriers from Latent Context Conditioning
by: Srivastava, Vishal
Published: (2026)
by: Srivastava, Vishal
Published: (2026)
Chatsparent: An Interactive System for Detecting and Mitigating Cognitive Fatigue in LLMs
by: Marwah, Riju, et al.
Published: (2025)
by: Marwah, Riju, et al.
Published: (2025)
Holistic Explainable AI (H-XAI): Extending Transparency Beyond Developers in AI-Driven Decision Making
by: Lakkaraju, Kausik, et al.
Published: (2025)
by: Lakkaraju, Kausik, et al.
Published: (2025)
Technology-assisted Personalized Yoga for Better Health -- Challenges and Outlook
by: Kumar, Vivek, et al.
Published: (2025)
by: Kumar, Vivek, et al.
Published: (2025)
Enterprise Large Language Model Evaluation Benchmark
by: Wang, Liya, et al.
Published: (2025)
by: Wang, Liya, et al.
Published: (2025)
Towards Information-Optimized Multi-Agent Path Finding: A Hybrid Framework with Reduced Inter-Agent Information Sharing
by: Muppasani, Bharath, et al.
Published: (2025)
by: Muppasani, Bharath, et al.
Published: (2025)
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
by: Li, Zhuoqun, et al.
Published: (2024)
by: Li, Zhuoqun, et al.
Published: (2024)
Human Evaluation of Procedural Knowledge Graph Extraction from Text with Large Language Models
by: Carriero, Valentina Anita, et al.
Published: (2024)
by: Carriero, Valentina Anita, et al.
Published: (2024)
BEACON: Balancing Convenience and Nutrition in Meals With Long-Term Group Recommendations and Reasoning on Multimodal Recipes
by: Nagpal, Vansh, et al.
Published: (2024)
by: Nagpal, Vansh, et al.
Published: (2024)
A Comprehensive Evaluation of Large Language Models on Benchmark Biomedical Text Processing Tasks
by: Jahan, Israt, et al.
Published: (2023)
by: Jahan, Israt, et al.
Published: (2023)
Quantifying Automation Risk in High-Automation AI Systems: A Bayesian Framework for Failure Propagation and Optimal Oversight
by: Srivastava, Vishal, et al.
Published: (2026)
by: Srivastava, Vishal, et al.
Published: (2026)
GAICo: A Deployed and Extensible Framework for Evaluating Diverse and Multimodal Generative AI Outputs
by: Gupta, Nitin, et al.
Published: (2025)
by: Gupta, Nitin, et al.
Published: (2025)
MILPaC: A Novel Benchmark for Evaluating Translation of Legal Text to Indian Languages
by: Mahapatra, Sayan, et al.
Published: (2023)
by: Mahapatra, Sayan, et al.
Published: (2023)
SpatialText: A Pure-Text Cognitive Benchmark for Spatial Understanding in Large Language Models
by: Jiang, Peiyao, et al.
Published: (2026)
by: Jiang, Peiyao, et al.
Published: (2026)
Text2GQL-Bench: A Text to Graph Query Language Benchmark [Experiment, Analysis & Benchmark]
by: Lyu, Songlin, et al.
Published: (2026)
by: Lyu, Songlin, et al.
Published: (2026)
Towards Effective Planning Strategies for Dynamic Opinion Networks
by: Muppasani, Bharath, et al.
Published: (2024)
by: Muppasani, Bharath, et al.
Published: (2024)
Promoting Research Collaboration with Open Data Driven Team Recommendation in Response to Call for Proposals
by: Valluru, Siva Likitha, et al.
Published: (2023)
by: Valluru, Siva Likitha, et al.
Published: (2023)
A Technical Survey of Reinforcement Learning Techniques for Large Language Models
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
by: Srivastava, Saksham Sahai, et al.
Published: (2025)
FinDABench: Benchmarking Financial Data Analysis Ability of Large Language Models
by: Liu, Shu, et al.
Published: (2024)
by: Liu, Shu, et al.
Published: (2024)
MEMO-Bench: A Multiple Benchmark for Text-to-Image and Multimodal Large Language Models on Human Emotion Analysis
by: Zhou, Yingjie, et al.
Published: (2024)
by: Zhou, Yingjie, et al.
Published: (2024)
MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
by: Srivastava, Varun, et al.
Published: (2025)
by: Srivastava, Varun, et al.
Published: (2025)
From Cloud to Edge: Rethinking Generative AI for Low-Resource Design Challenges
by: Vuruma, Sai Krishna Revanth, et al.
Published: (2024)
by: Vuruma, Sai Krishna Revanth, et al.
Published: (2024)
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
by: Bi, Zhenyu, et al.
Published: (2025)
by: Bi, Zhenyu, et al.
Published: (2025)
PhytoSynth: Leveraging Multi-modal Generative Models for Crop Disease Data Generation with Novel Benchmarking and Prompt Engineering Approach
by: Rai, Nitin, et al.
Published: (2025)
by: Rai, Nitin, et al.
Published: (2025)
ConDABench: Interactive Evaluation of Language Models for Data Analysis
by: Dutta, Avik, et al.
Published: (2025)
by: Dutta, Avik, et al.
Published: (2025)
A Cost-Benefit Analysis of On-Premise Large Language Model Deployment: Breaking Even with Commercial LLM Services
by: Pan, Guanzhong, et al.
Published: (2025)
by: Pan, Guanzhong, et al.
Published: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
by: Zhang, Xiaotian, et al.
Published: (2023)
by: Zhang, Xiaotian, et al.
Published: (2023)
Rating Multi-Modal Time-Series Forecasting Models (MM-TSFM) for Robustness Through a Causal Lens
by: Lakkaraju, Kausik, et al.
Published: (2024)
by: Lakkaraju, Kausik, et al.
Published: (2024)
DeepMath-Creative: A Benchmark for Evaluating Mathematical Creativity of Large Language Models
by: Chen, Xiaoyang, et al.
Published: (2025)
by: Chen, Xiaoyang, et al.
Published: (2025)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
by: Zhou, Xiyuan, et al.
Published: (2025)
by: Zhou, Xiyuan, et al.
Published: (2025)
MSDiagnosis: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis
by: Hou, Ruihui, et al.
Published: (2024)
by: Hou, Ruihui, et al.
Published: (2024)
Similar Items
-
PLANTS: A Novel Problem and Dataset for Summarization of Planning-Like (PL) Tasks
by: Pallagani, Vishal, et al.
Published: (2024) -
On Sample-Efficient Generalized Planning via Learned Transition Models
by: Gupta, Nitin, et al.
Published: (2026) -
The Case for Developing a Foundation Model for Planning-like Tasks from Scratch
by: Srivastava, Biplav, et al.
Published: (2024) -
A Neurosymbolic Fast and Slow Architecture for Graph Coloring
by: Khandelwal, Vedant, et al.
Published: (2024) -
A Planning Ontology to Represent and Exploit Planning Knowledge for Performance Efficiency
by: Muppasani, Bharath, et al.
Published: (2023)