MARPLE: A Benchmark for Long-Horizon Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Jin, Emily, Huang, Zhuoyi, Fränken, Jan-Philipp, Liu, Weiyu, Cha, Hannah, Brockbank, Erik, Wu, Sarah, Zhang, Ruohan, Wu, Jiajun, Gerstenberg, Tobias |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
di: Andukuri, Chinmaya, et al.
Pubblicazione: (2024)
di: Andukuri, Chinmaya, et al.
Pubblicazione: (2024)
Spot The Ball: A Benchmark for Visual Social Inference
di: Balamurugan, Neha, et al.
Pubblicazione: (2025)
di: Balamurugan, Neha, et al.
Pubblicazione: (2025)
Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
di: Cross, Logan, et al.
Pubblicazione: (2025)
di: Cross, Logan, et al.
Pubblicazione: (2025)
Why Someone Asked "Why": Foil Inference in Human and LLM Question Interpretation
di: Besch, Britt, et al.
Pubblicazione: (2026)
di: Besch, Britt, et al.
Pubblicazione: (2026)
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
di: Feng, Chun, et al.
Pubblicazione: (2024)
di: Feng, Chun, et al.
Pubblicazione: (2024)
Causal-PIK: Causality-based Physical Reasoning with a Physics-Informed Kernel
di: Parés-Morlans, Carlota, et al.
Pubblicazione: (2025)
di: Parés-Morlans, Carlota, et al.
Pubblicazione: (2025)
Learning Compositional Behaviors from Demonstration and Language
di: Liu, Weiyu, et al.
Pubblicazione: (2025)
di: Liu, Weiyu, et al.
Pubblicazione: (2025)
Predicate Hierarchies Improve Few-Shot State Classification
di: Jin, Emily, et al.
Pubblicazione: (2025)
di: Jin, Emily, et al.
Pubblicazione: (2025)
Post Reinforcement Learning Inference
di: Syrgkanis, Vasilis, et al.
Pubblicazione: (2023)
di: Syrgkanis, Vasilis, et al.
Pubblicazione: (2023)
CRAFT: Designing Creative and Functional 3D Objects
di: Guo, Michelle, et al.
Pubblicazione: (2024)
di: Guo, Michelle, et al.
Pubblicazione: (2024)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2026)
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2026)
Reasoning Cache: Continual Improvement Over Long Horizons via Short-Horizon RL
di: Wu, Ian, et al.
Pubblicazione: (2026)
di: Wu, Ian, et al.
Pubblicazione: (2026)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
di: Li, Manling, et al.
Pubblicazione: (2024)
di: Li, Manling, et al.
Pubblicazione: (2024)
Multi-Omics Analysis for Cancer Subtype Inference via Unrolling Graph Smoothness Priors
di: Lu, Jielong, et al.
Pubblicazione: (2025)
di: Lu, Jielong, et al.
Pubblicazione: (2025)
Deep Active Inference Agents for Delayed and Long-Horizon Environments
di: Yeganeh, Yavar Taheri, et al.
Pubblicazione: (2025)
di: Yeganeh, Yavar Taheri, et al.
Pubblicazione: (2025)
TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction
di: Jiang, Yunfan, et al.
Pubblicazione: (2024)
di: Jiang, Yunfan, et al.
Pubblicazione: (2024)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
di: Shen, Yuanzhe, et al.
Pubblicazione: (2026)
di: Shen, Yuanzhe, et al.
Pubblicazione: (2026)
Learning-Guided Rolling Horizon Optimization for Long-Horizon Flexible Job-Shop Scheduling
di: Li, Sirui, et al.
Pubblicazione: (2025)
di: Li, Sirui, et al.
Pubblicazione: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
di: Jang, Lawrence Keunho, et al.
Pubblicazione: (2026)
di: Jang, Lawrence Keunho, et al.
Pubblicazione: (2026)
HoTPP Benchmark: Are We Good at the Long Horizon Events Forecasting?
di: Karpukhin, Ivan, et al.
Pubblicazione: (2024)
di: Karpukhin, Ivan, et al.
Pubblicazione: (2024)
Composable Part-Based Manipulation
di: Liu, Weiyu, et al.
Pubblicazione: (2024)
di: Liu, Weiyu, et al.
Pubblicazione: (2024)
Human-like Affective Cognition in Foundation Models
di: Gandhi, Kanishk, et al.
Pubblicazione: (2024)
di: Gandhi, Kanishk, et al.
Pubblicazione: (2024)
MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
di: Li, Chengshu, et al.
Pubblicazione: (2025)
di: Li, Chengshu, et al.
Pubblicazione: (2025)
Hearing Anything Anywhere
di: Wang, Mason, et al.
Pubblicazione: (2024)
di: Wang, Mason, et al.
Pubblicazione: (2024)
Reinforcement Learning for Long-Horizon Interactive LLM Agents
di: Chen, Kevin, et al.
Pubblicazione: (2025)
di: Chen, Kevin, et al.
Pubblicazione: (2025)
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
di: Wu, Xixi, et al.
Pubblicazione: (2026)
di: Wu, Xixi, et al.
Pubblicazione: (2026)
λ: A Benchmark for Data-Efficiency in Long-Horizon Indoor Mobile Manipulation Robotics
di: Jaafar, Ahmed, et al.
Pubblicazione: (2024)
di: Jaafar, Ahmed, et al.
Pubblicazione: (2024)
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
di: Chen, Xianwei, et al.
Pubblicazione: (2026)
di: Chen, Xianwei, et al.
Pubblicazione: (2026)
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
di: Xu, Pei, et al.
Pubblicazione: (2025)
di: Xu, Pei, et al.
Pubblicazione: (2025)
Predicting Outcomes in Video Games with Long Short Term Memory Networks
di: Chulajata, Kittimate, et al.
Pubblicazione: (2024)
di: Chulajata, Kittimate, et al.
Pubblicazione: (2024)
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
di: Jin, Xutong, et al.
Pubblicazione: (2024)
di: Jin, Xutong, et al.
Pubblicazione: (2024)
Probe and Skip: Self-Predictive Token Skipping for Efficient Long-Context LLM Inference
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
di: Wu, Zimeng, et al.
Pubblicazione: (2026)
On Policy Evaluation Algorithms in Distributional Reinforcement Learning
di: Gerstenberg, Julian, et al.
Pubblicazione: (2024)
di: Gerstenberg, Julian, et al.
Pubblicazione: (2024)
The Best of Both Worlds: Hybridizing Neural Operators and Solvers for Stable Long-Horizon Inference
di: Roy, Rajyasri, et al.
Pubblicazione: (2025)
di: Roy, Rajyasri, et al.
Pubblicazione: (2025)
GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
di: Li, Yunfei, et al.
Pubblicazione: (2025)
di: Li, Yunfei, et al.
Pubblicazione: (2025)
RaC: Robot Learning for Long-Horizon Tasks by Scaling Recovery and Correction
di: Hu, Zheyuan, et al.
Pubblicazione: (2025)
di: Hu, Zheyuan, et al.
Pubblicazione: (2025)
Effective Explanations Support Planning Under Uncertainty
di: Zhou, Hanqi, et al.
Pubblicazione: (2026)
di: Zhou, Hanqi, et al.
Pubblicazione: (2026)
The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
di: Chung, Jae-Won, et al.
Pubblicazione: (2025)
di: Chung, Jae-Won, et al.
Pubblicazione: (2025)
Documenti analoghi
-
STaR-GATE: Teaching Language Models to Ask Clarifying Questions
di: Andukuri, Chinmaya, et al.
Pubblicazione: (2024) -
Spot The Ball: A Benchmark for Visual Social Inference
di: Balamurugan, Neha, et al.
Pubblicazione: (2025) -
Understanding Human Limits in Pattern Recognition: A Computational Model of Sequential Reasoning in Rock, Paper, Scissors
di: Cross, Logan, et al.
Pubblicazione: (2025) -
Why Someone Asked "Why": Foil Inference in Human and LLM Question Interpretation
di: Besch, Britt, et al.
Pubblicazione: (2026) -
Self-Supervised Alignment with Mutual Information: Learning to Follow Principles without Preference Labels
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)