Saved in:
| Main Authors: | Vira, Jash, Harris, Ashley |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.09594 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Class Boundary Extraction from Implicit Representations
by: Vira, Jash, et al.
Published: (2026)
by: Vira, Jash, et al.
Published: (2026)
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026)
by: Sayana, Krishna, et al.
Published: (2026)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
by: Bohacek, Maty, et al.
Published: (2025)
by: Bohacek, Maty, et al.
Published: (2025)
User Embedding Model for Personalized Language Prompting
by: Doddapaneni, Sumanth, et al.
Published: (2024)
by: Doddapaneni, Sumanth, et al.
Published: (2024)
REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives
by: Su, Kun, et al.
Published: (2025)
by: Su, Kun, et al.
Published: (2025)
Can AI Freelancers Compete? Benchmarking Earnings, Reliability, and Task Success at Scale
by: Noever, David, et al.
Published: (2025)
by: Noever, David, et al.
Published: (2025)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
by: Ruoss, Anian, et al.
Published: (2024)
by: Ruoss, Anian, et al.
Published: (2024)
AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
by: Nguyen, Hoang, et al.
Published: (2025)
by: Nguyen, Hoang, et al.
Published: (2025)
EVGeoQA: Benchmarking LLMs on Dynamic, Multi-Objective Geo-Spatial Exploration
by: Wu, Jianfei, et al.
Published: (2026)
by: Wu, Jianfei, et al.
Published: (2026)
Multi-Grained Temporal-Spatial Graph Learning for Stable Traffic Flow Forecasting
by: Lin, Zhenan, et al.
Published: (2025)
by: Lin, Zhenan, et al.
Published: (2025)
AgroFlux: A Spatial-Temporal Benchmark for Carbon and Nitrogen Flux Prediction in Agricultural Ecosystems
by: Cheng, Qi, et al.
Published: (2026)
by: Cheng, Qi, et al.
Published: (2026)
Do Attention Heads Compete or Cooperate during Counting?
by: Zsámboki, Pál, et al.
Published: (2025)
by: Zsámboki, Pál, et al.
Published: (2025)
STG4Traffic: A Survey and Benchmark of Spatial-Temporal Graph Neural Networks for Traffic Prediction
by: Luo, Xunlian, et al.
Published: (2023)
by: Luo, Xunlian, et al.
Published: (2023)
PACED: Distillation and On-Policy Self-Distillation at the Frontier of Student Competence
by: Xu, Yuanda, et al.
Published: (2026)
by: Xu, Yuanda, et al.
Published: (2026)
Comprehension Without Competence: Architectural Limits of LLMs in Symbolic Computation and Reasoning
by: Zhang, Zheng
Published: (2025)
by: Zhang, Zheng
Published: (2025)
Competence-Aware AI Agents with Metacognition for Unknown Situations and Environments (MUSE)
by: Valiente, Rodolfo, et al.
Published: (2024)
by: Valiente, Rodolfo, et al.
Published: (2024)
Survival Models: Proper Scoring Rule and Stochastic Optimization with Competing Risks
by: Alberge, Julie, et al.
Published: (2024)
by: Alberge, Julie, et al.
Published: (2024)
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
by: Sharma, Arun
Published: (2026)
by: Sharma, Arun
Published: (2026)
Can TabPFN Compete with GNNs for Node Classification via Graph Tabularization?
by: Choi, Jeongwhan, et al.
Published: (2025)
by: Choi, Jeongwhan, et al.
Published: (2025)
SpatialMAGIC: A Hybrid Framework Integrating Graph Diffusion and Spatial Attention for Spatial Transcriptomics Imputation
by: Zaman, Sayeem Bin, et al.
Published: (2026)
by: Zaman, Sayeem Bin, et al.
Published: (2026)
Explaining Low Perception Model Competency with High-Competency Counterfactuals
by: Pohland, Sara, et al.
Published: (2025)
by: Pohland, Sara, et al.
Published: (2025)
KIPPO: Koopman-Inspired Proximal Policy Optimization
by: Cozma, Andrei, et al.
Published: (2025)
by: Cozma, Andrei, et al.
Published: (2025)
Robust Real-Time Mortality Prediction in the Intensive Care Unit using Temporal Difference Learning
by: Frost, Thomas, et al.
Published: (2024)
by: Frost, Thomas, et al.
Published: (2024)
Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
by: Kong, Deyang, et al.
Published: (2025)
by: Kong, Deyang, et al.
Published: (2025)
Enabling Multi-Agent Transfer Reinforcement Learning via Scenario Independent Representation
by: Nipu, Ayesha Siddika, et al.
Published: (2024)
by: Nipu, Ayesha Siddika, et al.
Published: (2024)
MAIDCRL: Semi-centralized Multi-Agent Influence Dense-CNN Reinforcement Learning
by: Nipu, Ayesha Siddika, et al.
Published: (2024)
by: Nipu, Ayesha Siddika, et al.
Published: (2024)
Beyond Retrieval: Generating Narratives in Conversational Recommender Systems
by: Sayana, Krishna, et al.
Published: (2024)
by: Sayana, Krishna, et al.
Published: (2024)
SpatialTraceGen: High-Fidelity Traces for Efficient VLM Spatial Reasoning Distillation
by: Huh, Gio, et al.
Published: (2025)
by: Huh, Gio, et al.
Published: (2025)
STD-PLM: Understanding Both Spatial and Temporal Properties of Spatial-Temporal Data with PLM
by: Huang, YiHeng, et al.
Published: (2024)
by: Huang, YiHeng, et al.
Published: (2024)
PepBenchmark: A Standardized Benchmark for Peptide Machine Learning
by: Zhang, Jiahui, et al.
Published: (2026)
by: Zhang, Jiahui, et al.
Published: (2026)
Fast Benchmarking of Asynchronous Multi-Fidelity Optimization on Zero-Cost Benchmarks
by: Watanabe, Shuhei, et al.
Published: (2024)
by: Watanabe, Shuhei, et al.
Published: (2024)
Causal Masking on Spatial Data: An Information-Theoretic Case for Learning Spatial Datasets with Unimodal Language Models
by: Junkin, Jared, et al.
Published: (2025)
by: Junkin, Jared, et al.
Published: (2025)
CMP: Robust Whole-Body Tracking for Loco-Manipulation via Competence Manifold Projection
by: Cheng, Ziyang, et al.
Published: (2026)
by: Cheng, Ziyang, et al.
Published: (2026)
Spatially-Aware Transformer for Embodied Agents
by: Cho, Junmo, et al.
Published: (2024)
by: Cho, Junmo, et al.
Published: (2024)
FiSH: Fair Spatial Hotspots
by: P, Deepak, et al.
Published: (2021)
by: P, Deepak, et al.
Published: (2021)
Fast Optimizer Benchmark
by: Blauth, Simon, et al.
Published: (2024)
by: Blauth, Simon, et al.
Published: (2024)
Benchmark Success, Clinical Failure: When Reinforcement Learning Optimizes for Benchmarks, Not Patients
by: Berger, Armin, et al.
Published: (2025)
by: Berger, Armin, et al.
Published: (2025)
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
by: Pan, Haining, et al.
Published: (2025)
by: Pan, Haining, et al.
Published: (2025)
Model Ensembling for Constrained Optimization
by: Globus-Harris, Ira, et al.
Published: (2024)
by: Globus-Harris, Ira, et al.
Published: (2024)
Should You Use Your Large Language Model to Explore or Exploit?
by: Harris, Keegan, et al.
Published: (2025)
by: Harris, Keegan, et al.
Published: (2025)
Similar Items
-
Multi-Class Boundary Extraction from Implicit Representations
by: Vira, Jash, et al.
Published: (2026) -
Prompting Policies for Multi-step Reasoning and Tool-Use in Black-box LLMs with Iterative Distillation of Experience
by: Sayana, Krishna, et al.
Published: (2026) -
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
by: Bohacek, Maty, et al.
Published: (2025) -
User Embedding Model for Personalized Language Prompting
by: Doddapaneni, Sumanth, et al.
Published: (2024) -
REGEN: A Dataset and Benchmarks with Natural Language Critiques and Narratives
by: Su, Kun, et al.
Published: (2025)