SeekerGym: A Benchmark for Reliable Information Seeking
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Remy, Lee, Minseung, Li, Shuo, Bastani, Osbert |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Asymptotic Normality of Generalized Low-Rank Matrix Sensing via Riemannian Geometry
por: Bastani, Osbert
Publicado: (2024)
por: Bastani, Osbert
Publicado: (2024)
Conformal Structured Prediction
por: Zhang, Botong, et al.
Publicado: (2024)
por: Zhang, Botong, et al.
Publicado: (2024)
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
por: Ge, Haosen, et al.
Publicado: (2024)
por: Ge, Haosen, et al.
Publicado: (2024)
Rethinking Algorithmic Fairness for Human-AI Collaboration
por: Ge, Haosen, et al.
Publicado: (2023)
por: Ge, Haosen, et al.
Publicado: (2023)
Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching
por: Bastani, Hamsa, et al.
Publicado: (2026)
por: Bastani, Hamsa, et al.
Publicado: (2026)
RAPID: An Efficient Reinforcement Learning Algorithm for Small Language Models
por: Huang, Lianghuan, et al.
Publicado: (2025)
por: Huang, Lianghuan, et al.
Publicado: (2025)
Beating the Winner's Curse via Inference-Aware Policy Optimization
por: Bastani, Hamsa, et al.
Publicado: (2025)
por: Bastani, Hamsa, et al.
Publicado: (2025)
Improving Human Sequential Decision-Making with Reinforcement Learning
por: Bastani, Hamsa, et al.
Publicado: (2021)
por: Bastani, Hamsa, et al.
Publicado: (2021)
Group-Sparse Matrix Factorization for Transfer Learning of Word Embeddings
por: Xu, Kan, et al.
Publicado: (2021)
por: Xu, Kan, et al.
Publicado: (2021)
LLM Program Optimization via Retrieval Augmented Search
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
Stochastic Bandits with ReLU Neural Networks
por: Xu, Kan, et al.
Publicado: (2024)
por: Xu, Kan, et al.
Publicado: (2024)
Uncertainty Quantification for Neurosymbolic Programs via Compositional Conformal Prediction
por: Ramalingam, Ramya, et al.
Publicado: (2024)
por: Ramalingam, Ramya, et al.
Publicado: (2024)
SPARLING: Learning Latent Representations with Extremely Sparse Activations
por: Gupta, Kavi, et al.
Publicado: (2023)
por: Gupta, Kavi, et al.
Publicado: (2023)
Conformal Constrained Policy Optimization for Cost-Effective LLM Agents
por: Si, Wenwen, et al.
Publicado: (2025)
por: Si, Wenwen, et al.
Publicado: (2025)
Prior-Agnostic Incentive-Compatible Exploration
por: Ramalingam, Ramya, et al.
Publicado: (2026)
por: Ramalingam, Ramya, et al.
Publicado: (2026)
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions
por: Mell, Stephen, et al.
Publicado: (2025)
por: Mell, Stephen, et al.
Publicado: (2025)
Diversity By Design: Leveraging Distribution Matching for Offline Model-Based Optimization
por: Yao, Michael S., et al.
Publicado: (2025)
por: Yao, Michael S., et al.
Publicado: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
Improving Structural Diversity of Blackbox LLMs via Chain-of-Specification Prompting
por: Young, Halley, et al.
Publicado: (2024)
por: Young, Halley, et al.
Publicado: (2024)
Alignment of large language models with constrained learning
por: Zhang, Botong, et al.
Publicado: (2025)
por: Zhang, Botong, et al.
Publicado: (2025)
Generative Adversarial Model-Based Optimization via Source Critic Regularization
por: Yao, Michael S., et al.
Publicado: (2024)
por: Yao, Michael S., et al.
Publicado: (2024)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
por: Huang, Xinmeng, et al.
Publicado: (2024)
por: Huang, Xinmeng, et al.
Publicado: (2024)
Uncertainty in Language Models: Assessment through Rank-Calibration
por: Huang, Xinmeng, et al.
Publicado: (2024)
por: Huang, Xinmeng, et al.
Publicado: (2024)
Knowledgeable Language Models as Black-Box Optimizers for Personalized Medicine
por: Yao, Michael S., et al.
Publicado: (2025)
por: Yao, Michael S., et al.
Publicado: (2025)
AExGym: Benchmarks and Environments for Adaptive Experimentation
por: Wang, Jimmy, et al.
Publicado: (2024)
por: Wang, Jimmy, et al.
Publicado: (2024)
Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting
por: Kim, Seongmin, et al.
Publicado: (2025)
por: Kim, Seongmin, et al.
Publicado: (2025)
Neural Language of Thought Models
por: Wu, Yi-Fu, et al.
Publicado: (2024)
por: Wu, Yi-Fu, et al.
Publicado: (2024)
MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning
por: Wang, Yuepeng, et al.
Publicado: (2026)
por: Wang, Yuepeng, et al.
Publicado: (2026)
CrystalGym: A New Benchmark for Materials Discovery Using Reinforcement Learning
por: Govindarajan, Prashant, et al.
Publicado: (2025)
por: Govindarajan, Prashant, et al.
Publicado: (2025)
Eurekaverse: Environment Curriculum Generation via Large Language Models
por: Liang, William, et al.
Publicado: (2024)
por: Liang, William, et al.
Publicado: (2024)
BoxingGym: Benchmarking Progress in Automated Experimental Design and Model Discovery
por: Gandhi, Kanishk, et al.
Publicado: (2025)
por: Gandhi, Kanishk, et al.
Publicado: (2025)
BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing
por: Gu, Yunqi, et al.
Publicado: (2025)
por: Gu, Yunqi, et al.
Publicado: (2025)
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents
por: Pleines, Marco, et al.
Publicado: (2023)
por: Pleines, Marco, et al.
Publicado: (2023)
Adversarial Query Synthesis via Bayesian Optimization
por: Tao, Jeffrey, et al.
Publicado: (2026)
por: Tao, Jeffrey, et al.
Publicado: (2026)
TRAQ: Trustworthy Retrieval Augmented Question Answering via Conformal Prediction
por: Li, Shuo, et al.
Publicado: (2023)
por: Li, Shuo, et al.
Publicado: (2023)
Sliding Puzzles Gym: A Scalable Benchmark for State Representation in Visual Reinforcement Learning
por: de Oliveira, Bryan L. M., et al.
Publicado: (2024)
por: de Oliveira, Bryan L. M., et al.
Publicado: (2024)
SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems
por: Ramanujam, Asha, et al.
Publicado: (2025)
por: Ramanujam, Asha, et al.
Publicado: (2025)
TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
por: Cai, Yifu, et al.
Publicado: (2025)
por: Cai, Yifu, et al.
Publicado: (2025)
Are AI Capabilities Increasing Exponentially? A Competing Hypothesis
por: Ge, Haosen, et al.
Publicado: (2026)
por: Ge, Haosen, et al.
Publicado: (2026)
Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning
por: Salaorni, Davide, et al.
Publicado: (2025)
por: Salaorni, Davide, et al.
Publicado: (2025)
Ejemplares similares
-
Asymptotic Normality of Generalized Low-Rank Matrix Sensing via Riemannian Geometry
por: Bastani, Osbert
Publicado: (2024) -
Conformal Structured Prediction
por: Zhang, Botong, et al.
Publicado: (2024) -
Stochastic Online Conformal Prediction with Semi-Bandit Feedback
por: Ge, Haosen, et al.
Publicado: (2024) -
Rethinking Algorithmic Fairness for Human-AI Collaboration
por: Ge, Haosen, et al.
Publicado: (2023) -
Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching
por: Bastani, Hamsa, et al.
Publicado: (2026)