Identifying the Best Transition Law
Fuente:
arXiv
Salvato in:
| Autori principali: | Ahmadipour, Mehrasa, Crepon, élise, Garivier, Aurélien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EVaR-Optimal Arm Identification in Bandits
di: Ahmadipour, Mehrasa, et al.
Pubblicazione: (2025)
di: Ahmadipour, Mehrasa, et al.
Pubblicazione: (2025)
Sequential Learning of the Pareto Front for Multi-objective Bandits
di: Crépon, Elise, et al.
Pubblicazione: (2025)
di: Crépon, Elise, et al.
Pubblicazione: (2025)
Efficient Risk-sensitive Planning via Entropic Risk Measures
di: Marthe, Alexandre, et al.
Pubblicazione: (2025)
di: Marthe, Alexandre, et al.
Pubblicazione: (2025)
Identifying the Best Arm in the Presence of Global Environment Shifts
di: Srisawad, Phurinut, et al.
Pubblicazione: (2024)
di: Srisawad, Phurinut, et al.
Pubblicazione: (2024)
About the Cost of Central Privacy in Density Estimation
di: Lalanne, Clément, et al.
Pubblicazione: (2023)
di: Lalanne, Clément, et al.
Pubblicazione: (2023)
Stochastic Direct Search Method for Blind Resource Allocation
di: Achddou, Juliette, et al.
Pubblicazione: (2022)
di: Achddou, Juliette, et al.
Pubblicazione: (2022)
Beyond Average Return in Markov Decision Processes
di: Marthe, Alexandre, et al.
Pubblicazione: (2023)
di: Marthe, Alexandre, et al.
Pubblicazione: (2023)
Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
di: Huang, Audrey, et al.
Pubblicazione: (2025)
di: Huang, Audrey, et al.
Pubblicazione: (2025)
On the Statistical Complexity of Estimation and Testing under Privacy Constraints
di: Lalanne, Clément, et al.
Pubblicazione: (2022)
di: Lalanne, Clément, et al.
Pubblicazione: (2022)
Best-Arm Identification in Unimodal Bandits
di: Poiani, Riccardo, et al.
Pubblicazione: (2024)
di: Poiani, Riccardo, et al.
Pubblicazione: (2024)
Constrained Best Arm Identification with Tests for Feasibility
di: Cai, Ting, et al.
Pubblicazione: (2025)
di: Cai, Ting, et al.
Pubblicazione: (2025)
Fair Best Arm Identification with Fixed Confidence
di: Russo, Alessio, et al.
Pubblicazione: (2024)
di: Russo, Alessio, et al.
Pubblicazione: (2024)
Multi-Armed Bandits With Best-Action Queries
di: Bacchiocchi, Francesco, et al.
Pubblicazione: (2026)
di: Bacchiocchi, Francesco, et al.
Pubblicazione: (2026)
Identifying All ε-Best Arms in (Misspecified) Linear Bandits
di: Li, Zhekai, et al.
Pubblicazione: (2025)
di: Li, Zhekai, et al.
Pubblicazione: (2025)
Best-of-$\infty$ -- Asymptotic Performance of Test-Time LLM Ensembling
di: Komiyama, Junpei, et al.
Pubblicazione: (2025)
di: Komiyama, Junpei, et al.
Pubblicazione: (2025)
STEB: In Search of the Best Evaluation Approach for Synthetic Time Series
di: Stenger, Michael, et al.
Pubblicazione: (2025)
di: Stenger, Michael, et al.
Pubblicazione: (2025)
Revisiting the (Sub)Optimality of Best-of-N for Inference-Time Alignment
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
di: Sriraman, Ved, et al.
Pubblicazione: (2026)
Two-Fidelity Best-Action Identification for Stochastic Minimax Tree
di: Chen, Peter, et al.
Pubblicazione: (2026)
di: Chen, Peter, et al.
Pubblicazione: (2026)
Best-of-Tails: Bridging Optimism and Pessimism in Inference-Time Alignment
di: Hsu, Hsiang, et al.
Pubblicazione: (2026)
di: Hsu, Hsiang, et al.
Pubblicazione: (2026)
Learning to Select the Best Forecasting Tasks for Clinical Outcome Prediction
di: Xue, Yuan, et al.
Pubblicazione: (2024)
di: Xue, Yuan, et al.
Pubblicazione: (2024)
Adaptive Personalized Federated Learning via Multi-task Averaging of Kernel Mean Embeddings
di: Fermanian, Jean-Baptiste, et al.
Pubblicazione: (2026)
di: Fermanian, Jean-Baptiste, et al.
Pubblicazione: (2026)
On the Identifiability of Quantized Factors
di: Barin-Pacela, Vitória, et al.
Pubblicazione: (2023)
di: Barin-Pacela, Vitória, et al.
Pubblicazione: (2023)
Best-of-N Jailbreaking
di: Hughes, John, et al.
Pubblicazione: (2024)
di: Hughes, John, et al.
Pubblicazione: (2024)
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
di: Bushipaka, Praveen, et al.
Pubblicazione: (2025)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
di: Rho, Hyung Gyu, et al.
Pubblicazione: (2025)
di: Rho, Hyung Gyu, et al.
Pubblicazione: (2025)
Enhancing Maritime Trajectory Forecasting via H3 Index and Causal Language Modelling (CLM)
di: Drapier, Nicolas, et al.
Pubblicazione: (2024)
di: Drapier, Nicolas, et al.
Pubblicazione: (2024)
Majority of the Bests: Improving Best-of-N via Bootstrapping
di: Rakhsha, Amin, et al.
Pubblicazione: (2025)
di: Rakhsha, Amin, et al.
Pubblicazione: (2025)
Optimization Guarantees for Square-Root Natural-Gradient Variational Inference
di: Kumar, Navish, et al.
Pubblicazione: (2025)
di: Kumar, Navish, et al.
Pubblicazione: (2025)
BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
di: Reuel, Anka, et al.
Pubblicazione: (2024)
di: Reuel, Anka, et al.
Pubblicazione: (2024)
BoSS: A Best-of-Strategies Selector as an Oracle for Deep Active Learning
di: Huseljic, Denis, et al.
Pubblicazione: (2026)
di: Huseljic, Denis, et al.
Pubblicazione: (2026)
Disentanglement as Identifiable Pushforward Factorisation
di: Allen, Carl
Pubblicazione: (2024)
di: Allen, Carl
Pubblicazione: (2024)
Identifying Representations for Intervention Extrapolation
di: Saengkyongam, Sorawit, et al.
Pubblicazione: (2023)
di: Saengkyongam, Sorawit, et al.
Pubblicazione: (2023)
The Neural Pruning Law Hypothesis
di: Barbulescu, Eugen, et al.
Pubblicazione: (2025)
di: Barbulescu, Eugen, et al.
Pubblicazione: (2025)
Continuity Laws for Sequential Models
di: Yu, Annan, et al.
Pubblicazione: (2026)
di: Yu, Annan, et al.
Pubblicazione: (2026)
Variational Best-of-N Alignment
di: Amini, Afra, et al.
Pubblicazione: (2024)
di: Amini, Afra, et al.
Pubblicazione: (2024)
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
di: Tang, Yung-Chen, et al.
Pubblicazione: (2025)
Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
di: Xue, Bo, et al.
Pubblicazione: (2025)
di: Xue, Bo, et al.
Pubblicazione: (2025)
Reward Learning from Best-of-$N$ Preference Data: Targets, Tradeoffs, and Design Principles
di: Pukdee, Rattana, et al.
Pubblicazione: (2026)
di: Pukdee, Rattana, et al.
Pubblicazione: (2026)
Entropic Causal Inference: Graph Identifiability
di: Compton, Spencer, et al.
Pubblicazione: (2025)
di: Compton, Spencer, et al.
Pubblicazione: (2025)
Diversified Flow Matching with Translation Identifiability
di: Shrestha, Sagar, et al.
Pubblicazione: (2025)
di: Shrestha, Sagar, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EVaR-Optimal Arm Identification in Bandits
di: Ahmadipour, Mehrasa, et al.
Pubblicazione: (2025) -
Sequential Learning of the Pareto Front for Multi-objective Bandits
di: Crépon, Elise, et al.
Pubblicazione: (2025) -
Efficient Risk-sensitive Planning via Entropic Risk Measures
di: Marthe, Alexandre, et al.
Pubblicazione: (2025) -
Identifying the Best Arm in the Presence of Global Environment Shifts
di: Srisawad, Phurinut, et al.
Pubblicazione: (2024) -
About the Cost of Central Privacy in Density Estimation
di: Lalanne, Clément, et al.
Pubblicazione: (2023)