ShiQ: Bringing back Bellman to LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Clavier, Pierre, Grinsztajn, Nathan, Avalos, Raphael, Flet-Berliac, Yannis, Ergun, Irem, Domingues, Omar D., Tarassov, Eugene, Pietquin, Olivier, Richemond, Pierre H., Strub, Florian, Geist, Matthieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
by: Flet-Berliac, Yannis, et al.
Published: (2024)
by: Flet-Berliac, Yannis, et al.
Published: (2024)
Averaging log-likelihoods in direct alignment
by: Grinsztajn, Nathan, et al.
Published: (2024)
by: Grinsztajn, Nathan, et al.
Published: (2024)
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Bootstrapping Expectiles in Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2024)
by: Clavier, Pierre, et al.
Published: (2024)
RRLS : Robust Reinforcement Learning Suite
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Time-Constrained Robust MDPs
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
by: Magnino, Lorenzo, et al.
Published: (2026)
by: Magnino, Lorenzo, et al.
Published: (2026)
Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning
by: Wu, Zida, et al.
Published: (2025)
by: Wu, Zida, et al.
Published: (2025)
OPERA: Automatic Offline Policy Evaluation with Re-weighted Aggregates of Multiple Estimators
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Self-Improving Robust Preference Optimization
by: Choi, Eugene, et al.
Published: (2024)
by: Choi, Eugene, et al.
Published: (2024)
World Modelling Improves Language Model Agents
by: Guo, Shangmin, et al.
Published: (2025)
by: Guo, Shangmin, et al.
Published: (2025)
TRAPs, Generalisations of MZVs, Locality and Resurgence for Quantum Field Theories
by: Clavier, Pierre J.
Published: (2025)
by: Clavier, Pierre J.
Published: (2025)
Population-aware Online Mirror Descent for Mean-Field Games by Deep Reinforcement Learning
by: Wu, Zida, et al.
Published: (2024)
by: Wu, Zida, et al.
Published: (2024)
Compact Hypercube Embeddings for Fast Text-based Wildlife Observation Retrieval
by: Moummad, Ilyass, et al.
Published: (2026)
by: Moummad, Ilyass, et al.
Published: (2026)
A Survey of Temporal Credit Assignment in Deep Reinforcement Learning
by: Pignatelli, Eduardo, et al.
Published: (2023)
by: Pignatelli, Eduardo, et al.
Published: (2023)
Periodic agent-state based Q-learning for POMDPs
by: Sinha, Amit, et al.
Published: (2024)
by: Sinha, Amit, et al.
Published: (2024)
Convergence of regularized agent-state-based Q-learning in POMDPs
by: Sinha, Amit, et al.
Published: (2025)
by: Sinha, Amit, et al.
Published: (2025)
Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation
by: Moummad, Ilyass, et al.
Published: (2026)
by: Moummad, Ilyass, et al.
Published: (2026)
Learning in Mean Field Games: A Survey
by: Laurière, Mathieu, et al.
Published: (2022)
by: Laurière, Mathieu, et al.
Published: (2022)
Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning
by: Rita, Mathieu, et al.
Published: (2024)
by: Rita, Mathieu, et al.
Published: (2024)
Language Evolution with Deep Learning
by: Rita, Mathieu, et al.
Published: (2024)
by: Rita, Mathieu, et al.
Published: (2024)
Coalgebras, bialgebras and Rota-Baxter algebras from shuffles of rooted forests
by: Clavier, Pierre J., et al.
Published: (2025)
by: Clavier, Pierre J., et al.
Published: (2025)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
by: Robinson, David, et al.
Published: (2024)
by: Robinson, David, et al.
Published: (2024)
VITS : Variational Inference Thompson Sampling for contextual bandits
by: Clavier, Pierre, et al.
Published: (2023)
by: Clavier, Pierre, et al.
Published: (2023)
Tridendriform and dendriform Zeta Values from Schroeder trees
by: Catoire, Pierre, et al.
Published: (2025)
by: Catoire, Pierre, et al.
Published: (2025)
Bringing calorimetry (back) to life
by: Khodabandehlou, Faezeh, et al.
Published: (2026)
by: Khodabandehlou, Faezeh, et al.
Published: (2026)
Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification
by: Sarkar, Eklavya, et al.
Published: (2026)
by: Sarkar, Eklavya, et al.
Published: (2026)
Bringing the forest back: Restoration priorities in Colombia
by: Brooke A. Williams, et al.
Published: (2024)
by: Brooke A. Williams, et al.
Published: (2024)
SemPPL: Predicting pseudo-labels for better contrastive representations
by: Bošnjak, Matko, et al.
Published: (2023)
by: Bošnjak, Matko, et al.
Published: (2023)
A Diagrammatic Calculus for a Functional Model of Natural Language Semantics
by: Boyer, Matthieu Pierre
Published: (2025)
by: Boyer, Matthieu Pierre
Published: (2025)
Seawater carbonate chemistry and calcification during a study of barrier reef flat in Moorea, French Polynesia, 1998
by: Boucher, Guy, et al.
Published: (1998)
by: Boucher, Guy, et al.
Published: (1998)
Risk-seeking conservative policy iteration with agent-state based policies for Dec-POMDPs with guaranteed convergence
by: Sinha, Amit, et al.
Published: (2026)
by: Sinha, Amit, et al.
Published: (2026)
Solving robust MDPs as a sequence of static RL problems
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Bringing back the know-how: migrants and technology transfer
by: André Linard (Author)
Published: (2002)
by: André Linard (Author)
Published: (2002)
Bringing the state back in: Populism and economic nationalism in Europe
by: Paula D. Ganga
Published: (2024)
by: Paula D. Ganga
Published: (2024)
Success and failure of the sociology of culture? Bringing the arts back
by: Vera L. Zolberg
Published: (2005)
by: Vera L. Zolberg
Published: (2005)
Performance of far forward iceless blood storage containers in controlled cold environments
by: Antoine Vuong, et al.
Published: (2024)
by: Antoine Vuong, et al.
Published: (2024)
Temperature mediated back-action in micro- and nanomechanical resonators
by: Bellon, Ludovic, et al.
Published: (2024)
by: Bellon, Ludovic, et al.
Published: (2024)
Theoretical Barriers in Bellman-Based Reinforcement Learning
by: Pinon, Brieuc, et al.
Published: (2025)
by: Pinon, Brieuc, et al.
Published: (2025)
univ-lehavre/RAFALE: 1.1.1
by: Pierre-Olivier
Published: (2025)
by: Pierre-Olivier
Published: (2025)
Similar Items
-
Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion
by: Flet-Berliac, Yannis, et al.
Published: (2024) -
Averaging log-likelihoods in direct alignment
by: Grinsztajn, Nathan, et al.
Published: (2024) -
Towards Minimax Optimality of Model-based Robust Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2023) -
Bootstrapping Expectiles in Reinforcement Learning
by: Clavier, Pierre, et al.
Published: (2024) -
RRLS : Robust Reinforcement Learning Suite
by: Zouitine, Adil, et al.
Published: (2024)