R-ParVI: Particle-based variational inference through lens of rewards
Fuente:
arXiv
Guardado en:
| Autor principal: | Huang, Yongchao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Electrostatics-based particle sampling and approximate inference
por: Huang, Yongchao
Publicado: (2024)
por: Huang, Yongchao
Publicado: (2024)
Variational Inference via Smoothed Particle Hydrodynamics
por: Huang, Yongchao
Publicado: (2024)
por: Huang, Yongchao
Publicado: (2024)
Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
por: Molinghen, Yannick, et al.
Publicado: (2025)
por: Molinghen, Yannick, et al.
Publicado: (2025)
reward-lens: A Mechanistic Interpretability Library for Reward Models
por: Nadaf, Mohammed Suhail B
Publicado: (2026)
por: Nadaf, Mohammed Suhail B
Publicado: (2026)
Training data membership inference via Gaussian process meta-modeling: a post-hoc analysis approach
por: Huang, Yongchao, et al.
Publicado: (2025)
por: Huang, Yongchao, et al.
Publicado: (2025)
LLM-BI: Towards Fully Automated Bayesian Inference with Large Language Models
por: Huang, Yongchao
Publicado: (2025)
por: Huang, Yongchao
Publicado: (2025)
Noise-based reward-modulated learning
por: Fernández, Jesús García, et al.
Publicado: (2025)
por: Fernández, Jesús García, et al.
Publicado: (2025)
Bayesian Inference of Training Dataset Membership
por: Huang, Yongchao
Publicado: (2025)
por: Huang, Yongchao
Publicado: (2025)
Risk-averse Total-reward MDPs with ERM and EVaR
por: Su, Xihong, et al.
Publicado: (2024)
por: Su, Xihong, et al.
Publicado: (2024)
MOSLIM:Align with diverse preferences in prompts through reward classification
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Variational Inference Using Material Point Method
por: Huang, Yongchao
Publicado: (2024)
por: Huang, Yongchao
Publicado: (2024)
Semantic Fusion with Fuzzy-Membership Features for Controllable Language Modelling
por: Huang, Yongchao, et al.
Publicado: (2025)
por: Huang, Yongchao, et al.
Publicado: (2025)
EVAL: EigenVector-based Average-reward Learning
por: Adamczyk, Jacob, et al.
Publicado: (2025)
por: Adamczyk, Jacob, et al.
Publicado: (2025)
Enhancing Aspect-based Sentiment Analysis with ParsBERT in Persian Language
por: Ariai, Farid, et al.
Publicado: (2025)
por: Ariai, Farid, et al.
Publicado: (2025)
What should be observed for optimal reward in POMDPs?
por: Konsta, Alyzia-Maria, et al.
Publicado: (2024)
por: Konsta, Alyzia-Maria, et al.
Publicado: (2024)
Probing RLVR training instability through the lens of objective-level hacking
por: Dong, Yiming, et al.
Publicado: (2026)
por: Dong, Yiming, et al.
Publicado: (2026)
Batch and match: black-box variational inference with a score-based divergence
por: Cai, Diana, et al.
Publicado: (2024)
por: Cai, Diana, et al.
Publicado: (2024)
Project Prometheus: Bridging the Intent Gap in Agentic Program Repair via Reverse-Engineered Executable Specifications
por: Wang, Yongchao, et al.
Publicado: (2026)
por: Wang, Yongchao, et al.
Publicado: (2026)
Scalable and Accurate Graph Reasoning with LLM-based Multi-Agents
por: Hu, Yuwei, et al.
Publicado: (2024)
por: Hu, Yuwei, et al.
Publicado: (2024)
Towards better dense rewards in Reinforcement Learning Applications
por: Zhang, Shuyuan
Publicado: (2025)
por: Zhang, Shuyuan
Publicado: (2025)
Self-rewarding correction for mathematical reasoning
por: Xiong, Wei, et al.
Publicado: (2025)
por: Xiong, Wei, et al.
Publicado: (2025)
Active teacher selection for reward learning
por: Freedman, Rachel, et al.
Publicado: (2023)
por: Freedman, Rachel, et al.
Publicado: (2023)
ParBalans: Parallel Multi-Armed Bandits-based Adaptive Large Neighborhood Search
por: Yilmaz, Alican, et al.
Publicado: (2025)
por: Yilmaz, Alican, et al.
Publicado: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
por: Liu, Shih-Yang, et al.
Publicado: (2026)
por: Liu, Shih-Yang, et al.
Publicado: (2026)
Information-theoretic analysis of world models in optimal reward maximizers
por: Harwood, Alfred, et al.
Publicado: (2026)
por: Harwood, Alfred, et al.
Publicado: (2026)
Cognitive phantoms in LLMs through the lens of latent variables
por: Peereboom, Sanne, et al.
Publicado: (2024)
por: Peereboom, Sanne, et al.
Publicado: (2024)
The impact of intrinsic rewards on exploration in Reinforcement Learning
por: Kayal, Aya, et al.
Publicado: (2025)
por: Kayal, Aya, et al.
Publicado: (2025)
LLM generation novelty through the lens of semantic similarity
por: Davydov, Philipp, et al.
Publicado: (2025)
por: Davydov, Philipp, et al.
Publicado: (2025)
ParLS-PBO: A Parallel Local Search Solver for Pseudo Boolean Optimization
por: Chen, Zhihan, et al.
Publicado: (2024)
por: Chen, Zhihan, et al.
Publicado: (2024)
Streaming Looking Ahead with Token-level Self-reward
por: Zhang, Hongming, et al.
Publicado: (2025)
por: Zhang, Hongming, et al.
Publicado: (2025)
Episodic Reinforcement Learning with Expanded State-reward Space
por: Liang, Dayang, et al.
Publicado: (2024)
por: Liang, Dayang, et al.
Publicado: (2024)
Deep Reinforcement Learning with anticipatory reward in LSTM for Collision Avoidance of Mobile Robots
por: Poulet, Olivier, et al.
Publicado: (2025)
por: Poulet, Olivier, et al.
Publicado: (2025)
Self-supervised network distillation: an effective approach to exploration in sparse reward environments
por: Pecháč, Matej, et al.
Publicado: (2023)
por: Pecháč, Matej, et al.
Publicado: (2023)
Multi-agent cooperation through in-context co-player inference
por: Weis, Marissa A., et al.
Publicado: (2026)
por: Weis, Marissa A., et al.
Publicado: (2026)
InPars+: Supercharging Synthetic Data Generation for Information Retrieval Systems
por: Krastev, Matey, et al.
Publicado: (2025)
por: Krastev, Matey, et al.
Publicado: (2025)
Resolving space-sharing conflicts in road user interactions through uncertainty reduction: An active inference-based computational model
por: Schumann, Julian F., et al.
Publicado: (2026)
por: Schumann, Julian F., et al.
Publicado: (2026)
Adversarial robustness of VAEs through the lens of local geometry
por: Khan, Asif, et al.
Publicado: (2022)
por: Khan, Asif, et al.
Publicado: (2022)
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization
por: Xu, Ruijie, et al.
Publicado: (2024)
por: Xu, Ruijie, et al.
Publicado: (2024)
ParMod: A Parallel and Modular Framework for Learning Non-Markovian Tasks
por: Miao, Ruixuan, et al.
Publicado: (2024)
por: Miao, Ruixuan, et al.
Publicado: (2024)
ParFam -- (Neural Guided) Symbolic Regression Based on Continuous Global Optimization
por: Scholl, Philipp, et al.
Publicado: (2023)
por: Scholl, Philipp, et al.
Publicado: (2023)
Ejemplares similares
-
Electrostatics-based particle sampling and approximate inference
por: Huang, Yongchao
Publicado: (2024) -
Variational Inference via Smoothed Particle Hydrodynamics
por: Huang, Yongchao
Publicado: (2024) -
Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
por: Molinghen, Yannick, et al.
Publicado: (2025) -
reward-lens: A Mechanistic Interpretability Library for Reward Models
por: Nadaf, Mohammed Suhail B
Publicado: (2026) -
Training data membership inference via Gaussian process meta-modeling: a post-hoc analysis approach
por: Huang, Yongchao, et al.
Publicado: (2025)