The Reciprocity Gradient
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Yue, Poupart, Pascal, Zhu, Shuhui, Qiao, Dan, Li, Wenhao, Liu, Yuan, Zha, Hongyuan, Wang, Baoxiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Policy-Conditioned Policies for Multi-Agent Task Solving
di: Lin, Yue, et al.
Pubblicazione: (2025)
di: Lin, Yue, et al.
Pubblicazione: (2025)
Information Bargaining: Bilateral Commitment in Bayesian Persuasion
di: Lin, Yue, et al.
Pubblicazione: (2025)
di: Lin, Yue, et al.
Pubblicazione: (2025)
Learning to Negotiate via Voluntary Commitment
di: Zhu, Shuhui, et al.
Pubblicazione: (2025)
di: Zhu, Shuhui, et al.
Pubblicazione: (2025)
Carbon Market Simulation with Adaptive Mechanism Design
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Verbalized Bayesian Persuasion
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents
di: Zhu, Shuhui, et al.
Pubblicazione: (2026)
di: Zhu, Shuhui, et al.
Pubblicazione: (2026)
Measures of Variability for Risk-averse Policy Gradient
di: Luo, Yudong, et al.
Pubblicazione: (2025)
di: Luo, Yudong, et al.
Pubblicazione: (2025)
Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition
di: Qiao, Dan, et al.
Pubblicazione: (2025)
di: Qiao, Dan, et al.
Pubblicazione: (2025)
Why Online Reinforcement Learning is Causal
di: Schulte, Oliver, et al.
Pubblicazione: (2024)
di: Schulte, Oliver, et al.
Pubblicazione: (2024)
TDHook: A Lightweight Framework for Interpretability
di: Poupart, Yoann
Pubblicazione: (2025)
di: Poupart, Yoann
Pubblicazione: (2025)
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
di: Jeong, Jihwan, et al.
Pubblicazione: (2025)
di: Jeong, Jihwan, et al.
Pubblicazione: (2025)
A Comprehensive Survey on Inverse Constrained Reinforcement Learning: Definitions, Progress and Challenges
di: Liu, Guiliang, et al.
Pubblicazione: (2024)
di: Liu, Guiliang, et al.
Pubblicazione: (2024)
Transfer Learning for Diffusion Models
di: Ouyang, Yidong, et al.
Pubblicazione: (2024)
di: Ouyang, Yidong, et al.
Pubblicazione: (2024)
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
di: Wang, Haozhong, et al.
Pubblicazione: (2026)
di: Wang, Haozhong, et al.
Pubblicazione: (2026)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
di: Lu, Yiyang, et al.
Pubblicazione: (2026)
Decouple Graph Neural Networks: Train Multiple Simple GNNs Simultaneously Instead of One
di: Zhang, Hongyuan, et al.
Pubblicazione: (2023)
di: Zhang, Hongyuan, et al.
Pubblicazione: (2023)
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
di: Miao, Yanting, et al.
Pubblicazione: (2025)
di: Miao, Yanting, et al.
Pubblicazione: (2025)
Randomness and Interpolation Improve Gradient Descent
di: Li, Jiawen, et al.
Pubblicazione: (2025)
di: Li, Jiawen, et al.
Pubblicazione: (2025)
Causal invariant geographic network representations with feature and structural distribution shifts
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
di: Miao, Yanting, et al.
Pubblicazione: (2024)
di: Miao, Yanting, et al.
Pubblicazione: (2024)
TextAtari: 100K Frames Game Playing with Language Agents
di: Li, Wenhao, et al.
Pubblicazione: (2025)
di: Li, Wenhao, et al.
Pubblicazione: (2025)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
di: Li, Siyuan, et al.
Pubblicazione: (2025)
di: Li, Siyuan, et al.
Pubblicazione: (2025)
Does Flatness imply Generalization for Logistic Loss in Univariate Two-Layer ReLU Network?
di: Qiao, Dan, et al.
Pubblicazione: (2025)
di: Qiao, Dan, et al.
Pubblicazione: (2025)
Relative Policy-Transition Optimization for Fast Policy Transfer
di: Xu, Jiawei, et al.
Pubblicazione: (2022)
di: Xu, Jiawei, et al.
Pubblicazione: (2022)
Learning Soft Driving Constraints from Vectorized Scene Embeddings while Imitating Expert Trajectories
di: Mobarakeh, Niloufar Saeidi, et al.
Pubblicazione: (2024)
di: Mobarakeh, Niloufar Saeidi, et al.
Pubblicazione: (2024)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
di: Li, Yingru, et al.
Pubblicazione: (2025)
di: Li, Yingru, et al.
Pubblicazione: (2025)
Learning to Communicate Through Implicit Communication Channels
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
GradientStabilizer:Fix the Norm, Not the Gradient
di: Huang, Tianjin, et al.
Pubblicazione: (2025)
di: Huang, Tianjin, et al.
Pubblicazione: (2025)
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning
di: Liu, Zeyuan, et al.
Pubblicazione: (2025)
di: Liu, Zeyuan, et al.
Pubblicazione: (2025)
ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization
di: Su, Hongyuan, et al.
Pubblicazione: (2026)
di: Su, Hongyuan, et al.
Pubblicazione: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
di: Li, Yingru, et al.
Pubblicazione: (2025)
di: Li, Yingru, et al.
Pubblicazione: (2025)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
di: Yuan, Hui, et al.
Pubblicazione: (2024)
di: Yuan, Hui, et al.
Pubblicazione: (2024)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
di: Lei, Yunwen, et al.
Pubblicazione: (2026)
di: Lei, Yunwen, et al.
Pubblicazione: (2026)
FORLER: Federated Offline Reinforcement Learning with Q-Ensemble and Actor Rectification
di: Qiao, Nan, et al.
Pubblicazione: (2026)
di: Qiao, Nan, et al.
Pubblicazione: (2026)
GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
di: Luo, Wenying, et al.
Pubblicazione: (2025)
di: Luo, Wenying, et al.
Pubblicazione: (2025)
Interpretable Hybrid-Rule Temporal Point Processes
di: Cao, Yunyang, et al.
Pubblicazione: (2025)
di: Cao, Yunyang, et al.
Pubblicazione: (2025)
Tackling Data Corruption in Offline Reinforcement Learning via Sequence Modeling
di: Xu, Jiawei, et al.
Pubblicazione: (2024)
di: Xu, Jiawei, et al.
Pubblicazione: (2024)
TaxDistill: Improving Metagenomic Taxonomic Annotation via Distilled Genomic Foundation Models
di: Ye, Rongye, et al.
Pubblicazione: (2026)
di: Ye, Rongye, et al.
Pubblicazione: (2026)
Factor Decorrelation Enhanced Data Removal from Deep Predictive Models
di: Yang, Wenhao, et al.
Pubblicazione: (2025)
di: Yang, Wenhao, et al.
Pubblicazione: (2025)
Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints
di: Qiao, Dan, et al.
Pubblicazione: (2024)
di: Qiao, Dan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Policy-Conditioned Policies for Multi-Agent Task Solving
di: Lin, Yue, et al.
Pubblicazione: (2025) -
Information Bargaining: Bilateral Commitment in Bayesian Persuasion
di: Lin, Yue, et al.
Pubblicazione: (2025) -
Learning to Negotiate via Voluntary Commitment
di: Zhu, Shuhui, et al.
Pubblicazione: (2025) -
Carbon Market Simulation with Adaptive Mechanism Design
di: Wang, Han, et al.
Pubblicazione: (2024) -
Verbalized Bayesian Persuasion
di: Li, Wenhao, et al.
Pubblicazione: (2025)