The Reciprocity Gradient
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Yue, Poupart, Pascal, Zhu, Shuhui, Qiao, Dan, Li, Wenhao, Liu, Yuan, Zha, Hongyuan, Wang, Baoxiang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Policy-Conditioned Policies for Multi-Agent Task Solving
by: Lin, Yue, et al.
Published: (2025)
by: Lin, Yue, et al.
Published: (2025)
Information Bargaining: Bilateral Commitment in Bayesian Persuasion
by: Lin, Yue, et al.
Published: (2025)
by: Lin, Yue, et al.
Published: (2025)
Learning to Negotiate via Voluntary Commitment
by: Zhu, Shuhui, et al.
Published: (2025)
by: Zhu, Shuhui, et al.
Published: (2025)
Carbon Market Simulation with Adaptive Mechanism Design
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Verbalized Bayesian Persuasion
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Talk, Judge, Cooperate: Gossip-Driven Indirect Reciprocity in Self-Interested LLM Agents
by: Zhu, Shuhui, et al.
Published: (2026)
by: Zhu, Shuhui, et al.
Published: (2026)
Measures of Variability for Risk-averse Policy Gradient
by: Luo, Yudong, et al.
Published: (2025)
by: Luo, Yudong, et al.
Published: (2025)
Offline Multi-agent Reinforcement Learning via Sequential Score Decomposition
by: Qiao, Dan, et al.
Published: (2025)
by: Qiao, Dan, et al.
Published: (2025)
Why Online Reinforcement Learning is Causal
by: Schulte, Oliver, et al.
Published: (2024)
by: Schulte, Oliver, et al.
Published: (2024)
TDHook: A Lightweight Framework for Interpretability
by: Poupart, Yoann
Published: (2025)
by: Poupart, Yoann
Published: (2025)
Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens
by: Jeong, Jihwan, et al.
Published: (2025)
by: Jeong, Jihwan, et al.
Published: (2025)
A Comprehensive Survey on Inverse Constrained Reinforcement Learning: Definitions, Progress and Challenges
by: Liu, Guiliang, et al.
Published: (2024)
by: Liu, Guiliang, et al.
Published: (2024)
Transfer Learning for Diffusion Models
by: Ouyang, Yidong, et al.
Published: (2024)
by: Ouyang, Yidong, et al.
Published: (2024)
Safeguarding LLM Fine-tuning via Push-Pull Distributional Alignment
by: Wang, Haozhong, et al.
Published: (2026)
by: Wang, Haozhong, et al.
Published: (2026)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
by: Lu, Yiyang, et al.
Published: (2026)
by: Lu, Yiyang, et al.
Published: (2026)
Decouple Graph Neural Networks: Train Multiple Simple GNNs Simultaneously Instead of One
by: Zhang, Hongyuan, et al.
Published: (2023)
by: Zhang, Hongyuan, et al.
Published: (2023)
A Minimalist Method for Fine-tuning Text-to-Image Diffusion Models
by: Miao, Yanting, et al.
Published: (2025)
by: Miao, Yanting, et al.
Published: (2025)
Randomness and Interpolation Improve Gradient Descent
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
Causal invariant geographic network representations with feature and structural distribution shifts
by: Wang, Yuhan, et al.
Published: (2025)
by: Wang, Yuhan, et al.
Published: (2025)
Subject-driven Text-to-Image Generation via Preference-based Reinforcement Learning
by: Miao, Yanting, et al.
Published: (2024)
by: Miao, Yanting, et al.
Published: (2024)
TextAtari: 100K Frames Game Playing with Language Agents
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Does Flatness imply Generalization for Logistic Loss in Univariate Two-Layer ReLU Network?
by: Qiao, Dan, et al.
Published: (2025)
by: Qiao, Dan, et al.
Published: (2025)
Relative Policy-Transition Optimization for Fast Policy Transfer
by: Xu, Jiawei, et al.
Published: (2022)
by: Xu, Jiawei, et al.
Published: (2022)
Learning Soft Driving Constraints from Vectorized Scene Embeddings while Imitating Expert Trajectories
by: Mobarakeh, Niloufar Saeidi, et al.
Published: (2024)
by: Mobarakeh, Niloufar Saeidi, et al.
Published: (2024)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
Learning to Communicate Through Implicit Communication Channels
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
GradientStabilizer:Fix the Norm, Not the Gradient
by: Huang, Tianjin, et al.
Published: (2025)
by: Huang, Tianjin, et al.
Published: (2025)
ADG: Ambient Diffusion-Guided Dataset Recovery for Corruption-Robust Offline Reinforcement Learning
by: Liu, Zeyuan, et al.
Published: (2025)
by: Liu, Zeyuan, et al.
Published: (2025)
ContextEvolve: Multi-Agent Context Compression for Systems Code Optimization
by: Su, Hongyuan, et al.
Published: (2026)
by: Su, Hongyuan, et al.
Published: (2026)
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
by: Li, Yingru, et al.
Published: (2025)
by: Li, Yingru, et al.
Published: (2025)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
by: Yuan, Hui, et al.
Published: (2024)
by: Yuan, Hui, et al.
Published: (2024)
Stochastic Gradient Descent with Momentum is Algorithmically Stable
by: Lei, Yunwen, et al.
Published: (2026)
by: Lei, Yunwen, et al.
Published: (2026)
FORLER: Federated Offline Reinforcement Learning with Q-Ensemble and Actor Rectification
by: Qiao, Nan, et al.
Published: (2026)
by: Qiao, Nan, et al.
Published: (2026)
GSTM-HMU: Generative Spatio-Temporal Modeling for Human Mobility Understanding
by: Luo, Wenying, et al.
Published: (2025)
by: Luo, Wenying, et al.
Published: (2025)
Interpretable Hybrid-Rule Temporal Point Processes
by: Cao, Yunyang, et al.
Published: (2025)
by: Cao, Yunyang, et al.
Published: (2025)
Tackling Data Corruption in Offline Reinforcement Learning via Sequence Modeling
by: Xu, Jiawei, et al.
Published: (2024)
by: Xu, Jiawei, et al.
Published: (2024)
TaxDistill: Improving Metagenomic Taxonomic Annotation via Distilled Genomic Foundation Models
by: Ye, Rongye, et al.
Published: (2026)
by: Ye, Rongye, et al.
Published: (2026)
Factor Decorrelation Enhanced Data Removal from Deep Predictive Models
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Similar Items
-
Policy-Conditioned Policies for Multi-Agent Task Solving
by: Lin, Yue, et al.
Published: (2025) -
Information Bargaining: Bilateral Commitment in Bayesian Persuasion
by: Lin, Yue, et al.
Published: (2025) -
Learning to Negotiate via Voluntary Commitment
by: Zhu, Shuhui, et al.
Published: (2025) -
Carbon Market Simulation with Adaptive Mechanism Design
by: Wang, Han, et al.
Published: (2024) -
Verbalized Bayesian Persuasion
by: Li, Wenhao, et al.
Published: (2025)