Entire Space Counterfactual Learning for Reliable Content Recommendations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Hao, Chen, Zhichao, Liu, Zhaoran, Li, Haozhe, Yang, Degui, Liu, Xinggao, Li, Haoxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FreDF: Learning to Forecast in the Frequency Domain
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
von: Wang, Hao, et al.
Veröffentlicht: (2024)
von: Wang, Hao, et al.
Veröffentlicht: (2024)
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
von: Wang, Hao, et al.
Veröffentlicht: (2023)
von: Wang, Hao, et al.
Veröffentlicht: (2023)
Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting
von: Pan, Licheng, et al.
Veröffentlicht: (2025)
von: Pan, Licheng, et al.
Veröffentlicht: (2025)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
DeepFilter: A Transformer-style Framework for Accurate and Efficient Process Monitoring
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
ECAT: A Entire space Continual and Adaptive Transfer Learning Framework for Cross-Domain Recommendation
von: Hou, Chaoqun, et al.
Veröffentlicht: (2024)
von: Hou, Chaoqun, et al.
Veröffentlicht: (2024)
Enabling Agents to Communicate Entirely in Latent Space
von: Du, Zhuoyun, et al.
Veröffentlicht: (2025)
von: Du, Zhuoyun, et al.
Veröffentlicht: (2025)
From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive Optimization
von: Chen, Xinjie, et al.
Veröffentlicht: (2025)
von: Chen, Xinjie, et al.
Veröffentlicht: (2025)
Deep Time-series Forecasting Needs Kernelized Moment Balancing
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
von: Pan, Licheng, et al.
Veröffentlicht: (2026)
Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Estimating the Effects of Sample Training Orders for Large Language Models without Retraining
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
Task Priors: Enhancing Model Evaluation by Considering the Entire Space of Downstream Tasks
von: Patel, Niket, et al.
Veröffentlicht: (2025)
von: Patel, Niket, et al.
Veröffentlicht: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
Time-o1: Time-Series Forecasting Needs Transformed Label Alignment
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
Reason for Future, Act for Now: A Principled Framework for Autonomous LLM Agents with Provable Sample Efficiency
von: Liu, Zhihan, et al.
Veröffentlicht: (2023)
von: Liu, Zhihan, et al.
Veröffentlicht: (2023)
Entire Chain Uplift Modeling with Context-Enhanced Learning for Intelligent Marketing
von: Huang, Yinqiu, et al.
Veröffentlicht: (2024)
von: Huang, Yinqiu, et al.
Veröffentlicht: (2024)
Embed to Control Partially Observed Systems: Representation Learning with Provable Sample Efficiency
von: Wang, Lingxiao, et al.
Veröffentlicht: (2022)
von: Wang, Lingxiao, et al.
Veröffentlicht: (2022)
Self-Distilled Disentangled Learning for Counterfactual Prediction
von: Li, Xinshu, et al.
Veröffentlicht: (2024)
von: Li, Xinshu, et al.
Veröffentlicht: (2024)
IB-GRPO: Aligning LLM-based Learning Path Recommendation with Educational Objectives via Indicator-Based Group Relative Policy Optimization
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
von: Wang, Shuai, et al.
Veröffentlicht: (2026)
FDRMFL:Multi-modal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning
von: Wu, Haozhe
Veröffentlicht: (2025)
von: Wu, Haozhe
Veröffentlicht: (2025)
DDTime: Dataset Distillation with Spectral Alignment and Information Bottleneck for Time-Series Forecasting
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Revisiting Counterfactual Regression through the Lens of Gromov-Wasserstein Information Bottleneck
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
von: Zhang, Shenao, et al.
Veröffentlicht: (2024)
Latent-Space Contrastive Reinforcement Learning for Stable and Efficient LLM Reasoning
von: Shan, Lianlei, et al.
Veröffentlicht: (2026)
von: Shan, Lianlei, et al.
Veröffentlicht: (2026)
DKINet: Medication Recommendation via Domain Knowledge Informed Deep Learning
von: Liu, Sicen, et al.
Veröffentlicht: (2023)
von: Liu, Sicen, et al.
Veröffentlicht: (2023)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
von: Yin, Yutong, et al.
Veröffentlicht: (2025)
von: Yin, Yutong, et al.
Veröffentlicht: (2025)
HiPER: Hierarchical Reinforcement Learning with Explicit Credit Assignment for Large Language Model Agents
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
von: Peng, Jiangweizhi, et al.
Veröffentlicht: (2026)
Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation
von: Li, He, et al.
Veröffentlicht: (2025)
von: Li, He, et al.
Veröffentlicht: (2025)
Benchmarking Counterfactual Interpretability in Deep Learning Models for Time Series Classification
von: Kan, Ziwen, et al.
Veröffentlicht: (2024)
von: Kan, Ziwen, et al.
Veröffentlicht: (2024)
Can a Single Tree Outperform an Entire Forest?
von: Mao, Qiangqiang, et al.
Veröffentlicht: (2024)
von: Mao, Qiangqiang, et al.
Veröffentlicht: (2024)
LeapFactual: Reliable Visual Counterfactual Explanation Using Conditional Flow Matching
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
von: Cao, Zhuo, et al.
Veröffentlicht: (2025)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
von: Liu, Zhihan, et al.
Veröffentlicht: (2024)
Language Models for Controllable DNA Sequence Design
von: Su, Xingyu, et al.
Veröffentlicht: (2025)
von: Su, Xingyu, et al.
Veröffentlicht: (2025)
Contextual Dynamic Pricing with Strategic Buyers
von: Liu, Pangpang, et al.
Veröffentlicht: (2023)
von: Liu, Pangpang, et al.
Veröffentlicht: (2023)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
von: Lu, Miao, et al.
Veröffentlicht: (2022)
von: Lu, Miao, et al.
Veröffentlicht: (2022)
Decentralized Autoregressive Generation
von: Maschan, Stepan, et al.
Veröffentlicht: (2026)
von: Maschan, Stepan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FreDF: Learning to Forecast in the Frequency Domain
von: Wang, Hao, et al.
Veröffentlicht: (2024) -
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
von: Wang, Hao, et al.
Veröffentlicht: (2024) -
An Accurate and Interpretable Framework for Trustworthy Process Monitoring
von: Wang, Hao, et al.
Veröffentlicht: (2023) -
Mixture of Low Rank Adaptation with Partial Parameter Sharing for Time Series Forecasting
von: Pan, Licheng, et al.
Veröffentlicht: (2025) -
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
von: Wang, Hao, et al.
Veröffentlicht: (2026)