MOYU: A Theoretical Study on Massive Over-activation Yielded Uplifts in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Chi, Huang, Mincong, Wang, Chao, Wang, Yujie, Yu, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
by: Ma, Chi, et al.
Published: (2024)
by: Ma, Chi, et al.
Published: (2024)
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
by: Ma, Chi, et al.
Published: (2024)
by: Ma, Chi, et al.
Published: (2024)
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024)
by: Huang, Mincong, et al.
Published: (2024)
QQQ: Quality Quattuor-Bit Quantization for Large Language Models
by: Zhang, Ying, et al.
Published: (2024)
by: Zhang, Ying, et al.
Published: (2024)
Graph Neural Network with Two Uplift Estimators for Label-Scarcity Individual Uplift Modeling
by: Zhu, Dingyuan, et al.
Published: (2024)
by: Zhu, Dingyuan, et al.
Published: (2024)
Robustness-enhanced Uplift Modeling with Adversarial Feature Desensitization
by: Sun, Zexu, et al.
Published: (2023)
by: Sun, Zexu, et al.
Published: (2023)
UTBoost: Gradient Boosted Decision Trees for Uplift Modeling
by: Gao, Junjie, et al.
Published: (2023)
by: Gao, Junjie, et al.
Published: (2023)
A Comparative Study of Model Adaptation Strategies for Multi-Treatment Uplift Modeling
by: Zhang, Ruyue, et al.
Published: (2025)
by: Zhang, Ruyue, et al.
Published: (2025)
Orthogonal Uplift Learning with Permutation-Invariant Representations for Combinatorial Treatments
by: Su, Xinyan, et al.
Published: (2026)
by: Su, Xinyan, et al.
Published: (2026)
Rankability-enhanced Revenue Uplift Modeling Framework for Online Marketing
by: He, Bowei, et al.
Published: (2024)
by: He, Bowei, et al.
Published: (2024)
Less is More: on the Over-Globalizing Problem in Graph Transformers
by: Xing, Yujie, et al.
Published: (2024)
by: Xing, Yujie, et al.
Published: (2024)
Benchmarking for Deep Uplift Modeling in Online Marketing
by: Liu, Dugang, et al.
Published: (2024)
by: Liu, Dugang, et al.
Published: (2024)
Hierarchical Contextual Uplift Bandits for Catalog Personalization
by: Agrawal, Anupam, et al.
Published: (2026)
by: Agrawal, Anupam, et al.
Published: (2026)
Boosting Graph Robustness Against Backdoor Attacks: An Over-Similarity Perspective
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Evaluating Uplift Modeling under Structural Biases: Insights into Metric Stability and Model Robustness
by: Yang, Yuxuan, et al.
Published: (2026)
by: Yang, Yuxuan, et al.
Published: (2026)
Uplift modeling with continuous treatments: A predict-then-optimize approach
by: De Vos, Simon, et al.
Published: (2024)
by: De Vos, Simon, et al.
Published: (2024)
A New Transformation Approach for Uplift Modeling with Binary Outcome
by: Li, Kun, et al.
Published: (2023)
by: Li, Kun, et al.
Published: (2023)
FairUDT: Fairness-aware Uplift Decision Trees
by: Zahid, Anam, et al.
Published: (2025)
by: Zahid, Anam, et al.
Published: (2025)
A Theoretical Analysis of Self-Supervised Learning for Vision Transformers
by: Huang, Yu, et al.
Published: (2024)
by: Huang, Yu, et al.
Published: (2024)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Entire Chain Uplift Modeling with Context-Enhanced Learning for Intelligent Marketing
by: Huang, Yinqiu, et al.
Published: (2024)
by: Huang, Yinqiu, et al.
Published: (2024)
Can Slow-thinking LLMs Reason Over Time? Empirical Studies in Time Series Forecasting
by: Cheng, Mingyue, et al.
Published: (2025)
by: Cheng, Mingyue, et al.
Published: (2025)
Uplift Modeling Under Limited Supervision
by: Panagopoulos, George, et al.
Published: (2024)
by: Panagopoulos, George, et al.
Published: (2024)
Theoretical Analysis of Inductive Biases in Deep Convolutional Networks
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Which Company Adjustment Matter? Insights from Uplift Modeling on Financial Health
by: Wang, Xinlin, et al.
Published: (2025)
by: Wang, Xinlin, et al.
Published: (2025)
Is Inverse Reinforcement Learning Harder than Standard Reinforcement Learning? A Theoretical Perspective
by: Zhao, Lei, et al.
Published: (2023)
by: Zhao, Lei, et al.
Published: (2023)
Guardrailed Uplift Targeting: A Causal Optimization Playbook for Marketing Strategy
by: Sapru, Deepit
Published: (2025)
by: Sapru, Deepit
Published: (2025)
A Theoretical Study of Neural Network Expressive Power via Manifold Topology
by: Yao, Jiachen, et al.
Published: (2024)
by: Yao, Jiachen, et al.
Published: (2024)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
by: Wang, Jingyao, et al.
Published: (2025)
by: Wang, Jingyao, et al.
Published: (2025)
Polybasic Speculative Decoding Through a Theoretical Perspective
by: Wang, Ruilin, et al.
Published: (2025)
by: Wang, Ruilin, et al.
Published: (2025)
Robust Uplift Modeling with Large-Scale Contexts for Real-time Marketing
by: Sun, Zexu, et al.
Published: (2025)
by: Sun, Zexu, et al.
Published: (2025)
We Have It Covered: A Resampling-based Method for Uplift Model Comparison
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Jointly Optimizing Debiased CTR and Uplift for Coupons Marketing: A Unified Causal Framework
by: Yang, Siyun, et al.
Published: (2026)
by: Yang, Siyun, et al.
Published: (2026)
On the Convergence Analysis of Over-Parameterized Variational Autoencoders: A Neural Tangent Kernel Perspective
by: Wang, Li, et al.
Published: (2024)
by: Wang, Li, et al.
Published: (2024)
Direct Profit Estimation Using Uplift Modeling under Clustered Network Interference
by: Akker, Bram van den
Published: (2025)
by: Akker, Bram van den
Published: (2025)
ORGEval: Graph-Theoretic Evaluation of LLMs in Optimization Modeling
by: Wang, Zhuohan, et al.
Published: (2025)
by: Wang, Zhuohan, et al.
Published: (2025)
Secrets of GFlowNets' Learning Behavior: A Theoretical Study
by: Yu, Tianshu
Published: (2025)
by: Yu, Tianshu
Published: (2025)
Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm
by: Wei, Wen-Da, et al.
Published: (2026)
by: Wei, Wen-Da, et al.
Published: (2026)
TSC: A Simple Two-Sided Constraint against Over-Smoothing
by: Peng, Furong, et al.
Published: (2024)
by: Peng, Furong, et al.
Published: (2024)
Dataset Watermarking for Closed LLMs with Provable Detection
by: Huang, Pengrun, et al.
Published: (2026)
by: Huang, Pengrun, et al.
Published: (2026)
Similar Items
-
Dynamic Activation Pitfalls in LLaMA Models: An Empirical Study
by: Ma, Chi, et al.
Published: (2024) -
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
by: Ma, Chi, et al.
Published: (2024) -
Re-evaluating the Memory-balanced Pipeline Parallelism: BPipe
by: Huang, Mincong, et al.
Published: (2024) -
QQQ: Quality Quattuor-Bit Quantization for Large Language Models
by: Zhang, Ying, et al.
Published: (2024) -
Graph Neural Network with Two Uplift Estimators for Label-Scarcity Individual Uplift Modeling
by: Zhu, Dingyuan, et al.
Published: (2024)