Provable Benefit of Cutout and CutMix for Feature Learning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Oh, Junsoo, Yun, Chulhee |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
par: Oh, Junsoo, et autres
Publié: (2025)
par: Oh, Junsoo, et autres
Publié: (2025)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
par: Shin, Baekrok, et autres
Publié: (2024)
par: Shin, Baekrok, et autres
Publié: (2024)
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
par: Song, Minhak, et autres
Publié: (2025)
par: Song, Minhak, et autres
Publié: (2025)
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
par: Moon, Chaewon, et autres
Publié: (2026)
par: Moon, Chaewon, et autres
Publié: (2026)
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
par: Kim, Yujun, et autres
Publié: (2025)
par: Kim, Yujun, et autres
Publié: (2025)
Provable Benefits of In-Tool Learning for Large Language Models
par: Houliston, Sam, et autres
Publié: (2025)
par: Houliston, Sam, et autres
Publié: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
par: Cho, Hanseul, et autres
Publié: (2024)
par: Cho, Hanseul, et autres
Publié: (2024)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
par: Cho, Yeseul, et autres
Publié: (2025)
par: Cho, Yeseul, et autres
Publié: (2025)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
par: Kim, Jihwan, et autres
Publié: (2026)
par: Kim, Jihwan, et autres
Publié: (2026)
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
par: Baek, Beomhan, et autres
Publié: (2025)
par: Baek, Beomhan, et autres
Publié: (2025)
Provable Long-Range Benefits of Next-Token Prediction
par: Cao, Xinyuan, et autres
Publié: (2025)
par: Cao, Xinyuan, et autres
Publié: (2025)
Parameter Expanded Stochastic Gradient Markov Chain Monte Carlo
par: Kim, Hyunsu, et autres
Publié: (2025)
par: Kim, Hyunsu, et autres
Publié: (2025)
Provable Effects of Data Replay in Continual Learning: A Feature Learning Perspective
par: Ding, Meng, et autres
Publié: (2026)
par: Ding, Meng, et autres
Publié: (2026)
Provable Scaling Laws of Feature Emergence from Learning Dynamics of Grokking
par: Tian, Yuandong
Publié: (2025)
par: Tian, Yuandong
Publié: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
par: Kim, Juno, et autres
Publié: (2025)
par: Kim, Juno, et autres
Publié: (2025)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
par: Gaitonde, Jason, et autres
Publié: (2026)
par: Gaitonde, Jason, et autres
Publié: (2026)
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
par: Lee, Joongkyu, et autres
Publié: (2025)
par: Lee, Joongkyu, et autres
Publié: (2025)
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
par: Yadav, Robin, et autres
Publié: (2025)
par: Yadav, Robin, et autres
Publié: (2025)
Provable Benefits of Complex Parameterizations for Structured State Space Models
par: Ran-Milo, Yuval, et autres
Publié: (2024)
par: Ran-Milo, Yuval, et autres
Publié: (2024)
Privacy-Preserving Split Learning with Vision Transformers using Patch-Wise Random and Noisy CutMix
par: Oh, Seungeun, et autres
Publié: (2024)
par: Oh, Seungeun, et autres
Publié: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
par: Cho, Hanseul, et autres
Publié: (2024)
par: Cho, Hanseul, et autres
Publié: (2024)
Linear attention is (maybe) all you need (to understand transformer optimization)
par: Ahn, Kwangjun, et autres
Publié: (2023)
par: Ahn, Kwangjun, et autres
Publié: (2023)
Joint Embedding vs Reconstruction: Provable Benefits of Latent Space Prediction for Self Supervised Learning
par: Van Assel, Hugues, et autres
Publié: (2025)
par: Van Assel, Hugues, et autres
Publié: (2025)
Stochastic Extragradient with Flip-Flop Shuffling & Anchoring: Provable Improvements
par: Chae, Jiseok, et autres
Publié: (2024)
par: Chae, Jiseok, et autres
Publié: (2024)
Provably Efficient Exploration in Inverse Constrained Reinforcement Learning
par: Yue, Bo, et autres
Publié: (2024)
par: Yue, Bo, et autres
Publié: (2024)
The Path Not Taken: RLVR Provably Learns Off the Principals
par: Zhu, Hanqing, et autres
Publié: (2025)
par: Zhu, Hanqing, et autres
Publié: (2025)
Provable Zero-Shot Generalization in Offline Reinforcement Learning
par: Wang, Zhiyong, et autres
Publié: (2025)
par: Wang, Zhiyong, et autres
Publié: (2025)
Oldie but Goodie: Re-illuminating Label Propagation on Graphs with Partially Observed Features
par: Yun, Sukwon, et autres
Publié: (2025)
par: Yun, Sukwon, et autres
Publié: (2025)
Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
par: Kar, Avik, et autres
Publié: (2024)
par: Kar, Avik, et autres
Publié: (2024)
Dynamic Model Predictive Shielding for Provably Safe Reinforcement Learning
par: Banerjee, Arko, et autres
Publié: (2024)
par: Banerjee, Arko, et autres
Publié: (2024)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
par: Hu, Jerry Yao-Chieh, et autres
Publié: (2025)
Deciphering Raw Data in Neuro-Symbolic Learning with Provable Guarantees
par: Tao, Lue, et autres
Publié: (2023)
par: Tao, Lue, et autres
Publié: (2023)
Transformers Provably Implement In-Context Reinforcement Learning with Policy Improvement
par: Liang, Haodong, et autres
Publié: (2026)
par: Liang, Haodong, et autres
Publié: (2026)
Provable Representation with Efficient Planning for Partial Observable Reinforcement Learning
par: Zhang, Hongming, et autres
Publié: (2023)
par: Zhang, Hongming, et autres
Publié: (2023)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
par: Oh, Junsoo, et autres
Publié: (2025)
par: Oh, Junsoo, et autres
Publié: (2025)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
par: Chen, Siyu, et autres
Publié: (2024)
par: Chen, Siyu, et autres
Publié: (2024)
Fundamental Benefit of Alternating Updates in Minimax Optimization
par: Lee, Jaewook, et autres
Publié: (2024)
par: Lee, Jaewook, et autres
Publié: (2024)
Provably Extracting the Features from a General Superposition
par: Liu, Allen
Publié: (2025)
par: Liu, Allen
Publié: (2025)
Cut Less, Fold More: Model Compression through the Lens of Projection Geometry
par: Saukh, Olga, et autres
Publié: (2026)
par: Saukh, Olga, et autres
Publié: (2026)
Provably Efficient Action-Manipulation Attack Against Continuous Reinforcement Learning
par: Luo, Zhi, et autres
Publié: (2024)
par: Luo, Zhi, et autres
Publié: (2024)
Documents similaires
-
From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature Learning
par: Oh, Junsoo, et autres
Publié: (2025) -
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
par: Shin, Baekrok, et autres
Publié: (2024) -
Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model Training
par: Song, Minhak, et autres
Publié: (2025) -
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
par: Moon, Chaewon, et autres
Publié: (2026) -
The Cost of Robustness: Tighter Bounds on Parameter Complexity for Robust Memorization in ReLU Nets
par: Kim, Yujun, et autres
Publié: (2025)