MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
Fuente:
arXiv
Salvato in:
| Autori principali: | Xiao, Da, Meng, Qingye, Li, Shengping, Yuan, Xingyuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024)
di: Xiao, Da, et al.
Pubblicazione: (2024)
CNCast: Leveraging 3D Swin Transformer and DiT for Enhanced Regional Weather Forecasting
di: Liang, Hongli, et al.
Pubblicazione: (2025)
di: Liang, Hongli, et al.
Pubblicazione: (2025)
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026)
di: Jian, Yichang, et al.
Pubblicazione: (2026)
Residual Stream Duality in Modern Transformer Architectures
di: Zhang, Yifan
Pubblicazione: (2026)
di: Zhang, Yifan
Pubblicazione: (2026)
Revisiting the Shape Convention of Transformer Language Models
di: Liao, Feng-Ting, et al.
Pubblicazione: (2026)
di: Liao, Feng-Ting, et al.
Pubblicazione: (2026)
Policy Learning with a Language Bottleneck
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
di: Srivastava, Megha, et al.
Pubblicazione: (2024)
Learning Adapter Rank via Symmetry Breaking
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
di: Doyle, Cooper, et al.
Pubblicazione: (2025)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
di: Yuan, Yurun, et al.
Pubblicazione: (2026)
di: Yuan, Yurun, et al.
Pubblicazione: (2026)
Are Transformers Able to Reason by Connecting Separated Knowledge in Training Data?
di: Yin, Yutong, et al.
Pubblicazione: (2025)
di: Yin, Yutong, et al.
Pubblicazione: (2025)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
Solo Connection: A Parameter Efficient Fine-Tuning Technique for Transformers
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
di: Pathak, Harsh Nilesh, et al.
Pubblicazione: (2025)
Language Bottleneck Models for Qualitative Knowledge State Modeling
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
di: Berthon, Antonin, et al.
Pubblicazione: (2025)
CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
di: Pan, Teng, et al.
Pubblicazione: (2026)
di: Pan, Teng, et al.
Pubblicazione: (2026)
PCDCNet: A Surrogate Model for Air Quality Forecasting with Physical-Chemical Dynamics and Constraints
di: Wang, Shuo, et al.
Pubblicazione: (2025)
di: Wang, Shuo, et al.
Pubblicazione: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
di: Zhao, Weilin, et al.
Pubblicazione: (2025)
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
di: Guo, Yiran, et al.
Pubblicazione: (2026)
di: Guo, Yiran, et al.
Pubblicazione: (2026)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
di: Hui, Tingfeng, et al.
Pubblicazione: (2024)
JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
di: Tian, Yuandong, et al.
Pubblicazione: (2023)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
Efficient Mathematical Reasoning Models via Dynamic Pruning and Knowledge Distillation
di: Yu, Fengming, et al.
Pubblicazione: (2025)
di: Yu, Fengming, et al.
Pubblicazione: (2025)
Trajectory Bellman Residual Minimization: A Simple Value-Based Method for LLM Reasoning
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
di: Yuan, Yurun, et al.
Pubblicazione: (2025)
ABBEL: LLM Agents Acting through Belief Bottlenecks Expressed in Language
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
di: Lidayan, Aly, et al.
Pubblicazione: (2025)
Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2026)
di: Renduchintala, H S V N S Kowndinya, et al.
Pubblicazione: (2026)
LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM Safety
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
di: Yang, Junxiao, et al.
Pubblicazione: (2026)
MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate
di: Wang, Jianze, et al.
Pubblicazione: (2026)
di: Wang, Jianze, et al.
Pubblicazione: (2026)
Dense SAE Latents Are Features, Not Bugs
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
Frac-Connections: Fractional Extension of Hyper-Connections
di: Zhu, Defa, et al.
Pubblicazione: (2025)
di: Zhu, Defa, et al.
Pubblicazione: (2025)
Dynamic Universal Approximation Theory: The Basic Theory for Transformer-based Large Language Models
di: Wang, Wei, et al.
Pubblicazione: (2024)
di: Wang, Wei, et al.
Pubblicazione: (2024)
Benchmarking and Understanding Compositional Relational Reasoning of LLMs
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
di: Ni, Ruikang, et al.
Pubblicazione: (2024)
Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models
di: Yin, Huifeng, et al.
Pubblicazione: (2025)
di: Yin, Huifeng, et al.
Pubblicazione: (2025)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
di: Ildiz, M. Emrullah, et al.
Pubblicazione: (2024)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
di: Luo, Sijia, et al.
Pubblicazione: (2026)
di: Luo, Sijia, et al.
Pubblicazione: (2026)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
di: Li, Jin, et al.
Pubblicazione: (2025)
di: Li, Jin, et al.
Pubblicazione: (2025)
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
di: Shi, Jiajun, et al.
Pubblicazione: (2025)
mHC: Manifold-Constrained Hyper-Connections
di: Xie, Zhenda, et al.
Pubblicazione: (2025)
di: Xie, Zhenda, et al.
Pubblicazione: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2026)
di: Moosa, Ibraheem Muhammad, et al.
Pubblicazione: (2026)
On the Robustness of Transformers against Context Hijacking for Linear Classification
di: Li, Tianle, et al.
Pubblicazione: (2025)
di: Li, Tianle, et al.
Pubblicazione: (2025)
Deconstructing Pre-training: Knowledge Attribution Analysis in MoE and Dense Models
di: Wang, Bo, et al.
Pubblicazione: (2026)
di: Wang, Bo, et al.
Pubblicazione: (2026)
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
di: Gong, Zixuan, et al.
Pubblicazione: (2025)
Neutral Residues: Revisiting Adapters for Model Extension
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
di: Talla, Franck Signe, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Improving Transformers with Dynamically Composable Multi-Head Attention
di: Xiao, Da, et al.
Pubblicazione: (2024) -
CNCast: Leveraging 3D Swin Transformer and DiT for Enhanced Regional Weather Forecasting
di: Liang, Hongli, et al.
Pubblicazione: (2025) -
Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Induction
di: Jian, Yichang, et al.
Pubblicazione: (2026) -
Residual Stream Duality in Modern Transformer Architectures
di: Zhang, Yifan
Pubblicazione: (2026) -
Revisiting the Shape Convention of Transformer Language Models
di: Liao, Feng-Ting, et al.
Pubblicazione: (2026)