DenseFormer: Enhancing Information Flow in Transformers via Depth Weighted Averaging
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pagliardini, Matteo, Mohtashami, Amirkeivan, Fleuret, Francois, Jaggi, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
Leveraging the true depth of LLMs
von: González, Ramón Calvo, et al.
Veröffentlicht: (2025)
von: González, Ramón Calvo, et al.
Veröffentlicht: (2025)
DoGE: Domain Reweighting with Generalization Estimation
von: Fan, Simin, et al.
Veröffentlicht: (2023)
von: Fan, Simin, et al.
Veröffentlicht: (2023)
Social Learning: Towards Collaborative Learning with Large Language Models
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023)
Benchmarking Optimizers for Large Language Model Pretraining
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)
Enhancing Multilingual LLM Pretraining with Model-Based Data Selection
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
von: Messmer, Bettina, et al.
Veröffentlicht: (2025)
DenseFormer: Learning Dense Depth Map from Sparse Depth and Image via Conditional Diffusion Model
von: Yuan, Ming, et al.
Veröffentlicht: (2025)
von: Yuan, Ming, et al.
Veröffentlicht: (2025)
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
Personalized Collaborative Fine-Tuning for On-Device Large Language Models
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
von: Wagner, Nicolas, et al.
Veröffentlicht: (2024)
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
von: Ashkboos, Saleh, et al.
Veröffentlicht: (2024)
Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
von: Makkuva, Ashok Vardhan, et al.
Veröffentlicht: (2024)
ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
von: Shao, Jintian, et al.
Veröffentlicht: (2025)
Towards an empirical understanding of MoE design choices
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
von: Fan, Dongyang, et al.
Veröffentlicht: (2024)
TyphoFormer: Language-Augmented Transformer for Accurate Typhoon Track Forecasting
von: Li, Lincan, et al.
Veröffentlicht: (2025)
von: Li, Lincan, et al.
Veröffentlicht: (2025)
The Journey Matters: Average Parameter Count over Pre-training Unifies Sparse and Dense Scaling Laws
von: Jin, Tian, et al.
Veröffentlicht: (2025)
von: Jin, Tian, et al.
Veröffentlicht: (2025)
WARM: On the Benefits of Weight Averaged Reward Models
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
von: Ramé, Alexandre, et al.
Veröffentlicht: (2024)
SecFormer: Fast and Accurate Privacy-Preserving Inference for Transformer Models via SMPC
von: Luo, Jinglong, et al.
Veröffentlicht: (2024)
von: Luo, Jinglong, et al.
Veröffentlicht: (2024)
MatFormer: Nested Transformer for Elastic Inference
von: Devvrit, et al.
Veröffentlicht: (2023)
von: Devvrit, et al.
Veröffentlicht: (2023)
Beyond Simple Averaging: Improving NLP Ensemble Performance with Topological-Data-Analysis-Based Weighting
von: Proskura, Polina, et al.
Veröffentlicht: (2024)
von: Proskura, Polina, et al.
Veröffentlicht: (2024)
The Free Transformer
von: Fleuret, François
Veröffentlicht: (2025)
von: Fleuret, François
Veröffentlicht: (2025)
RecurFormer: Not All Transformer Heads Need Self-Attention
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
von: Xiao, Da, et al.
Veröffentlicht: (2025)
von: Xiao, Da, et al.
Veröffentlicht: (2025)
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
von: London, Charles, et al.
Veröffentlicht: (2025)
von: London, Charles, et al.
Veröffentlicht: (2025)
Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging
von: Ablin, Pierre, et al.
Veröffentlicht: (2025)
von: Ablin, Pierre, et al.
Veröffentlicht: (2025)
Hierarchical Token Prepending: Enhancing Information Flow in Decoder-based LLM Embeddings
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
von: Ding, Xueying, et al.
Veröffentlicht: (2025)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
von: Fan, Dongyang, et al.
Veröffentlicht: (2025)
Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
von: Yang, Hongru, et al.
Veröffentlicht: (2024)
Weights to Code: Extracting Interpretable Algorithms from the Discrete Transformer
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
von: Zhang, Yifan, et al.
Veröffentlicht: (2026)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting
von: Weers, Alexander, et al.
Veröffentlicht: (2026)
von: Weers, Alexander, et al.
Veröffentlicht: (2026)
Do Transformers Use their Depth Adaptively? Evidence from a Relational Reasoning Task
von: Curth, Alicia, et al.
Veröffentlicht: (2026)
von: Curth, Alicia, et al.
Veröffentlicht: (2026)
Transformers on Markov Data: Constant Depth Suffices
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
von: Rajaraman, Nived, et al.
Veröffentlicht: (2024)
Mitigating Unintended Memorization with LoRA in Federated Learning for LLMs
von: Bossy, Thierry, et al.
Veröffentlicht: (2025)
von: Bossy, Thierry, et al.
Veröffentlicht: (2025)
FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
von: Nrusimha, Aniruddha, et al.
Veröffentlicht: (2025)
Contextual Graph Transformer: A Small Language Model for Enhanced Engineering Document Information Extraction
von: Reddy, Karan, et al.
Veröffentlicht: (2025)
von: Reddy, Karan, et al.
Veröffentlicht: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CoTFormer: A Chain-of-Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023) -
Leveraging the true depth of LLMs
von: González, Ramón Calvo, et al.
Veröffentlicht: (2025) -
DoGE: Domain Reweighting with Generalization Estimation
von: Fan, Simin, et al.
Veröffentlicht: (2023) -
Social Learning: Towards Collaborative Learning with Large Language Models
von: Mohtashami, Amirkeivan, et al.
Veröffentlicht: (2023) -
Benchmarking Optimizers for Large Language Model Pretraining
von: Semenov, Andrei, et al.
Veröffentlicht: (2025)