Differential Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Tianzhu, Dong, Li, Xia, Yuqing, Sun, Yutao, Zhu, Yi, Huang, Gao, Wei, Furu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Universal YOCO for Efficient Depth Scaling
von: Sun, Yutao, et al.
Veröffentlicht: (2026)
von: Sun, Yutao, et al.
Veröffentlicht: (2026)
Rectified Sparse Attention
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
von: Sun, Yutao, et al.
Veröffentlicht: (2025)
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
Reinforcement Pre-Training
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025)
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025)
von: Wu, Xun, et al.
Veröffentlicht: (2025)
Thinking Augmented Pre-training
von: Wang, Liang, et al.
Veröffentlicht: (2025)
von: Wang, Liang, et al.
Veröffentlicht: (2025)
On-Policy RL with Optimal Reward Baseline
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
von: Hao, Yaru, et al.
Veröffentlicht: (2025)
On-Policy Context Distillation for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
Mixture of LoRA Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Online Experiential Learning for Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2026)
Learning Novel Transformer Architecture for Time-series Forecasting
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Juyuan, et al.
Veröffentlicht: (2025)
BitNet b1.58 2B4T Technical Report
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
Multiplex Thinking: Reasoning via Token-wise Branch-and-Merge
von: Tang, Yao, et al.
Veröffentlicht: (2026)
von: Tang, Yao, et al.
Veröffentlicht: (2026)
Fighting Spurious Correlations in Text Classification via a Causal Learning Perspective
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
Towards Stable and Effective Reinforcement Learning for Mixture-of-Experts
von: Zhang, Di, et al.
Veröffentlicht: (2025)
von: Zhang, Di, et al.
Veröffentlicht: (2025)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
von: Zhu, Dawei, et al.
Veröffentlicht: (2023)
Textual Aesthetics in Large Language Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2024)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2024)
Multi-Head Mixture-of-Experts
von: Wu, Xun, et al.
Veröffentlicht: (2024)
von: Wu, Xun, et al.
Veröffentlicht: (2024)
Scaling Optimal LR Across Token Horizons
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
von: Bjorck, Johan, et al.
Veröffentlicht: (2024)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Maximum Score Routing For Mixture-of-Experts
von: Dong, Bowen, et al.
Veröffentlicht: (2025)
von: Dong, Bowen, et al.
Veröffentlicht: (2025)
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
von: Ma, Shuming, et al.
Veröffentlicht: (2024)
On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability
von: Zheng, Chenyu, et al.
Veröffentlicht: (2024)
von: Zheng, Chenyu, et al.
Veröffentlicht: (2024)
Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models
von: Li, Zongqian, et al.
Veröffentlicht: (2026)
von: Li, Zongqian, et al.
Veröffentlicht: (2026)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
von: Ram, Dhananjay, et al.
Veröffentlicht: (2025)
Black-Box On-Policy Distillation of Large Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
von: Jiang, Ting, et al.
Veröffentlicht: (2024)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Trainable Transformer in Transformer
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2023)
MathScale: Scaling Instruction Tuning for Mathematical Reasoning
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Tang, Zhengyang, et al.
Veröffentlicht: (2024)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
Text Diffusion with Reinforced Conditioning
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Liu, Yuxuan, et al.
Veröffentlicht: (2024)
CAST: Continuous and Differentiable Semi-Structured Sparsity-Aware Training for Large Language Models
von: Huang, Weiyu, et al.
Veröffentlicht: (2025)
von: Huang, Weiyu, et al.
Veröffentlicht: (2025)
Deterministic Differentiable Structured Pruning for Large Language Models
von: Huang, Weiyu, et al.
Veröffentlicht: (2026)
von: Huang, Weiyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Universal YOCO for Efficient Depth Scaling
von: Sun, Yutao, et al.
Veröffentlicht: (2026) -
Rectified Sparse Attention
von: Sun, Yutao, et al.
Veröffentlicht: (2025) -
Multimodal Latent Language Modeling with Next-Token Diffusion
von: Sun, Yutao, et al.
Veröffentlicht: (2024) -
Reinforcement Pre-Training
von: Dong, Qingxiu, et al.
Veröffentlicht: (2025) -
BitNet v2: Native 4-bit Activations with Hadamard Transformation for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2025)