Mixture of Hidden-Dimensions Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yilong, Shang, Junyuan, Zhang, Zhengyu, Sheng, Jiawei, Liu, Tingwen, Wang, Shuohuan, Sun, Yu, Wu, Hua, Wang, Haifeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026)
by: Chen, Yilong, et al.
Published: (2026)
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
by: Chen, Yilong, et al.
Published: (2024)
by: Chen, Yilong, et al.
Published: (2024)
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
MoR: Mixture of Ranks for Low-Rank Adaptation Tuning
by: Tang, Chuanyu, et al.
Published: (2024)
by: Tang, Chuanyu, et al.
Published: (2024)
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
Improving Reasoning Capabilities in Small Models through Mixture-of-Layers Distillation with Stepwise Attention on Key Information
by: Chen, Yao, et al.
Published: (2026)
by: Chen, Yao, et al.
Published: (2026)
Curiosity-Driven Reinforcement Learning from Human Feedback
by: Sun, Haoran, et al.
Published: (2025)
by: Sun, Haoran, et al.
Published: (2025)
Advantageous Parameter Expansion Training Makes Better Large Language Models
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
by: Ji, Longbin, et al.
Published: (2026)
by: Ji, Longbin, et al.
Published: (2026)
BeamLoRA: Beam-Constraint Low-Rank Adaptation
by: Gu, Naibin, et al.
Published: (2025)
by: Gu, Naibin, et al.
Published: (2025)
HFT: Half Fine-Tuning for Large Language Models
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Autoregressive Pre-Training on Pixels and Texts
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
On Training Data Influence of GPT Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Tool-Augmented Reward Modeling
by: Li, Lei, et al.
Published: (2023)
by: Li, Lei, et al.
Published: (2023)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
by: Su, Taoyu, et al.
Published: (2024)
by: Su, Taoyu, et al.
Published: (2024)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
by: Su, Taoyu, et al.
Published: (2024)
by: Su, Taoyu, et al.
Published: (2024)
Optimal Transport Guided Correlation Assignment for Multimodal Entity Linking
by: Zhang, Zefeng, et al.
Published: (2024)
by: Zhang, Zefeng, et al.
Published: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
Revealing and Mitigating the Challenge of Detecting Character Knowledge Errors in LLM Role-Playing
by: Zhang, Wenyuan, et al.
Published: (2024)
by: Zhang, Wenyuan, et al.
Published: (2024)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
by: Sun, Shaoning, et al.
Published: (2026)
by: Sun, Shaoning, et al.
Published: (2026)
ATTNPO: Attention-Guided Process Supervision for Efficient Reasoning
by: Nie, Shuaiyi, et al.
Published: (2026)
by: Nie, Shuaiyi, et al.
Published: (2026)
Enhancing Multimodal Entity and Relation Extraction with Variational Information Bottleneck
by: Cui, Shiyao, et al.
Published: (2023)
by: Cui, Shiyao, et al.
Published: (2023)
Maximum Score Routing For Mixture-of-Experts
by: Dong, Bowen, et al.
Published: (2025)
by: Dong, Bowen, et al.
Published: (2025)
LLM Hallucination Detection: A Fast Fourier Transform Method Based on Hidden Layer Temporal Signals
by: Li, Jinxin, et al.
Published: (2025)
by: Li, Jinxin, et al.
Published: (2025)
Towards Boosting Many-to-Many Multilingual Machine Translation with Large Language Models
by: Gao, Pengzhi, et al.
Published: (2024)
by: Gao, Pengzhi, et al.
Published: (2024)
HyperMem: Hypergraph Memory for Long-Term Conversations
by: Yue, Juwei, et al.
Published: (2026)
by: Yue, Juwei, et al.
Published: (2026)
AIMDiT: Modality Augmentation and Interaction via Multimodal Dimension Transformation for Emotion Recognition in Conversations
by: Wu, Sheng, et al.
Published: (2024)
by: Wu, Sheng, et al.
Published: (2024)
A Survey on Parallel Reasoning
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions
by: Pan, Wenbo, et al.
Published: (2025)
by: Pan, Wenbo, et al.
Published: (2025)
Debiasing Multimodal Large Language Models via Noise-Aware Preference Optimization
by: Zhang, Zefeng, et al.
Published: (2025)
by: Zhang, Zefeng, et al.
Published: (2025)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
Seed-Guided Topic Discovery with Out-of-Vocabulary Seeds
by: Zhang, Yu, et al.
Published: (2022)
by: Zhang, Yu, et al.
Published: (2022)
Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal Perspective
by: Su, Taoyu, et al.
Published: (2025)
by: Su, Taoyu, et al.
Published: (2025)
BEE-RAG: Balanced Entropy Engineering for Retrieval-Augmented Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Unveiling Knowledge Utilization Mechanisms in LLM-based Retrieval-Augmented Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Similar Items
-
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
by: Chen, Yilong, et al.
Published: (2026) -
Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking
by: Chen, Yilong, et al.
Published: (2025) -
DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
by: Chen, Yilong, et al.
Published: (2024) -
NACL: A General and Effective KV Cache Eviction Framework for LLMs at Inference Time
by: Chen, Yilong, et al.
Published: (2024) -
Sparse Growing Transformer: Training-Time Sparse Depth Allocation via Progressive Attention Looping
by: Chen, Yao, et al.
Published: (2026)