mHC: Manifold-Constrained Hyper-Connections
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Zhenda, Wei, Yixuan, Cao, Huanqi, Zhao, Chenggang, Deng, Chengqi, Li, Jiashi, Dai, Damai, Gao, Huazuo, Chang, Jiang, Yu, Kuai, Zhao, Liang, Zhou, Shangyan, Xu, Zhean, Zhang, Zhengyan, Zeng, Wangding, Hu, Shengding, Wang, Yuqing, Yuan, Jingyang, Wang, Lean, Liang, Wenfeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025)
by: Yuan, Jingyang, et al.
Published: (2025)
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
by: Wang, Lean, et al.
Published: (2024)
by: Wang, Lean, et al.
Published: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025)
by: Zhao, Chenggang, et al.
Published: (2025)
mHC-GNN: Manifold-Constrained Hyper-Connections for Graph Neural Networks
by: Mishra, Subhankar
Published: (2026)
by: Mishra, Subhankar
Published: (2026)
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
by: Dai, Damai, et al.
Published: (2024)
by: Dai, Damai, et al.
Published: (2024)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
by: Mutlu, Abdulvahap, et al.
Published: (2026)
by: Mutlu, Abdulvahap, et al.
Published: (2026)
mHC-HSI: Clustering-Guided Hyper-Connection Mamba for Hyperspectral Image Classification
by: Zhu, Yimin, et al.
Published: (2026)
by: Zhu, Yimin, et al.
Published: (2026)
mHC-lite: You Don't Need 20 Sinkhorn-Knopp Iterations
by: Yang, Yongyi, et al.
Published: (2026)
by: Yang, Yongyi, et al.
Published: (2026)
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
by: Cheng, Xin, et al.
Published: (2026)
by: Cheng, Xin, et al.
Published: (2026)
White-Box mHC: Electromagnetic Spectrum-Aware and Interpretable Stream Interactions for Hyperspectral Image Classification
by: Zhu, Yimin, et al.
Published: (2026)
by: Zhu, Yimin, et al.
Published: (2026)
KromHC: Manifold-Constrained Hyper-Connections with Kronecker-Product Residual Matrices
by: Zhou, Wuyang, et al.
Published: (2026)
by: Zhou, Wuyang, et al.
Published: (2026)
JPmHC Dynamical Isometry via Orthogonal Hyper-Connections
by: Sengupta, Biswa, et al.
Published: (2026)
by: Sengupta, Biswa, et al.
Published: (2026)
go-$m$HC: Direct Parameterization of Manifold-Constrained Hyper-Connections via Generalized Orthostochastic Matrices
by: Dandachi, Torque, et al.
Published: (2026)
by: Dandachi, Torque, et al.
Published: (2026)
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
by: DeepSeek-AI, et al.
Published: (2024)
by: DeepSeek-AI, et al.
Published: (2024)
Quantum Data Structure for Range Minimum Query
by: Wang, Qisheng, et al.
Published: (2026)
by: Wang, Qisheng, et al.
Published: (2026)
Simple and Faster Algorithms for Knapsack
by: He, Qizheng, et al.
Published: (2023)
by: He, Qizheng, et al.
Published: (2023)
HC-GST: Heterophily-aware Distribution Consistency based Graph Self-training
by: Wang, Fali, et al.
Published: (2024)
by: Wang, Fali, et al.
Published: (2024)
ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models
by: Song, Chenyang, et al.
Published: (2024)
by: Song, Chenyang, et al.
Published: (2024)
Innentitelbild: Activating and Stabilizing a Reversible four Electron Redox Reaction of I−/I+ for Aqueous Zn‐Iodine Battery (Angew. Chem. 25/2024)
by: Chenggang Wang, et al.
Published: (2024)
by: Chenggang Wang, et al.
Published: (2024)
Activating and Stabilizing a Reversible four Electron Redox Reaction of I−/I+ for Aqueous Zn‐Iodine Battery
by: Chenggang Wang, et al.
Published: (2024)
by: Chenggang Wang, et al.
Published: (2024)
Exploring Activation Patterns of Parameters in Language Models
by: Wang, Yudong, et al.
Published: (2024)
by: Wang, Yudong, et al.
Published: (2024)
Recursive Optimal Stopping with Poisson Stopping Constraints
by: Liang, Gechun, et al.
Published: (2024)
by: Liang, Gechun, et al.
Published: (2024)
Matrix Fejér-Riesz type theorem for a union of an interval and a point
by: Sun, Shengding, et al.
Published: (2025)
by: Sun, Shengding, et al.
Published: (2025)
VULCAN: Vision-Language-Model Enhanced Multi-Agent Cooperative Navigation for Indoor Fire-Disaster Response
by: Liu, Shengding, et al.
Published: (2026)
by: Liu, Shengding, et al.
Published: (2026)
Nonfundamental‐Driven Price Shocks and Corporate Climate Risk Disclosure
by: Hu Wang, et al.
Published: (2026)
by: Hu Wang, et al.
Published: (2026)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
LLMKey: LLM-Powered Wireless Key Generation Scheme for Next-Gen IoV Systems
by: Yang, Huanqi, et al.
Published: (2025)
by: Yang, Huanqi, et al.
Published: (2025)
Monte Carlo Tree Search for Execution-Guided Program Repair with Large Language Models
by: Liang, Yixuan
Published: (2026)
by: Liang, Yixuan
Published: (2026)
Investigating Chinese Parents' Growth Mindset and Their Parenting Practices in Children's English Learning
by: Chenggang Liang, et al.
Published: (2025)
by: Chenggang Liang, et al.
Published: (2025)
Enhanced Approximation Algorithms for the Capacitated Location Routing Problem
by: Zhao, Jingyang, et al.
Published: (2025)
by: Zhao, Jingyang, et al.
Published: (2025)
FAME: Forecasting Academic Impact via Continuous-Time Manifold Evolution
by: Ding, Jianrong, et al.
Published: (2026)
by: Ding, Jianrong, et al.
Published: (2026)
Metacognitive Prompting Improves Understanding in Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
TRAM: Benchmarking Temporal Reasoning for Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
by: Wu, Zhiyu, et al.
Published: (2024)
by: Wu, Zhiyu, et al.
Published: (2024)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
by: Wang, Yuqing, et al.
Published: (2024)
by: Wang, Yuqing, et al.
Published: (2024)
The spin measurement of the black hole SLX 1746-331 using Insight-HXMT observations
by: Chen, Jiashi, et al.
Published: (2025)
by: Chen, Jiashi, et al.
Published: (2025)
Quasi-periodic oscillations and reflection feature evolution in 4U 1630-47 observed with Insight-HXMT
by: Chen, Jiashi, et al.
Published: (2025)
by: Chen, Jiashi, et al.
Published: (2025)
QUIJOTE discovery of the cation radicals HC5N+ and HC7N+
by: Cernicharo, J., et al.
Published: (2024)
by: Cernicharo, J., et al.
Published: (2024)
Similar Items
-
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
by: Yuan, Jingyang, et al.
Published: (2025) -
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
by: Wang, Lean, et al.
Published: (2024) -
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025) -
mHC-GNN: Manifold-Constrained Hyper-Connections for Graph Neural Networks
by: Mishra, Subhankar
Published: (2026) -
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
by: Dai, Damai, et al.
Published: (2024)