Attention Residuals
Fuente:
arXiv
Salvato in:
| Autori principali: | Kimi Team, Chen, Guangyu, Zhang, Yu, Su, Jianlin, Xu, Weixin, Pan, Siyuan, Wang, Yaoyu, Wang, Yucheng, Chen, Guanduo, Yin, Bohong, Chen, Yutian, Yan, Junjie, Wei, Ming, Zhang, Y., Meng, Fanqing, Hong, Chao, Xie, Xiaotong, Liu, Shaowei, Lu, Enzhe, Tai, Yunpeng, Chen, Yanru, Men, Xin, Guo, Haiqing, Charles, Y., Lu, Haoyu, Sui, Lin, Zhu, Jinguo, Zhou, Zaida, He, Weiran, Huang, Weixiao, Xu, Xinran, Wang, Yuzhi, Lai, Guokun, Du, Yulun, Wu, Yuxin, Yang, Zhilin, Zhou, Xinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Muon is Scalable for LLM Training
di: Liu, Jingyuan, et al.
Pubblicazione: (2025)
di: Liu, Jingyuan, et al.
Pubblicazione: (2025)
MoBA: Mixture of Block Attention for Long-Context LLMs
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
di: Lu, Enzhe, et al.
Pubblicazione: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
di: Kimi Team, et al.
Pubblicazione: (2025)
di: Kimi Team, et al.
Pubblicazione: (2025)
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
di: Zhao, Pukun, et al.
Pubblicazione: (2025)
di: Zhao, Pukun, et al.
Pubblicazione: (2025)
STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO
di: Zhao, Pukun, et al.
Pubblicazione: (2026)
di: Zhao, Pukun, et al.
Pubblicazione: (2026)
On the geometric side of the Jacquet-Rallis relative trace formula
di: Lu, Weixiao
Pubblicazione: (2024)
di: Lu, Weixiao
Pubblicazione: (2024)
Platform investment and creators' quality choice
di: Limei Chen, et al.
Pubblicazione: (2024)
di: Limei Chen, et al.
Pubblicazione: (2024)
Poincaré type J-equation
di: Chen, Xiuxiong, et al.
Pubblicazione: (2026)
di: Chen, Xiuxiong, et al.
Pubblicazione: (2026)
Robust and Efficient Adversarial Defense in SNNs via Image Purification and Joint Detection
di: Chen, Weiran, et al.
Pubblicazione: (2024)
di: Chen, Weiran, et al.
Pubblicazione: (2024)
Simulating Field Experiments with Large Language Models
di: Chen, Yaoyu, et al.
Pubblicazione: (2024)
di: Chen, Yaoyu, et al.
Pubblicazione: (2024)
Predicting Field Experiments with Large Language Models
di: Chen, Yaoyu, et al.
Pubblicazione: (2025)
di: Chen, Yaoyu, et al.
Pubblicazione: (2025)
Kimi-Audio Technical Report
di: KimiTeam, et al.
Pubblicazione: (2025)
di: KimiTeam, et al.
Pubblicazione: (2025)
GEOCHEMICAL ANALYSES OF EOCENE OILS IN DEEPLY BURIED SANDSTONE RESERVOIRS IN THE DONGYING DEPRESSION, BOHAI BAY BASIN, NE CHINA
di: Xiaoxiao Zhou, et al.
Pubblicazione: (2024)
di: Xiaoxiao Zhou, et al.
Pubblicazione: (2024)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
di: Chen, Chen, et al.
Pubblicazione: (2025)
di: Chen, Chen, et al.
Pubblicazione: (2025)
DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive Model
di: Chen, Bohong, et al.
Pubblicazione: (2025)
di: Chen, Bohong, et al.
Pubblicazione: (2025)
Option Market Making via Reinforcement Learning
di: Fang, Zhou, et al.
Pubblicazione: (2023)
di: Fang, Zhou, et al.
Pubblicazione: (2023)
Convergence of the Planewave Approximations for Quantum Incommensurate Systems
di: Wang, Ting, et al.
Pubblicazione: (2022)
di: Wang, Ting, et al.
Pubblicazione: (2022)
On the relative Langlands duality for $\operatorname{Sp}_{2n} \backslash \operatorname{GL}_{2n+1}$ (with an appendix by Zeyu Wang)
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
Central values of Asai L-functions and twisted Gan--Gross--Prasad conjecture
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
AirIMU: Learning Uncertainty Propagation for Inertial Odometry
di: Qiu, Yuheng, et al.
Pubblicazione: (2023)
di: Qiu, Yuheng, et al.
Pubblicazione: (2023)
机法二深信与判教智慧——善导"要弘二门"与太虚"大乘三系"的会通及圆觉、极乐、华藏三重境界阐释
di: Chen, Jianlin
Pubblicazione: (2026)
di: Chen, Jianlin
Pubblicazione: (2026)
LMStyle Benchmark: Evaluating Text Style Transfer for Chatbots
di: Chen, Jianlin
Pubblicazione: (2024)
di: Chen, Jianlin
Pubblicazione: (2024)
Efficient Video Diffusion with Sparse Information Transmission for Video Compression
di: Zhou, Mingde, et al.
Pubblicazione: (2026)
di: Zhou, Mingde, et al.
Pubblicazione: (2026)
Penalty-Based First-Order Methods for Bilevel Optimization with Minimax and Constrained Lower-Level Problems
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
di: Shen, Yiyang, et al.
Pubblicazione: (2026)
Sharp criteria for a degenerate diffusion-aggregation system with the intermediate exponent
di: Zhou, Tiantian, et al.
Pubblicazione: (2025)
di: Zhou, Tiantian, et al.
Pubblicazione: (2025)
TurboRAG: Accelerating Retrieval-Augmented Generation with Precomputed KV Caches for Chunked Text
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
di: Lu, Songshuo, et al.
Pubblicazione: (2024)
SCRNet: Spatial-Channel Regulation Network for Medical Ultrasound Image Segmentation
di: Xu, Weixin, et al.
Pubblicazione: (2025)
di: Xu, Weixin, et al.
Pubblicazione: (2025)
After Admission: The Emotional Suffering of Students Enrolled Through the Rural Students Quota Plan in China's Elite Universities
di: Songdi Wang, et al.
Pubblicazione: (2024)
di: Songdi Wang, et al.
Pubblicazione: (2024)
Periods detecting Eisenstein series and sums of $L$-values I
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
di: Lu, Weixiao, et al.
Pubblicazione: (2025)
Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
di: Wang, Jiahao, et al.
Pubblicazione: (2025)
Evolutionary Greedy Algorithm for Optimal Sensor Placement Problem in Urban Sewage Surveillance
di: Wang, Sunyu, et al.
Pubblicazione: (2024)
di: Wang, Sunyu, et al.
Pubblicazione: (2024)
EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
CE-LoRA: Computation-Efficient LoRA Fine-Tuning for Language Models
di: Chen, Guanduo, et al.
Pubblicazione: (2025)
di: Chen, Guanduo, et al.
Pubblicazione: (2025)
Performance assessment of the effective core potentials under the Fermionic neural network: first and second row elements
di: Wang, Mengsa, et al.
Pubblicazione: (2024)
di: Wang, Mengsa, et al.
Pubblicazione: (2024)
Kimi-VL Technical Report
di: Kimi Team, et al.
Pubblicazione: (2025)
di: Kimi Team, et al.
Pubblicazione: (2025)
Low‐Dose Golidocitinib for Triple‐Mutant (JAK1/STAT3/TET2) ITLPD‐GI After DLBCL Remission
di: Wenwen Wang, et al.
Pubblicazione: (2026)
di: Wenwen Wang, et al.
Pubblicazione: (2026)
Canagliflozin Delays Aortic Valve Calcification by Enhancing the AMPK /Nrf2/ HO ‐1 Antioxidant Signaling Pathway in Valvular Interstitial Cells
di: Quangong Zhao, et al.
Pubblicazione: (2025)
di: Quangong Zhao, et al.
Pubblicazione: (2025)
Reconstruction of multiple strings of constant weight from prefix-suffix compositions
di: Yang, Yaoyu, et al.
Pubblicazione: (2024)
di: Yang, Yaoyu, et al.
Pubblicazione: (2024)
Arbitrage on Decentralized Exchanges
di: He, Xue Dong, et al.
Pubblicazione: (2025)
di: He, Xue Dong, et al.
Pubblicazione: (2025)
Optimal Design of Automated Market Makers on Decentralized Exchanges
di: He, Xue Dong, et al.
Pubblicazione: (2024)
di: He, Xue Dong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Muon is Scalable for LLM Training
di: Liu, Jingyuan, et al.
Pubblicazione: (2025) -
MoBA: Mixture of Block Attention for Long-Context LLMs
di: Lu, Enzhe, et al.
Pubblicazione: (2025) -
Kimi Linear: An Expressive, Efficient Attention Architecture
di: Kimi Team, et al.
Pubblicazione: (2025) -
EvoEmpirBench: Dynamic Spatial Reasoning with Agent-ExpVer
di: Zhao, Pukun, et al.
Pubblicazione: (2025) -
STAR: Mitigating Cascading Errors in Spatial Reasoning via Turn-point Alignment and Segment-level DPO
di: Zhao, Pukun, et al.
Pubblicazione: (2026)