Gespeichert in:
| Hauptverfasser: | Kim, Yongwan, Park, Sungchul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.25813 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025)
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025)
von: Wu, Xun, et al.
Veröffentlicht: (2025)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026)
Improving SMOTE via Fusing Conditional VAE for Data-adaptive Noise Filtering
von: Hong, Sungchul, et al.
Veröffentlicht: (2024)
von: Hong, Sungchul, et al.
Veröffentlicht: (2024)
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
von: Zhao, Yushu, et al.
Veröffentlicht: (2025)
Bilevel Autoresearch: Meta-Autoresearching Itself
von: Qu, Yaonan, et al.
Veröffentlicht: (2026)
von: Qu, Yaonan, et al.
Veröffentlicht: (2026)
LittleBit: Ultra Low-Bit Quantization via Latent Factorization
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
von: Lee, Banseok, et al.
Veröffentlicht: (2025)
LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment
von: Lee, Banseok, et al.
Veröffentlicht: (2026)
von: Lee, Banseok, et al.
Veröffentlicht: (2026)
Outlier-Safe Pre-Training for Robust 4-Bit Quantization of Large Language Models
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
von: Park, Jungwoo, et al.
Veröffentlicht: (2025)
Transcendence: Generative Models Can Outperform The Experts That Train Them
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
von: Zhang, Edwin, et al.
Veröffentlicht: (2024)
Federated Domain Generalization with Label Smoothing and Balanced Decentralized Training
von: Soltany, Milad, et al.
Veröffentlicht: (2024)
von: Soltany, Milad, et al.
Veröffentlicht: (2024)
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
von: Farhat, Yehya, et al.
Veröffentlicht: (2023)
BitNet b1.58 2B4T Technical Report
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
von: Ma, Shuming, et al.
Veröffentlicht: (2025)
How Many Experts Are Enough? Towards Optimal Semantic Specialization for Mixture-of-Experts
von: Park, Sumin, et al.
Veröffentlicht: (2025)
von: Park, Sumin, et al.
Veröffentlicht: (2025)
dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis
von: Xie, Luyuan, et al.
Veröffentlicht: (2025)
von: Xie, Luyuan, et al.
Veröffentlicht: (2025)
Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
von: Park, Jongseok, et al.
Veröffentlicht: (2026)
Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks
von: Kim, Jeongmo, et al.
Veröffentlicht: (2025)
von: Kim, Jeongmo, et al.
Veröffentlicht: (2025)
Decentralized Adversarial Training over Graphs
von: Cao, Ying, et al.
Veröffentlicht: (2023)
von: Cao, Ying, et al.
Veröffentlicht: (2023)
DAMEL: Dual-Axis Multi-Expert Learning for Class-Imbalanced Learning
von: Lee, Hyuck, et al.
Veröffentlicht: (2026)
von: Lee, Hyuck, et al.
Veröffentlicht: (2026)
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Peiyuan, et al.
Veröffentlicht: (2026)
AdaQAT: Adaptive Bit-Width Quantization-Aware Training
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
von: Gernigon, Cédric, et al.
Veröffentlicht: (2024)
MixtureKit: A General Framework for Composing, Training, and Visualizing Mixture-of-Experts Models
von: Chamma, Ahmad, et al.
Veröffentlicht: (2025)
von: Chamma, Ahmad, et al.
Veröffentlicht: (2025)
Local Mixtures of Experts: Essentially Free Test-Time Training via Model Merging
von: Bertolissi, Ryo, et al.
Veröffentlicht: (2025)
von: Bertolissi, Ryo, et al.
Veröffentlicht: (2025)
BitNet a4.8: 4-bit Activations for 1-bit LLMs
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Forecasting VIX using interpretable Kolmogorov-Arnold networks
von: Cho, So-Yoon, et al.
Veröffentlicht: (2025)
von: Cho, So-Yoon, et al.
Veröffentlicht: (2025)
Decentralized Autoregressive Generation
von: Maschan, Stepan, et al.
Veröffentlicht: (2026)
von: Maschan, Stepan, et al.
Veröffentlicht: (2026)
Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language Models
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
von: Zhang, Tianao, et al.
Veröffentlicht: (2025)
Load-Aware Training Scheduling for Model Circulation-based Decentralized Federated Learning
von: Kainuma, Haruki, et al.
Veröffentlicht: (2025)
von: Kainuma, Haruki, et al.
Veröffentlicht: (2025)
Characterization of Multi-Model Agentic AI Systems on General Tasks via Trace-Driven Simulation
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
von: Kim, Donghwan, et al.
Veröffentlicht: (2026)
FRAIN to Train: A Fast-and-Reliable Solution for Decentralized Federated Learning
von: Park, Sanghyeon, et al.
Veröffentlicht: (2025)
von: Park, Sanghyeon, et al.
Veröffentlicht: (2025)
Loop Corrections to the Training Error and Generalization Gap of Random Feature Models
von: Kim, Taeyoung
Veröffentlicht: (2026)
von: Kim, Taeyoung
Veröffentlicht: (2026)
Exploring Expert Specialization through Unsupervised Training in Sparse Mixture of Experts
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
von: Nikolic, Strahinja, et al.
Veröffentlicht: (2025)
Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
von: Nguyen-Nhat, Minh-Khoi, et al.
Veröffentlicht: (2025)
Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation
von: Tian, Hanlin, et al.
Veröffentlicht: (2024)
von: Tian, Hanlin, et al.
Veröffentlicht: (2024)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
Towards Improving Long-Tail Entity Predictions in Temporal Knowledge Graphs through Global Similarity and Weighted Sampling
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Mehrnoosh, et al.
Veröffentlicht: (2025)
Efficient Epistemic Uncertainty Estimation for Large Language Models via Knowledge Distillation
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
von: Park, Seonghyeon, et al.
Veröffentlicht: (2026)
PFedDST: Personalized Federated Learning with Decentralized Selection Training
von: Fan, Mengchen, et al.
Veröffentlicht: (2025)
von: Fan, Mengchen, et al.
Veröffentlicht: (2025)
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
von: Zhang, Jinhao Zhang Yunquan, et al.
Veröffentlicht: (2026)
von: Zhang, Jinhao Zhang Yunquan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
von: Nielsen, Jacob, et al.
Veröffentlicht: (2025) -
BitNet Distillation
von: Wu, Xun, et al.
Veröffentlicht: (2025) -
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024) -
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
von: Panfilov, Alexander, et al.
Veröffentlicht: (2026) -
Improving SMOTE via Fusing Conditional VAE for Data-adaptive Noise Filtering
von: Hong, Sungchul, et al.
Veröffentlicht: (2024)