TernaryLLM: Ternarized Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Tianqi, Li, Zhe, Xu, Weixiang, Zhu, Zeyu, Li, Dong, Tian, Lu, Barsoum, Emad, Wang, Peisong, Cheng, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
von: Ke, Wenjin, et al.
Veröffentlicht: (2025)
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
von: Yan, Xianglong, et al.
Veröffentlicht: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025)
von: An, Zihao, et al.
Veröffentlicht: (2025)
Ternarization of Vision Language Models for use on edge devices
von: Crulis, Ben, et al.
Veröffentlicht: (2025)
von: Crulis, Ben, et al.
Veröffentlicht: (2025)
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Enhancing One-shot Pruned Pre-trained Language Models through Sparse-Dense-Sparse Mechanism
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
von: Li, Guanchen, et al.
Veröffentlicht: (2024)
FastGL: A GPU-Efficient Framework for Accelerating Sampling-Based GNN Training at Large Scale
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2024)
Theory-optimal Quantization Based on Flatness
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
von: Huang, Xiusheng, et al.
Veröffentlicht: (2026)
Týr-the-Pruner: Structural Pruning LLMs via Global Sparsity Distribution Optimization
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
von: Li, Guanchen, et al.
Veröffentlicht: (2025)
Learnable Permutation for Structured Sparsity on Transformer Models
von: Li, Zekai, et al.
Veröffentlicht: (2026)
von: Li, Zekai, et al.
Veröffentlicht: (2026)
HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
von: Xie, Zhinan, et al.
Veröffentlicht: (2025)
von: Xie, Zhinan, et al.
Veröffentlicht: (2025)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
von: Yang, Mingyu, et al.
Veröffentlicht: (2025)
A Survey of Graph Meets Large Language Model: Progress and Future Directions
von: Li, Yuhan, et al.
Veröffentlicht: (2023)
von: Li, Yuhan, et al.
Veröffentlicht: (2023)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
von: Dukler, Yonatan, et al.
Veröffentlicht: (2025)
von: Dukler, Yonatan, et al.
Veröffentlicht: (2025)
Tequila: Trapping-free Ternary Quantization for Large Language Models
von: Huang, Hong, et al.
Veröffentlicht: (2025)
von: Huang, Hong, et al.
Veröffentlicht: (2025)
GLBench: A Comprehensive Benchmark for Graph with Large Language Models
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
Policy Disruption in Reinforcement Learning:Adversarial Attack with Large Language Models and Critical State Identification
von: Jiang, Junyong, et al.
Veröffentlicht: (2025)
von: Jiang, Junyong, et al.
Veröffentlicht: (2025)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
DALI: A Workload-Aware Offloading Framework for Efficient MoE Inference on Local PCs
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
von: Zhu, Zeyu, et al.
Veröffentlicht: (2026)
Preserving Diversity in Supervised Fine-Tuning of Large Language Models
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
von: Li, Ziniu, et al.
Veröffentlicht: (2024)
Understanding the Role of Textual Prompts in LLM for Time Series Forecasting: an Adapter View
von: Niu, Peisong, et al.
Veröffentlicht: (2023)
von: Niu, Peisong, et al.
Veröffentlicht: (2023)
Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding
von: Li, Jinze, et al.
Veröffentlicht: (2025)
von: Li, Jinze, et al.
Veröffentlicht: (2025)
LoFT-LLM: Low-Frequency Time-Series Forecasting with Large Language Models
von: You, Jiacheng, et al.
Veröffentlicht: (2025)
von: You, Jiacheng, et al.
Veröffentlicht: (2025)
Binary and Ternary Quantization Can Enhance Feature Discrimination
von: Lu, Weizhi, et al.
Veröffentlicht: (2025)
von: Lu, Weizhi, et al.
Veröffentlicht: (2025)
LLM4XCE: Large Language Models for Extremely Large-Scale Massive MIMO Channel Estimation
von: Li, Renbin, et al.
Veröffentlicht: (2025)
von: Li, Renbin, et al.
Veröffentlicht: (2025)
Verbalized Graph Representation Learning: A Fully Interpretable Graph Model Based on Large Language Models Throughout the Entire Process
von: Ji, Xingyu, et al.
Veröffentlicht: (2024)
von: Ji, Xingyu, et al.
Veröffentlicht: (2024)
Ban&Pick: Ehancing Performance and Efficiency of MoE-LLMs via Smarter Routing
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)
Block Rotation is All You Need for MXFP4 Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
DartQuant: Efficient Rotational Distribution Calibration for LLM Quantization
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
von: Shao, Yuantian, et al.
Veröffentlicht: (2025)
ResPrune: Text-Conditioned Subspace Reconstruction for Visual Token Pruning in Large Vision-Language Models
von: Li, Xu, et al.
Veröffentlicht: (2026)
von: Li, Xu, et al.
Veröffentlicht: (2026)
SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
von: Liao, Huanxuan, et al.
Veröffentlicht: (2025)
EfficientLLM: Efficiency in Large Language Models
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2025)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2025)
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
von: Li, Yuhan, et al.
Veröffentlicht: (2024)
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
von: Li, Xiangman, et al.
Veröffentlicht: (2025)
von: Li, Xiangman, et al.
Veröffentlicht: (2025)
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
von: Wang, Zhaoxin, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoxin, et al.
Veröffentlicht: (2026)
CultureLLM: Incorporating Cultural Differences into Large Language Models
von: Li, Cheng, et al.
Veröffentlicht: (2024)
von: Li, Cheng, et al.
Veröffentlicht: (2024)
Certain Head, Uncertain Tail: Expert-Sample for Test-Time Scaling in Fine-Grained MoE
von: Chen, Yuanteng, et al.
Veröffentlicht: (2026)
von: Chen, Yuanteng, et al.
Veröffentlicht: (2026)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
von: Fashi, Parsa Ashrafi, et al.
Veröffentlicht: (2026)
von: Fashi, Parsa Ashrafi, et al.
Veröffentlicht: (2026)
ADO-LLM: Analog Design Bayesian Optimization with In-Context Learning of Large Language Models
von: Yin, Yuxuan, et al.
Veröffentlicht: (2024)
von: Yin, Yuxuan, et al.
Veröffentlicht: (2024)
Astro: Activation-guided Structured Regularization for Outlier-Robust LLM Post-Training Quantization
von: Chen, Xi, et al.
Veröffentlicht: (2026)
von: Chen, Xi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
von: Ke, Wenjin, et al.
Veröffentlicht: (2025) -
PT$^2$-LLM: Post-Training Ternarization for Large Language Models
von: Yan, Xianglong, et al.
Veröffentlicht: (2025) -
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
von: An, Zihao, et al.
Veröffentlicht: (2025) -
Ternarization of Vision Language Models for use on edge devices
von: Crulis, Ben, et al.
Veröffentlicht: (2025) -
EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
von: Chen, Yuanteng, et al.
Veröffentlicht: (2025)