STBLLM: Breaking the 1-Bit Barrier with Structured Binary LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Peijie, Li, Lujun, Zhong, Yuedong, Du, Dayou, Fan, Ruibo, Chen, Yuhan, Tang, Zhenheng, Wang, Qiang, Xue, Wei, Guo, Yike, Chu, Xiaowen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
LPZero: Language Model Zero-cost Proxy Search from Zero
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
by: Du, Dayou, et al.
Published: (2024)
by: Du, Dayou, et al.
Published: (2024)
VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting
by: Tang, Yujin, et al.
Published: (2024)
by: Tang, Yujin, et al.
Published: (2024)
Benchmarking and Dissecting the Nvidia Hopper GPU Architecture
by: Luo, Weile, et al.
Published: (2024)
by: Luo, Weile, et al.
Published: (2024)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
by: Luo, Weile, et al.
Published: (2025)
by: Luo, Weile, et al.
Published: (2025)
Model Quantization and Hardware Acceleration for Vision Transformers: A Comprehensive Survey
by: Du, Dayou, et al.
Published: (2024)
by: Du, Dayou, et al.
Published: (2024)
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production
by: Liu, Xiang, et al.
Published: (2026)
by: Liu, Xiang, et al.
Published: (2026)
ParZC: Parametric Zero-Cost Proxies for Efficient NAS
by: Dong, Peijie, et al.
Published: (2024)
by: Dong, Peijie, et al.
Published: (2024)
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?
by: Tang, Zhenheng, et al.
Published: (2025)
by: Tang, Zhenheng, et al.
Published: (2025)
Should We Really Edit Language Models? On the Evaluation of Edited Language Models
by: Li, Qi, et al.
Published: (2024)
by: Li, Qi, et al.
Published: (2024)
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
by: Su, Xiaoyan, et al.
Published: (2026)
by: Su, Xiaoyan, et al.
Published: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
by: Xu, Binxing, et al.
Published: (2026)
by: Xu, Binxing, et al.
Published: (2026)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
by: Du, Dayou, et al.
Published: (2025)
by: Du, Dayou, et al.
Published: (2025)
FuseFL: One-Shot Federated Learning through the Lens of Causality with Progressive Model Fusion
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)
by: Tang, Zichen, et al.
Published: (2026)
by: Tang, Zichen, et al.
Published: (2026)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
by: Lai, Kunfeng, et al.
Published: (2025)
by: Lai, Kunfeng, et al.
Published: (2025)
LongGenBench: Long-context Generation Benchmark
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
AnTKV: Anchor Token-Aware Sub-Bit Vector Quantization for KV Cache in Large Language Models
by: Li, Zeyu, et al.
Published: (2025)
by: Li, Zeyu, et al.
Published: (2025)
Dissecting Outlier Dynamics in LLM NVFP4 Pretraining
by: Dong, Peijie, et al.
Published: (2026)
by: Dong, Peijie, et al.
Published: (2026)
OmniReview: A Large-scale Benchmark and LLM-enhanced Framework for Realistic Reviewer Recommendation
by: Huang, Yehua, et al.
Published: (2026)
by: Huang, Yehua, et al.
Published: (2026)
Delta Decompression for MoE-based LLMs Compression
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
Are Dilemmas and Conflicts in LLM Alignment Solvable? A View from Priority Graph
by: Tang, Zhenheng, et al.
Published: (2026)
by: Tang, Zhenheng, et al.
Published: (2026)
Rethinking Deep Research from the Perspective of Web Content Distribution Matching
by: Yu, Zixuan, et al.
Published: (2026)
by: Yu, Zixuan, et al.
Published: (2026)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
by: Ai, Yuang, et al.
Published: (2026)
by: Ai, Yuang, et al.
Published: (2026)
Reasoning Language Model Inference Serving Unveiled: An Empirical Study
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
You Know What I'm Saying: Jailbreak Attack via Implicit Reference
by: Wu, Tianyu, et al.
Published: (2024)
by: Wu, Tianyu, et al.
Published: (2024)
NoRA: Nested Low-Rank Adaptation for Efficient Fine-Tuning Large Models
by: Lin, Cheng, et al.
Published: (2024)
by: Lin, Cheng, et al.
Published: (2024)
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM Training
by: Lin, Wenxiang, et al.
Published: (2026)
by: Lin, Wenxiang, et al.
Published: (2026)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
by: Guo, Jingkai, et al.
Published: (2025)
by: Guo, Jingkai, et al.
Published: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
From ChatGPT to DeepSeek: Can LLMs Simulate Humanity?
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
by: Pan, Xinglin, et al.
Published: (2025)
by: Pan, Xinglin, et al.
Published: (2025)
FedImpro: Measuring and Improving Client Update in Federated Learning
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Improving Text-to-Image Generation with Input-Side Inference-Time Scaling
by: Chen, Ruibo, et al.
Published: (2025)
by: Chen, Ruibo, et al.
Published: (2025)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
by: Xie, Rui, et al.
Published: (2025)
by: Xie, Rui, et al.
Published: (2025)
Similar Items
-
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025) -
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
by: Dong, Peijie, et al.
Published: (2024) -
LPZero: Language Model Zero-cost Proxy Search from Zero
by: Dong, Peijie, et al.
Published: (2024) -
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
by: Du, Dayou, et al.
Published: (2024) -
VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting
by: Tang, Yujin, et al.
Published: (2024)