OneBit: Towards Extremely Low-bit Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yuzhuang, Han, Xu, Yang, Zonghan, Wang, Shuo, Zhu, Qingfu, Liu, Zhiyuan, Liu, Weidong, Che, Wanxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024)
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2026)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2026)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
von: Liu, Yijun, et al.
Veröffentlicht: (2025)
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2025)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2025)
HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
von: Shan, Baocai, et al.
Veröffentlicht: (2026)
von: Shan, Baocai, et al.
Veröffentlicht: (2026)
Advancing Tool-Augmented Large Language Models via Meta-Verification and Reflection Learning
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2025)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2026)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2026)
Exploring Large Language Models for Communication Games: An Empirical Study on Werewolf
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
When Does Context Help? Error Dynamics of Contextual Information in Large Language Models
von: Wang, Dingzirui, et al.
Veröffentlicht: (2026)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2026)
Semi-Instruct: Bridging Natural-Instruct and Self-Instruct for Code Large Language Models
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
von: Luo, Xianzhen, et al.
Veröffentlicht: (2024)
A Survey of Table Reasoning with Large Language Models
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
How Do Language Models Understand Tables? A Mechanistic Analysis of Cell Location
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2026)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
von: Niu, Tianhao, et al.
Veröffentlicht: (2026)
AlignEvoSkill: Towards Knowledge-Aware and Task-Aligned Agent Skill Evolution
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
Pluggable Neural Machine Translation Models via Memory-augmented Adapters
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2023)
Scaling Laws for Agent Harnesses via Effective Feedback Compute
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2026)
Abacus-SQL: A Text-to-SQL System Empowering Cross-Domain and Open-Domain Database Retrieval
von: Xu, Keyan, et al.
Veröffentlicht: (2025)
von: Xu, Keyan, et al.
Veröffentlicht: (2025)
RoT: Enhancing Table Reasoning with Iterative Row-Wise Traversals
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2025)
Delta-CoMe: Training-Free Delta-Compression with Mixed-Precision for Large Language Models
von: Ping, Bowen, et al.
Veröffentlicht: (2024)
von: Ping, Bowen, et al.
Veröffentlicht: (2024)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
von: Wang, Yixuan, et al.
Veröffentlicht: (2025)
Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling
von: Ji, Shiyu, et al.
Veröffentlicht: (2025)
von: Ji, Shiyu, et al.
Veröffentlicht: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
von: Ji, Shiyu, et al.
Veröffentlicht: (2026)
Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzriui, et al.
Veröffentlicht: (2025)
Exploring Hybrid Question Answering via Program-based Prompting
von: Shi, Qi, et al.
Veröffentlicht: (2024)
von: Shi, Qi, et al.
Veröffentlicht: (2024)
Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring
von: Mu, Honglin, et al.
Veröffentlicht: (2024)
von: Mu, Honglin, et al.
Veröffentlicht: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2025)
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
Know More, Know Clearer: A Meta-Cognitive Framework for Knowledge Augmentation in Large Language Models
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Improving Demonstration Diversity by Human-Free Fusing for Text-to-SQL
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
DAC: Decomposed Automation Correction for Text-to-SQL
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQL
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
Language Anisotropic Cross-Lingual Model Editing
von: Xu, Yang, et al.
Veröffentlicht: (2022)
von: Xu, Yang, et al.
Veröffentlicht: (2022)
Improving Grammatical Error Correction via Contextual Data Augmentation
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
von: Wang, Yixuan, et al.
Veröffentlicht: (2024)
Enabling Real-Time Conversations with Minimal Training Costs
von: Xu, Wang, et al.
Veröffentlicht: (2024)
von: Xu, Wang, et al.
Veröffentlicht: (2024)
Improving Language Model Reasoning with Self-motivated Learning
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
Perspective Transition of Large Language Models for Solving Subjective Tasks
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2024) -
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2026) -
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
von: Wang, Yixuan, et al.
Veröffentlicht: (2025) -
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
von: Liu, Yijun, et al.
Veröffentlicht: (2025) -
CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
von: Xu, Yuzhuang, et al.
Veröffentlicht: (2025)