Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Seungcheol, Bae, Jeongin, Kwon, Beomseok, Kim, Minjun, Kim, Byeongwook, Kwon, Se Jung, Kang, U, Lee, Dongsoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2023)
von: Park, Seungcheol, et al.
Veröffentlicht: (2023)
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
von: Park, Seungcheol, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Compression Algorithms for Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2024)
von: Park, Seungcheol, et al.
Veröffentlicht: (2024)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025)
von: Park, Gunho, et al.
Veröffentlicht: (2025)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
von: Kim, Heejun, et al.
Veröffentlicht: (2026)
Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
von: Seo, Yeongbin, et al.
Veröffentlicht: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026)
von: Johnson, Warren
Veröffentlicht: (2026)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
von: Kim, Jaemin, et al.
Veröffentlicht: (2024)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
von: Johnson, Warren, et al.
Veröffentlicht: (2026)
"As Eastern Powers, I will veto." : An Investigation of Nation-level Bias of Large Language Models in International Relations
von: Choi, Jonghyeon, et al.
Veröffentlicht: (2025)
von: Choi, Jonghyeon, et al.
Veröffentlicht: (2025)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
von: Sun, Jingyi, et al.
Veröffentlicht: (2024)
von: Sun, Jingyi, et al.
Veröffentlicht: (2024)
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
Trusted Uncertainty in Large Language Models: A Unified Framework for Confidence Calibration and Risk-Controlled Refusal
von: Oehri, Markus, et al.
Veröffentlicht: (2025)
von: Oehri, Markus, et al.
Veröffentlicht: (2025)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
von: Yang, June Yong, et al.
Veröffentlicht: (2024)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
von: Jin, Heegon, et al.
Veröffentlicht: (2024)
von: Jin, Heegon, et al.
Veröffentlicht: (2024)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
von: Lee, Yejin, et al.
Veröffentlicht: (2025)
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
von: Vasilev, Stefan, et al.
Veröffentlicht: (2025)
von: Vasilev, Stefan, et al.
Veröffentlicht: (2025)
Tokenization Is More Than Compression
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
von: Schmidt, Craig W., et al.
Veröffentlicht: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks
von: Luo, Jianwen, et al.
Veröffentlicht: (2025)
von: Luo, Jianwen, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
SAGE: Sign-Adaptive Gradient for Memory-Efficient LLM Optimization
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
von: Lee, Wooin, et al.
Veröffentlicht: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hoang, et al.
Veröffentlicht: (2025)
Improving Commonsense Bias Classification by Mitigating the Influence of Demographic Terms
von: Lee, JinKyu, et al.
Veröffentlicht: (2024)
von: Lee, JinKyu, et al.
Veröffentlicht: (2024)
Revealing the Parametric Knowledge of Language Models: A Unified Framework for Attribution Methods
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
von: Yu, Haeun, et al.
Veröffentlicht: (2024)
Towards Probabilistic Question Answering Over Tabular Data
von: Shen, Chen, et al.
Veröffentlicht: (2025)
von: Shen, Chen, et al.
Veröffentlicht: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
von: Tessema, Bethel Melesse, et al.
Veröffentlicht: (2024)
TCProF: Time-Complexity Prediction SSL Framework
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
von: Haque, Md. Asraful, et al.
Veröffentlicht: (2026)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
von: Goldin, Gili, et al.
Veröffentlicht: (2024)
Math Natural Language Inference: this should be easy!
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
von: de Paiva, Valeria, et al.
Veröffentlicht: (2025)
Fast Quiet-STaR: Thinking Without Thought Tokens
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
von: Wang, Zhilin, et al.
Veröffentlicht: (2026)
Pitfalls in Evaluating Interpretability Agents
von: Haklay, Tal, et al.
Veröffentlicht: (2026)
von: Haklay, Tal, et al.
Veröffentlicht: (2026)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
von: Gómez-Rodríguez, Carlos, et al.
Veröffentlicht: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2023) -
Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information
von: Park, Seungcheol, et al.
Veröffentlicht: (2025) -
A Comprehensive Survey of Compression Algorithms for Language Models
von: Park, Seungcheol, et al.
Veröffentlicht: (2024) -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
von: Park, Gunho, et al.
Veröffentlicht: (2025) -
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
von: Kim, Heejun, et al.
Veröffentlicht: (2026)