TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Junhan, Park, Yeo Jeong, Son, Seungwoo, Lee, Chungman, Kim, Ho-young, Kim, Joonyoung, Jeon, Yongkweon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
BoA: Attention-aware Post-training Quantization without Backpropagation
por: Kim, Junhan, et al.
Publicado: (2024)
por: Kim, Junhan, et al.
Publicado: (2024)
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
por: Kim, Junhan, et al.
Publicado: (2024)
por: Kim, Junhan, et al.
Publicado: (2024)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
por: Kim, Junhan, et al.
Publicado: (2026)
por: Kim, Junhan, et al.
Publicado: (2026)
On the Importance of a Multi-Scale Calibration for Quantization
por: Son, Seungwoo, et al.
Publicado: (2026)
por: Son, Seungwoo, et al.
Publicado: (2026)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
por: Ahn, Jinwoo, et al.
Publicado: (2026)
por: Ahn, Jinwoo, et al.
Publicado: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
por: Son, Seungwoo, et al.
Publicado: (2024)
por: Son, Seungwoo, et al.
Publicado: (2024)
HLQ: Fast and Efficient Backpropagation via Hadamard Low-rank Quantization
por: Kim, Seonggon, et al.
Publicado: (2024)
por: Kim, Seonggon, et al.
Publicado: (2024)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
por: Park, Juneyoung, et al.
Publicado: (2026)
por: Park, Juneyoung, et al.
Publicado: (2026)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
por: Kim, Taesu, et al.
Publicado: (2024)
por: Kim, Taesu, et al.
Publicado: (2024)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
por: Park, Juneyoung, et al.
Publicado: (2026)
por: Park, Juneyoung, et al.
Publicado: (2026)
OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance
por: Park, Yeo Jeong, et al.
Publicado: (2026)
por: Park, Yeo Jeong, et al.
Publicado: (2026)
Efficient Personalization of Quantized Diffusion Model without Backpropagation
por: Seo, Hoigi, et al.
Publicado: (2025)
por: Seo, Hoigi, et al.
Publicado: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
por: Cho, Yoonjun, et al.
Publicado: (2026)
por: Cho, Yoonjun, et al.
Publicado: (2026)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
por: Jeon, Hyesung, et al.
Publicado: (2024)
por: Jeon, Hyesung, et al.
Publicado: (2024)
FEATHer: Fourier-Efficient Adaptive Temporal Hierarchy Forecaster for Time-Series Forecasting
por: Lee, Jaehoon, et al.
Publicado: (2026)
por: Lee, Jaehoon, et al.
Publicado: (2026)
Activation Quantization of Vision Encoders Needs Prefixing Registers
por: Kim, Seunghyeon, et al.
Publicado: (2025)
por: Kim, Seunghyeon, et al.
Publicado: (2025)
Are Self-Attentions Effective for Time Series Forecasting?
por: Kim, Dongbin, et al.
Publicado: (2024)
por: Kim, Dongbin, et al.
Publicado: (2024)
DAPD: Dependency-Aware Parallel Decoding via Attention for Diffusion LLMs
por: Kim, Bumjun, et al.
Publicado: (2026)
por: Kim, Bumjun, et al.
Publicado: (2026)
Difficulty-aware Balancing Margin Loss for Long-tailed Recognition
por: Son, Minseok, et al.
Publicado: (2024)
por: Son, Minseok, et al.
Publicado: (2024)
Learning to Continually Learn with the Bayesian Principle
por: Lee, Soochan, et al.
Publicado: (2024)
por: Lee, Soochan, et al.
Publicado: (2024)
Attention-aware Semantic Communications for Collaborative Inference
por: Im, Jiwoong, et al.
Publicado: (2024)
por: Im, Jiwoong, et al.
Publicado: (2024)
U-Net for crab image semantic segmentation with PyTorch : UAV imaging and deep learning approach can id entify Brachyura in tidal flats
por: Dongwoo Kim, et al.
Publicado: (2024)
por: Dongwoo Kim, et al.
Publicado: (2024)
FedWSQ: Efficient Federated Learning with Weight Standardization and Distribution-Aware Non-Uniform Quantization
por: Kim, Seung-Wook, et al.
Publicado: (2025)
por: Kim, Seung-Wook, et al.
Publicado: (2025)
DeepHQ: Learned Hierarchical Quantizer for Progressive Deep Image Coding
por: Lee, Jooyoung, et al.
Publicado: (2024)
por: Lee, Jooyoung, et al.
Publicado: (2024)
Quantization-Aware Imitation-Learning for Resource-Efficient Robotic Control
por: Park, Seongmin, et al.
Publicado: (2024)
por: Park, Seongmin, et al.
Publicado: (2024)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
por: Lee, Heejun, et al.
Publicado: (2024)
por: Lee, Heejun, et al.
Publicado: (2024)
Thickness-aware E(3)-Equivariant 3D Mesh Neural Networks
por: Kim, Sungwon, et al.
Publicado: (2025)
por: Kim, Sungwon, et al.
Publicado: (2025)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
por: Kim, Jiyoon, et al.
Publicado: (2025)
por: Kim, Jiyoon, et al.
Publicado: (2025)
IRA: Adaptive Interest-aware Representation and Alignment for Personalized Multi-interest Retrieval
por: Lee, Youngjune, et al.
Publicado: (2025)
por: Lee, Youngjune, et al.
Publicado: (2025)
An Experimental Study of the Relative Sensitivity to Oxidative Aging of NaBH 4 ‐promoted Green Hypergolic Fuels
por: Kyounghwan Lee, et al.
Publicado: (2025)
por: Kyounghwan Lee, et al.
Publicado: (2025)
ASAP: Attention Sink Anchored Pruning
por: Lee, Jaehyuk, et al.
Publicado: (2026)
por: Lee, Jaehyuk, et al.
Publicado: (2026)
SCOPE-FE: Structured Control of Operator and Pairwise Exploration for Feature Engineering
por: Park, Minhee, et al.
Publicado: (2026)
por: Park, Minhee, et al.
Publicado: (2026)
Spline Dimensional Decomposition with Interpolation-based Optimal Knot Selection for Stochastic Dynamic Analysis
por: Kim, Yeonsu, et al.
Publicado: (2025)
por: Kim, Yeonsu, et al.
Publicado: (2025)
Grouped Differential Attention
por: Lim, Junghwan, et al.
Publicado: (2025)
por: Lim, Junghwan, et al.
Publicado: (2025)
Assigning Distinct Roles to Quantized and Low-Rank Matrices Toward Optimal Weight Decomposition
por: Cho, Yoonjun, et al.
Publicado: (2025)
por: Cho, Yoonjun, et al.
Publicado: (2025)
Bridging the Gap Between Molecule and Textual Descriptions via Substructure-aware Alignment
por: Park, Hyuntae, et al.
Publicado: (2025)
por: Park, Hyuntae, et al.
Publicado: (2025)
Dynamic Time-aware Continual User Representation Learning
por: Choi, Seungyoon, et al.
Publicado: (2025)
por: Choi, Seungyoon, et al.
Publicado: (2025)
Hierarchical Attention-based Graph Neural Network with Relevance-driven Pruning
por: Kum, Seungwoo
Publicado: (2026)
por: Kum, Seungwoo
Publicado: (2026)
"I Choose to Live, for Life Itself": Understanding Agency of Home-Based Care Patients Through Information Practices and Relational Dynamics in Care Networks
por: Kim, Sung-In, et al.
Publicado: (2026)
por: Kim, Sung-In, et al.
Publicado: (2026)
Residual MPC: Blending Reinforcement Learning with GPU-Parallelized Model Predictive Control
por: Jeon, Se Hwan, et al.
Publicado: (2025)
por: Jeon, Se Hwan, et al.
Publicado: (2025)
Ejemplares similares
-
BoA: Attention-aware Post-training Quantization without Backpropagation
por: Kim, Junhan, et al.
Publicado: (2024) -
Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers
por: Kim, Junhan, et al.
Publicado: (2024) -
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
por: Kim, Junhan, et al.
Publicado: (2026) -
On the Importance of a Multi-Scale Calibration for Quantization
por: Son, Seungwoo, et al.
Publicado: (2026) -
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
por: Ahn, Jinwoo, et al.
Publicado: (2026)