Affine-Scaled Attention: Towards Flexible and Stable Transformer Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Bae, Jeongin, Park, Baeseong, Park, Gunho, Kim, Minsub, Lee, Joonhyung, Yoo, Junhee, Woo, Sunghyeon, Ryu, Jiwon, Kwon, Se Jung, Lee, Dongsoo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
por: Park, Gunho, et al.
Publicado: (2025)
por: Park, Gunho, et al.
Publicado: (2025)
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
por: Yoon, Kanghoon, et al.
Publicado: (2025)
por: Yoon, Kanghoon, et al.
Publicado: (2025)
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
por: Park, Gunho, et al.
Publicado: (2022)
por: Park, Gunho, et al.
Publicado: (2022)
To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
por: Lee, Joonhyung, et al.
Publicado: (2024)
por: Lee, Joonhyung, et al.
Publicado: (2024)
ICaRus: Identical Cache Reuse for Efficient Multi Model Inference
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
DropBP: Accelerating Fine-Tuning of Large Language Models by Dropping Backward Propagation
por: Woo, Sunghyeon, et al.
Publicado: (2024)
por: Woo, Sunghyeon, et al.
Publicado: (2024)
Training-free Dropout Sampling for Semantic Token Acceptance in Speculative Decoding
por: Lee, Jeongtae, et al.
Publicado: (2026)
por: Lee, Jeongtae, et al.
Publicado: (2026)
An Inquiry into Datacenter TCO for LLM Inference with FP8
por: Kim, Jiwoo, et al.
Publicado: (2025)
por: Kim, Jiwoo, et al.
Publicado: (2025)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
No Token Left Behind: Reliable KV Cache Compression via Importance-Aware Mixed Precision Quantization
por: Yang, June Yong, et al.
Publicado: (2024)
por: Yang, June Yong, et al.
Publicado: (2024)
PrefillShare: A Shared Prefill Module for KV Reuse in Multi-LLM Disaggregated Serving
por: Woo, Sunghyeon, et al.
Publicado: (2026)
por: Woo, Sunghyeon, et al.
Publicado: (2026)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
por: Park, Seungcheol, et al.
Publicado: (2025)
por: Park, Seungcheol, et al.
Publicado: (2025)
Faster Inference of LLMs using FP8 on the Intel Gaudi
por: Lee, Joonhyung, et al.
Publicado: (2025)
por: Lee, Joonhyung, et al.
Publicado: (2025)
FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
por: Lee, Jung Hyun, et al.
Publicado: (2023)
por: Lee, Jung Hyun, et al.
Publicado: (2023)
Attention-Based Reading, Highlighting, and Forecasting of the Limit Order Book
por: Jung, Jiwon, et al.
Publicado: (2024)
por: Jung, Jiwon, et al.
Publicado: (2024)
LRQ: Optimizing Post-Training Quantization for Large Language Models by Learning Low-Rank Weight-Scaling Matrices
por: Lee, Jung Hyun, et al.
Publicado: (2024)
por: Lee, Jung Hyun, et al.
Publicado: (2024)
Debunking the CUDA Myth Towards GPU-based AI Systems
por: Lee, Yunjae, et al.
Publicado: (2024)
por: Lee, Yunjae, et al.
Publicado: (2024)
Rethinking Channel Dimensions to Isolate Outliers for Low-bit Weight Quantization of Large Language Models
por: Heo, Jung Hwan, et al.
Publicado: (2023)
por: Heo, Jung Hwan, et al.
Publicado: (2023)
Attention-aware Semantic Communications for Collaborative Inference
por: Im, Jiwoong, et al.
Publicado: (2024)
por: Im, Jiwoong, et al.
Publicado: (2024)
Organic Ferroelectrics for Regulation of Electronic and Ionic Transport Toward Neuromorphic Applications
por: Minsub Lee, et al.
Publicado: (2024)
por: Minsub Lee, et al.
Publicado: (2024)
BioArtlas: Mapping Bioart's Polysemy with an Embedding-Derived Lexicon
por: Bae, Joonhyung
Publicado: (2025)
por: Bae, Joonhyung
Publicado: (2025)
BioArtlas: Computational Clustering of Multi‑Dimensional Complexity in Bioart
por: Bae, Joonhyung
Publicado: (2025)
por: Bae, Joonhyung
Publicado: (2025)
BioArtlas: Computational Clustering of Multi-Dimensional Complexity in Bioart
por: Bae, Joonhyung
Publicado: (2025)
por: Bae, Joonhyung
Publicado: (2025)
ASTRA: Mapping Art-Technology Institutions via Conceptual Axes, Text Embeddings, and Unsupervised Clustering
por: Bae, Joonhyung
Publicado: (2026)
por: Bae, Joonhyung
Publicado: (2026)
Thief of Truth: VR comics about the relationship between AI and humans
por: Bae, Joonhyung
Publicado: (2025)
por: Bae, Joonhyung
Publicado: (2025)
Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier
por: Bae, Joonhyung
Publicado: (2026)
por: Bae, Joonhyung
Publicado: (2026)
Addressing Low E/S and N/P Ratio Challenges in Li–S Batteries with a Multifunctional Interlayer
por: Cinthya Paulina, et al.
Publicado: (2026)
por: Cinthya Paulina, et al.
Publicado: (2026)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
por: Park, Geon, et al.
Publicado: (2025)
por: Park, Geon, et al.
Publicado: (2025)
Visual Preference Inference: An Image Sequence-Based Preference Reasoning in Tabletop Object Manipulation
por: Lee, Joonhyung, et al.
Publicado: (2024)
por: Lee, Joonhyung, et al.
Publicado: (2024)
Where to invest? Effects of technological capabilities on corporate venture capital investments
por: Wonsang Ryu, et al.
Publicado: (2024)
por: Wonsang Ryu, et al.
Publicado: (2024)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
por: Choi, Kanghyun, et al.
Publicado: (2024)
por: Choi, Kanghyun, et al.
Publicado: (2024)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
por: Kim, Sohee, et al.
Publicado: (2025)
por: Kim, Sohee, et al.
Publicado: (2025)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
por: Back, Kyungryul, et al.
Publicado: (2025)
por: Back, Kyungryul, et al.
Publicado: (2025)
Physical Containers as Framing Conditions for Visualization in Augmented Reality
por: Bae, Jiyeon, et al.
Publicado: (2026)
por: Bae, Jiyeon, et al.
Publicado: (2026)
Serial Quantitative Evaluation of Load Redistribution and Osteotomy Gap After Medial Open‐Wedge High Tibial Osteotomy
por: Jae Woo Cho, et al.
Publicado: (2025)
por: Jae Woo Cho, et al.
Publicado: (2025)
LampQ: Towards Accurate Layer-wise Mixed Precision Quantization for Vision Transformers
por: Kim, Minjun, et al.
Publicado: (2025)
por: Kim, Minjun, et al.
Publicado: (2025)
PAFormer: Part Aware Transformer for Person Re-identification
por: Jung, Hyeono, et al.
Publicado: (2024)
por: Jung, Hyeono, et al.
Publicado: (2024)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
por: Oh, Sehyeon, et al.
Publicado: (2026)
por: Oh, Sehyeon, et al.
Publicado: (2026)
Ejemplares similares
-
CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMs
por: Park, Gunho, et al.
Publicado: (2025) -
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
por: Park, Gunho, et al.
Publicado: (2025) -
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
por: Park, Gunho, et al.
Publicado: (2025) -
SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification
por: Yoon, Kanghoon, et al.
Publicado: (2025) -
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
por: Park, Gunho, et al.
Publicado: (2022)