QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Noh, Kanghyun, Choi, Jinheon, Kim, Yulhwa |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
di: Jo, Dongwon, et al.
Pubblicazione: (2024)
di: Jo, Dongwon, et al.
Pubblicazione: (2024)
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
di: Jeon, Hyesung, et al.
Pubblicazione: (2024)
di: Jeon, Hyesung, et al.
Pubblicazione: (2024)
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
di: Kim, Jiyoon, et al.
Pubblicazione: (2025)
Optimizing LLMs Using Quantization for Mobile Execution
di: Yadav, Agatsya, et al.
Pubblicazione: (2025)
di: Yadav, Agatsya, et al.
Pubblicazione: (2025)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
di: Song, Jiwon, et al.
Pubblicazione: (2024)
di: Song, Jiwon, et al.
Pubblicazione: (2024)
Robust Hallucination Detection in LLMs via Adaptive Token Selection
di: Niu, Mengjia, et al.
Pubblicazione: (2025)
di: Niu, Mengjia, et al.
Pubblicazione: (2025)
Adaptive Teaching in Heterogeneous Agents: Balancing Surprise in Sparse Reward Scenarios
di: Clark, Emma, et al.
Pubblicazione: (2024)
di: Clark, Emma, et al.
Pubblicazione: (2024)
FastKV: Decoupling of Context Reduction and KV Cache Compression for Prefill-Decoding Acceleration
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
di: Jo, Dongwon, et al.
Pubblicazione: (2025)
Mitigating Premature Discretization with Progressive Quantization for Robust Vector Tokenization
di: Zhao, Wenhao, et al.
Pubblicazione: (2026)
di: Zhao, Wenhao, et al.
Pubblicazione: (2026)
Low-Bit Quantization Favors Undertrained LLMs: Scaling Laws for Quantized LLMs with 100T Training Tokens
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
di: Ouyang, Xu, et al.
Pubblicazione: (2024)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
di: Choi, Yumin, et al.
Pubblicazione: (2025)
di: Choi, Yumin, et al.
Pubblicazione: (2025)
MimiQ: Low-Bit Data-Free Quantization of Vision Transformers with Encouraging Inter-Head Attention Similarity
di: Choi, Kanghyun, et al.
Pubblicazione: (2024)
di: Choi, Kanghyun, et al.
Pubblicazione: (2024)
Adaptive Branch Specialization in Spectral-Spatial Graph Neural Networks for Certified Robustness
di: Choi, Yoonhyuk, et al.
Pubblicazione: (2025)
di: Choi, Yoonhyuk, et al.
Pubblicazione: (2025)
Importance Analysis for Dynamic Control of Balancing Parameter in a Simple Knowledge Distillation Setting
di: Kim, Seongmin, et al.
Pubblicazione: (2025)
di: Kim, Seongmin, et al.
Pubblicazione: (2025)
System Prompt Optimization with Meta-Learning
di: Choi, Yumin, et al.
Pubblicazione: (2025)
di: Choi, Yumin, et al.
Pubblicazione: (2025)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
STILL: Selecting Tokens for Intra-Layer Hybrid Attention to Linearize LLMs
di: Meng, Weikang, et al.
Pubblicazione: (2026)
di: Meng, Weikang, et al.
Pubblicazione: (2026)
Let the Void Be Void: Robust Open-Set Semi-Supervised Learning via Selective Non-Alignment
di: Choi, You Rim, et al.
Pubblicazione: (2025)
di: Choi, You Rim, et al.
Pubblicazione: (2025)
Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching
di: Aytes, Simon A., et al.
Pubblicazione: (2025)
di: Aytes, Simon A., et al.
Pubblicazione: (2025)
ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning
di: Lee, Yuna, et al.
Pubblicazione: (2026)
di: Lee, Yuna, et al.
Pubblicazione: (2026)
ASMR: Adaptive Skeleton-Mesh Rigging and Skinning via 2D Generative Prior
di: Hong, Seokhyeon, et al.
Pubblicazione: (2025)
di: Hong, Seokhyeon, et al.
Pubblicazione: (2025)
Adaptive Layer-Wise Transformations for Post-Training Quantization of Large Language Models
di: Pham, Cuong, et al.
Pubblicazione: (2025)
di: Pham, Cuong, et al.
Pubblicazione: (2025)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
di: Taniguchi, Rei, et al.
Pubblicazione: (2026)
MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
di: Wang, Dongwei, et al.
Pubblicazione: (2026)
LAQuant: A Simple Overhead-free Large Reasoning Model Quantization by Layer-wise Lookahead Loss
di: Choi, Euntae, et al.
Pubblicazione: (2026)
di: Choi, Euntae, et al.
Pubblicazione: (2026)
How Robustly do LLMs Understand Execution Semantics?
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
di: Spiess, Claudio, et al.
Pubblicazione: (2026)
Adaptive Token-Weighted Differential Privacy for LLMs: Not All Tokens Require Equal Protection
di: Yu, Manjiang, et al.
Pubblicazione: (2025)
di: Yu, Manjiang, et al.
Pubblicazione: (2025)
APreQEL: Adaptive Mixed Precision Quantization For Edge LLMs
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
di: Bouzouad, Meriem, et al.
Pubblicazione: (2026)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
di: Cho, Yoonjun, et al.
Pubblicazione: (2026)
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
di: Park, Yeonsik, et al.
Pubblicazione: (2026)
di: Park, Yeonsik, et al.
Pubblicazione: (2026)
NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
di: Chong, Hyochan, et al.
Pubblicazione: (2026)
di: Chong, Hyochan, et al.
Pubblicazione: (2026)
LayerShuffle: Enhancing Robustness in Vision Transformers by Randomizing Layer Execution Order
di: Freiberger, Matthias, et al.
Pubblicazione: (2024)
di: Freiberger, Matthias, et al.
Pubblicazione: (2024)
FALQON: Accelerating LoRA Fine-tuning with Low-Bit Floating-Point Arithmetic
di: Choi, Kanghyun, et al.
Pubblicazione: (2025)
di: Choi, Kanghyun, et al.
Pubblicazione: (2025)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
di: Noh, Kangjun, et al.
Pubblicazione: (2026)
di: Noh, Kangjun, et al.
Pubblicazione: (2026)
Mitigating Quantization Errors Due to Activation Spikes in GLU-Based LLMs
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
di: Yang, Jaewoo, et al.
Pubblicazione: (2024)
SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
di: Choi, Hyeonbeom, et al.
Pubblicazione: (2026)
di: Choi, Hyeonbeom, et al.
Pubblicazione: (2026)
Two-Stage Grid Optimization for Group-wise Quantization of LLMs
di: Kim, Junhan, et al.
Pubblicazione: (2026)
di: Kim, Junhan, et al.
Pubblicazione: (2026)
Pyramid Vector Quantization for LLMs
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
di: van der Ouderaa, Tycho F. A., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Mixture of Scales: Memory-Efficient Token-Adaptive Binarization for Large Language Models
di: Jo, Dongwon, et al.
Pubblicazione: (2024) -
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
di: Jeon, Hyesung, et al.
Pubblicazione: (2024) -
Column-wise Quantization of Weights and Partial Sums for Accurate and Efficient Compute-In-Memory Accelerators
di: Kim, Jiyoon, et al.
Pubblicazione: (2025) -
Optimizing LLMs Using Quantization for Mobile Execution
di: Yadav, Agatsya, et al.
Pubblicazione: (2025) -
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
di: Song, Jiwon, et al.
Pubblicazione: (2024)