TTQ: Activation-Aware Test-Time Quantization to Accelerate LLM Inference On The Fly
Fuente:
arXiv
Guardado en:
| Autores principales: | Koike-Akino, Toshiaki, Liu, Jing, Wang, Ye |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
Directional Embedding Smoothing for Robust Vision Language Models
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
LatentLLM: Attention-Aware Joint Tensor Compression
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025)
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
por: Liu, Jing, et al.
Publicado: (2025)
por: Liu, Jing, et al.
Publicado: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
por: Liu, Sheng, et al.
Publicado: (2025)
por: Liu, Sheng, et al.
Publicado: (2025)
Geo-ADAPT-VQE: Quantum Information Metric-Aware Circuit Optimization for Quantum Chemistry
por: Sohail, Mohammad Aamir, et al.
Publicado: (2026)
por: Sohail, Mohammad Aamir, et al.
Publicado: (2026)
Quantum Implicit Neural Compression
por: Fujihashi, Takuya, et al.
Publicado: (2024)
por: Fujihashi, Takuya, et al.
Publicado: (2024)
Amplification Effects in Test-Time Reinforcement Learning: Safety and Reasoning Vulnerabilities
por: Khattar, Vanshaj, et al.
Publicado: (2026)
por: Khattar, Vanshaj, et al.
Publicado: (2026)
MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation
por: Wan, Zhongwei, et al.
Publicado: (2024)
por: Wan, Zhongwei, et al.
Publicado: (2024)
Quantization-Aware Collaborative Inference for Large Embodied AI Models
por: Lyu, Zhonghao, et al.
Publicado: (2026)
por: Lyu, Zhonghao, et al.
Publicado: (2026)
Winning Big with Small Models: Knowledge Distillation vs. Self-Training for Reducing Hallucination in Product QA Agents
por: Lewis, Ashley, et al.
Publicado: (2025)
por: Lewis, Ashley, et al.
Publicado: (2025)
Can VLM Pseudo-Labels Train a Time-Series QA Model That Outperforms the VLM?
por: Fujimura, Takuya, et al.
Publicado: (2025)
por: Fujimura, Takuya, et al.
Publicado: (2025)
When marine radar target detection meets pretrained large language models
por: Hu, Qiying, et al.
Publicado: (2025)
por: Hu, Qiying, et al.
Publicado: (2025)
TQ-DiT: Efficient Time-Aware Quantization for Diffusion Transformers
por: Hwang, Younghye, et al.
Publicado: (2025)
por: Hwang, Younghye, et al.
Publicado: (2025)
Dynamic Fault Characteristics Evaluation in Power Grid
por: Pei, Hao, et al.
Publicado: (2023)
por: Pei, Hao, et al.
Publicado: (2023)
LLMs as High-Dimensional Nonlinear Autoregressive Models with Attention: Training, Alignment and Inference
por: Krishnamurthy, Vikram
Publicado: (2026)
por: Krishnamurthy, Vikram
Publicado: (2026)
Beyond Sinusoids: A Morlet Wavelet Framework for Transformer Positional Encoding
por: Zeris, Athanasios
Publicado: (2026)
por: Zeris, Athanasios
Publicado: (2026)
Energy-Gated Attention and Wavelet Positional Encoding: Complementary Inductive Biases for Transformer Attention
por: Zeris, Athanasios
Publicado: (2026)
por: Zeris, Athanasios
Publicado: (2026)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
por: Zeris, Athanasios
Publicado: (2026)
por: Zeris, Athanasios
Publicado: (2026)
A Graph Signal Processing Framework for Hallucination Detection in Large Language Models
por: Noël, Valentin
Publicado: (2025)
por: Noël, Valentin
Publicado: (2025)
Bridging Brain Signals and Language: A Deep Learning Approach to EEG-to-Text Decoding
por: Gedawy, Mostafa El, et al.
Publicado: (2025)
por: Gedawy, Mostafa El, et al.
Publicado: (2025)
EEG-CLIP : Learning EEG representations from natural language descriptions
por: Ndir, Tidiane Camaret, et al.
Publicado: (2025)
por: Ndir, Tidiane Camaret, et al.
Publicado: (2025)
Lightweight Conceptual Dictionary Learning for Text Classification Using Information Compression
por: Wan, Li, et al.
Publicado: (2024)
por: Wan, Li, et al.
Publicado: (2024)
Text me the data: Generating Ground Pressure Sequence from Textual Descriptions for HAR
por: Ray, Lala Shakti Swarup, et al.
Publicado: (2024)
por: Ray, Lala Shakti Swarup, et al.
Publicado: (2024)
Training-Free Spectral Fingerprints of Voice Processing in Transformers
por: Noël, Valentin
Publicado: (2025)
por: Noël, Valentin
Publicado: (2025)
Random Channel Ablation for Robust Hand Gesture Classification with Multimodal Biosignals
por: Bimbraw, Keshav, et al.
Publicado: (2024)
por: Bimbraw, Keshav, et al.
Publicado: (2024)
NeuroNarrator: A Generalist EEG-to-Text Foundation Model for Clinical Interpretation via Spectro-Spatial Grounding and Temporal State-Space Reasoning
por: Wang, Guoan, et al.
Publicado: (2026)
por: Wang, Guoan, et al.
Publicado: (2026)
sEEG-based Encoding for Sentence Retrieval: A Contrastive Learning Approach to Brain-Language Alignment
por: Liu, Yijun
Publicado: (2025)
por: Liu, Yijun
Publicado: (2025)
ECG-Expert-QA: A Benchmark for Evaluating Medical Large Language Models in Heart Disease Diagnosis
por: Wang, Xu, et al.
Publicado: (2025)
por: Wang, Xu, et al.
Publicado: (2025)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
por: Sankar, Ashwin, et al.
Publicado: (2024)
por: Sankar, Ashwin, et al.
Publicado: (2024)
Power-of-Two Quantization-Aware-Training (PoT-QAT) in Large Language Models (LLMs)
por: Elgenedy, Mahmoud
Publicado: (2026)
por: Elgenedy, Mahmoud
Publicado: (2026)
Multi-Band Wi-Fi Neural Dynamic Fusion
por: Kato, Sorachi, et al.
Publicado: (2024)
por: Kato, Sorachi, et al.
Publicado: (2024)
GPT Sonograpy: Hand Gesture Decoding from Forearm Ultrasound Images via VLM
por: Bimbraw, Keshav, et al.
Publicado: (2024)
por: Bimbraw, Keshav, et al.
Publicado: (2024)
SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning
por: Cai, Mingsheng, et al.
Publicado: (2025)
por: Cai, Mingsheng, et al.
Publicado: (2025)
Enabling ASR for Low-Resource Languages: A Comprehensive Dataset Creation Approach
por: Yeroyan, Ara, et al.
Publicado: (2024)
por: Yeroyan, Ara, et al.
Publicado: (2024)
EEG-Based Brain-LLM Interface for Human Preference Aligned Generation
por: Zhang, Junzi, et al.
Publicado: (2026)
por: Zhang, Junzi, et al.
Publicado: (2026)
Conformal Inference for Time Series over Graphs
por: Dua, Sonakshi, et al.
Publicado: (2025)
por: Dua, Sonakshi, et al.
Publicado: (2025)
Sequential Inference for Gaussian Processes: A Signal Processing Perspective
por: Waxman, Daniel, et al.
Publicado: (2026)
por: Waxman, Daniel, et al.
Publicado: (2026)
Leveraging Large Language Models for Wireless Symbol Detection via In-Context Learning
por: Abbas, Momin, et al.
Publicado: (2024)
por: Abbas, Momin, et al.
Publicado: (2024)
Ejemplares similares
-
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026) -
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025) -
Directional Embedding Smoothing for Robust Vision Language Models
por: Wang, Ye, et al.
Publicado: (2026) -
LatentLLM: Attention-Aware Joint Tensor Compression
por: Koike-Akino, Toshiaki, et al.
Publicado: (2025) -
AWP: Activation-Aware Weight Pruning and Quantization with Projected Gradient Descent
por: Liu, Jing, et al.
Publicado: (2025)