Matryoshka Quantization
Fuente:
arXiv
Saved in:
| Main Authors: | Nair, Pranav, Datta, Puranjay, Dean, Jeff, Jain, Prateek, Kusupati, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Matryoshka Model Learning for Improved Elastic Student Models
by: Verma, Chetan, et al.
Published: (2025)
by: Verma, Chetan, et al.
Published: (2025)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024)
by: Jain, Gagan, et al.
Published: (2024)
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022)
by: Kusupati, Aditya, et al.
Published: (2022)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
by: Nair, Pranav Ajit, et al.
Published: (2024)
by: Nair, Pranav Ajit, et al.
Published: (2024)
Compressing Many-Shots in In-Context Learning
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
Estimating Causal Effects in Gaussian Linear SCMs with Finite Data
by: Maiti, Aurghya, et al.
Published: (2026)
by: Maiti, Aurghya, et al.
Published: (2026)
MatMamba: A Matryoshka State Space Model
by: Shukla, Abhinav, et al.
Published: (2024)
by: Shukla, Abhinav, et al.
Published: (2024)
Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs
by: Kaliaperumal, Pranav Kumar
Published: (2026)
by: Kaliaperumal, Pranav Kumar
Published: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
by: Bussmann, Bart, et al.
Published: (2025)
by: Bussmann, Bart, et al.
Published: (2025)
Runtime-Certified Bounded-Error Quantized Attention
by: Calver, Dean
Published: (2026)
by: Calver, Dean
Published: (2026)
Arithmetic-Intensity-Aware Quantization
by: Singh, Taig, et al.
Published: (2025)
by: Singh, Taig, et al.
Published: (2025)
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
by: Goldie, Anna, et al.
Published: (2024)
by: Goldie, Anna, et al.
Published: (2024)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Transitive RL: Value Learning via Divide and Conquer
by: Park, Seohong, et al.
Published: (2025)
by: Park, Seohong, et al.
Published: (2025)
Matryoshka Multimodal Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
by: Li, Changhao, et al.
Published: (2024)
by: Li, Changhao, et al.
Published: (2024)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Interleaved Gibbs Diffusion: Generating Discrete-Continuous Data with Implicit Constraints
by: Anil, Gautham Govind, et al.
Published: (2025)
by: Anil, Gautham Govind, et al.
Published: (2025)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
by: Lin, Bokai, et al.
Published: (2024)
by: Lin, Bokai, et al.
Published: (2024)
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
by: Kumar, Ramnath, et al.
Published: (2023)
by: Kumar, Ramnath, et al.
Published: (2023)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
by: Sung, Yi-Lin, et al.
Published: (2025)
by: Sung, Yi-Lin, et al.
Published: (2025)
LookupViT: Compressing visual information to a limited number of tokens
by: Koner, Rajat, et al.
Published: (2024)
by: Koner, Rajat, et al.
Published: (2024)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Machine Learning-Based Security Policy Analysis
by: Jain, Krish, et al.
Published: (2024)
by: Jain, Krish, et al.
Published: (2024)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
by: Li, Cen-Jhih, et al.
Published: (2025)
by: Li, Cen-Jhih, et al.
Published: (2025)
Succeeding at Scale: Automated Dataset Construction and Query-Side Adaptation for Multi-Tenant Search
by: Jain, Prateek, et al.
Published: (2026)
by: Jain, Prateek, et al.
Published: (2026)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
by: L, Yashas Samaga B, et al.
Published: (2024)
by: L, Yashas Samaga B, et al.
Published: (2024)
Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
by: Nair, Pravin
Published: (2025)
by: Nair, Pravin
Published: (2025)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
by: Tiwari, Rishabh, et al.
Published: (2025)
by: Tiwari, Rishabh, et al.
Published: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
by: Meterez, Alexandru, et al.
Published: (2026)
by: Meterez, Alexandru, et al.
Published: (2026)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
by: Wen, Tiansheng, et al.
Published: (2025)
by: Wen, Tiansheng, et al.
Published: (2025)
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025)
by: Goyal, Sahil, et al.
Published: (2025)
Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
by: Jain, Shreyansh, et al.
Published: (2025)
by: Jain, Shreyansh, et al.
Published: (2025)
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
by: Mahajan, Pranav, et al.
Published: (2026)
by: Mahajan, Pranav, et al.
Published: (2026)
SolarTformer: A Transformer Based Deep Learning Approach for Short Term Solar Power Forecasting
by: Basu, Ankan, et al.
Published: (2026)
by: Basu, Ankan, et al.
Published: (2026)
How to Parameterize Asymmetric Quantization Ranges for Quantization-Aware Training
by: You, Jaeseong, et al.
Published: (2024)
by: You, Jaeseong, et al.
Published: (2024)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
by: Ghadia, Ravi, et al.
Published: (2025)
by: Ghadia, Ravi, et al.
Published: (2025)
Preserve-Then-Quantize: Balancing Rank Budgets for Quantization Error Reconstruction in LLMs
by: Cho, Yoonjun, et al.
Published: (2026)
by: Cho, Yoonjun, et al.
Published: (2026)
Similar Items
-
Matryoshka Model Learning for Improved Elastic Student Models
by: Verma, Chetan, et al.
Published: (2025) -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
by: Jain, Gagan, et al.
Published: (2024) -
Matryoshka Representation Learning
by: Kusupati, Aditya, et al.
Published: (2022) -
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
by: Nair, Pranav Ajit, et al.
Published: (2024) -
Compressing Many-Shots in In-Context Learning
by: Khatri, Devvrit, et al.
Published: (2025)