Salvato in:
| Autori principali: | Nair, Pranav, Datta, Puranjay, Dean, Jeff, Jain, Prateek, Kusupati, Aditya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.06786 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Matryoshka Model Learning for Improved Elastic Student Models
di: Verma, Chetan, et al.
Pubblicazione: (2025)
di: Verma, Chetan, et al.
Pubblicazione: (2025)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
di: Jain, Gagan, et al.
Pubblicazione: (2024)
di: Jain, Gagan, et al.
Pubblicazione: (2024)
Matryoshka Representation Learning
di: Kusupati, Aditya, et al.
Pubblicazione: (2022)
di: Kusupati, Aditya, et al.
Pubblicazione: (2022)
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
di: Nair, Pranav Ajit, et al.
Pubblicazione: (2024)
di: Nair, Pranav Ajit, et al.
Pubblicazione: (2024)
Compressing Many-Shots in In-Context Learning
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)
MatMamba: A Matryoshka State Space Model
di: Shukla, Abhinav, et al.
Pubblicazione: (2024)
di: Shukla, Abhinav, et al.
Pubblicazione: (2024)
Estimating Causal Effects in Gaussian Linear SCMs with Finite Data
di: Maiti, Aurghya, et al.
Pubblicazione: (2026)
di: Maiti, Aurghya, et al.
Pubblicazione: (2026)
Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs
di: Kaliaperumal, Pranav Kumar
Pubblicazione: (2026)
di: Kaliaperumal, Pranav Kumar
Pubblicazione: (2026)
Learning Multi-Level Features with Matryoshka Sparse Autoencoders
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
di: Bussmann, Bart, et al.
Pubblicazione: (2025)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
di: Goru, Ritesh, et al.
Pubblicazione: (2025)
di: Goru, Ritesh, et al.
Pubblicazione: (2025)
Runtime-Certified Bounded-Error Quantized Attention
di: Calver, Dean
Pubblicazione: (2026)
di: Calver, Dean
Pubblicazione: (2026)
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
di: Kumar, Ramnath, et al.
Pubblicazione: (2023)
di: Kumar, Ramnath, et al.
Pubblicazione: (2023)
That Chip Has Sailed: A Critique of Unfounded Skepticism Around AI for Chip Design
di: Goldie, Anna, et al.
Pubblicazione: (2024)
di: Goldie, Anna, et al.
Pubblicazione: (2024)
Arithmetic-Intensity-Aware Quantization
di: Singh, Taig, et al.
Pubblicazione: (2025)
di: Singh, Taig, et al.
Pubblicazione: (2025)
Matryoshka Multimodal Models
di: Cai, Mu, et al.
Pubblicazione: (2024)
di: Cai, Mu, et al.
Pubblicazione: (2024)
Transitive RL: Value Learning via Divide and Conquer
di: Park, Seohong, et al.
Pubblicazione: (2025)
di: Park, Seohong, et al.
Pubblicazione: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
di: Li, Changhao, et al.
Pubblicazione: (2024)
di: Li, Changhao, et al.
Pubblicazione: (2024)
Interleaved Gibbs Diffusion: Generating Discrete-Continuous Data with Implicit Constraints
di: Anil, Gautham Govind, et al.
Pubblicazione: (2025)
di: Anil, Gautham Govind, et al.
Pubblicazione: (2025)
RSQ: Learning from Important Tokens Leads to Better Quantized LLMs
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
di: Sung, Yi-Lin, et al.
Pubblicazione: (2025)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
di: Lin, Bokai, et al.
Pubblicazione: (2024)
di: Lin, Bokai, et al.
Pubblicazione: (2024)
Succeeding at Scale: Automated Dataset Construction and Query-Side Adaptation for Multi-Tenant Search
di: Jain, Prateek, et al.
Pubblicazione: (2026)
di: Jain, Prateek, et al.
Pubblicazione: (2026)
LookupViT: Compressing visual information to a limited number of tokens
di: Koner, Rajat, et al.
Pubblicazione: (2024)
di: Koner, Rajat, et al.
Pubblicazione: (2024)
Machine Learning-Based Security Policy Analysis
di: Jain, Krish, et al.
Pubblicazione: (2024)
di: Jain, Krish, et al.
Pubblicazione: (2024)
Systematic Optimization of Open Source Large Language Models for Mathematical Reasoning
di: Pawar, Pranav, et al.
Pubblicazione: (2025)
di: Pawar, Pranav, et al.
Pubblicazione: (2025)
Masked Generative Nested Transformers with Decode Time Scaling
di: Goyal, Sahil, et al.
Pubblicazione: (2025)
di: Goyal, Sahil, et al.
Pubblicazione: (2025)
Anytime Pretraining: Horizon-Free Learning-Rate Schedules with Weight Averaging
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
di: Meterez, Alexandru, et al.
Pubblicazione: (2026)
An Efficient Sparse Fine-Tuning with Low Quantization Error via Neural Network Pruning
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
di: Li, Cen-Jhih, et al.
Pubblicazione: (2025)
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
di: Modoranu, Ionut-Vlad, et al.
Pubblicazione: (2026)
Beyond Matryoshka: Revisiting Sparse Coding for Adaptive Representation
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
di: Wen, Tiansheng, et al.
Pubblicazione: (2025)
HiRE: High Recall Approximate Top-$k$ Estimation for Efficient LLM Inference
di: L, Yashas Samaga B, et al.
Pubblicazione: (2024)
di: L, Yashas Samaga B, et al.
Pubblicazione: (2024)
Mathematical Framework for Custom Reward Functions in Job Application Evaluation using Reinforcement Learning
di: Jain, Shreyansh, et al.
Pubblicazione: (2025)
di: Jain, Shreyansh, et al.
Pubblicazione: (2025)
SolarTformer: A Transformer Based Deep Learning Approach for Short Term Solar Power Forecasting
di: Basu, Ankan, et al.
Pubblicazione: (2026)
di: Basu, Ankan, et al.
Pubblicazione: (2026)
Softmax is $1/2$-Lipschitz: A tight bound across all $\ell_p$ norms
di: Nair, Pravin
Pubblicazione: (2025)
di: Nair, Pravin
Pubblicazione: (2025)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
di: Nair, Lakshmi
Pubblicazione: (2024)
di: Nair, Lakshmi
Pubblicazione: (2024)
QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
di: Tiwari, Rishabh, et al.
Pubblicazione: (2025)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
di: Ghadia, Ravi, et al.
Pubblicazione: (2025)
di: Ghadia, Ravi, et al.
Pubblicazione: (2025)
Assigning Credit with Partial Reward Decoupling in Multi-Agent Proximal Policy Optimization
di: Kapoor, Aditya, et al.
Pubblicazione: (2024)
di: Kapoor, Aditya, et al.
Pubblicazione: (2024)
Soft $Q(λ)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
di: Mahajan, Pranav, et al.
Pubblicazione: (2026)
di: Mahajan, Pranav, et al.
Pubblicazione: (2026)
Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models
di: Kallem, Pranav
Pubblicazione: (2026)
di: Kallem, Pranav
Pubblicazione: (2026)
Documenti analoghi
-
Matryoshka Model Learning for Improved Elastic Student Models
di: Verma, Chetan, et al.
Pubblicazione: (2025) -
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
di: Jain, Gagan, et al.
Pubblicazione: (2024) -
Matryoshka Representation Learning
di: Kusupati, Aditya, et al.
Pubblicazione: (2022) -
CDQuant: Greedy Coordinate Descent for Accurate LLM Quantization
di: Nair, Pranav Ajit, et al.
Pubblicazione: (2024) -
Compressing Many-Shots in In-Context Learning
di: Khatri, Devvrit, et al.
Pubblicazione: (2025)