Matryoshka Model Learning for Improved Elastic Student Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Verma, Chetan, Timmaraju, Aditya Srinivas, Hsieh, Cho-Jui, Damle, Suyash, Bui, Ngot, Zhang, Yang, Chen, Wen, Liu, Xin, Jain, Prateek, Dhillon, Inderjit S |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
por: Gupta, Nilesh, et al.
Publicado: (2025)
por: Gupta, Nilesh, et al.
Publicado: (2025)
Matryoshka Quantization
por: Nair, Pranav, et al.
Publicado: (2025)
por: Nair, Pranav, et al.
Publicado: (2025)
Large Language Models are Interpretable Learners
por: Wang, Ruochen, et al.
Publicado: (2024)
por: Wang, Ruochen, et al.
Publicado: (2024)
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
por: Kumar, Ramnath, et al.
Publicado: (2023)
por: Kumar, Ramnath, et al.
Publicado: (2023)
FastLane: Efficient Routed Systems for Late-Interaction Retrieval
por: Kumar, Ramnath, et al.
Publicado: (2026)
por: Kumar, Ramnath, et al.
Publicado: (2026)
Compressing Many-Shots in In-Context Learning
por: Khatri, Devvrit, et al.
Publicado: (2025)
por: Khatri, Devvrit, et al.
Publicado: (2025)
Matryoshka Representation Learning
por: Kusupati, Aditya, et al.
Publicado: (2022)
por: Kusupati, Aditya, et al.
Publicado: (2022)
MatFormer: Nested Transformer for Elastic Inference
por: Devvrit, et al.
Publicado: (2023)
por: Devvrit, et al.
Publicado: (2023)
Dual-Encoders for Extreme Multi-Label Classification
por: Gupta, Nilesh, et al.
Publicado: (2023)
por: Gupta, Nilesh, et al.
Publicado: (2023)
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
por: Yen, Jui-Nan, et al.
Publicado: (2024)
por: Yen, Jui-Nan, et al.
Publicado: (2024)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
por: Wang, Yihan, et al.
Publicado: (2022)
por: Wang, Yihan, et al.
Publicado: (2022)
LASER: Attention with Exponential Transformation
por: Duvvuri, Sai Surya, et al.
Publicado: (2024)
por: Duvvuri, Sai Surya, et al.
Publicado: (2024)
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review
por: Prakriya, Neha, et al.
Publicado: (2024)
por: Prakriya, Neha, et al.
Publicado: (2024)
AI Co-Scientist for Ranking: Discovering Novel Search Ranking Models alongside LLM-based AI Agents with Cloud Computing Access
por: Wu, Liwei, et al.
Publicado: (2026)
por: Wu, Liwei, et al.
Publicado: (2026)
Wavelet GPT: Wavelet Inspired Large Language Models
por: Verma, Prateek
Publicado: (2024)
por: Verma, Prateek
Publicado: (2024)
A Language Model With Million Context Length For Raw Audio
por: Verma, Prateek
Publicado: (2022)
por: Verma, Prateek
Publicado: (2022)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
por: Verma, Prateek
Publicado: (2024)
por: Verma, Prateek
Publicado: (2024)
Matryoshka Multimodal Models
por: Cai, Mu, et al.
Publicado: (2024)
por: Cai, Mu, et al.
Publicado: (2024)
ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization
por: Patel, Nirmal, et al.
Publicado: (2026)
por: Patel, Nirmal, et al.
Publicado: (2026)
Geometric Median (GM) Matching for Robust Data Pruning
por: Acharya, Anish, et al.
Publicado: (2024)
por: Acharya, Anish, et al.
Publicado: (2024)
MatMamba: A Matryoshka State Space Model
por: Shukla, Abhinav, et al.
Publicado: (2024)
por: Shukla, Abhinav, et al.
Publicado: (2024)
Large Language Models Implicitly Learn to See and Hear Just By Reading
por: Verma, Prateek, et al.
Publicado: (2025)
por: Verma, Prateek, et al.
Publicado: (2025)
ELT: Elastic Looped Transformers for Visual Generation
por: Goyal, Sahil, et al.
Publicado: (2026)
por: Goyal, Sahil, et al.
Publicado: (2026)
Matryoshka Diffusion Models
por: Gu, Jiatao, et al.
Publicado: (2023)
por: Gu, Jiatao, et al.
Publicado: (2023)
Towards Signal Processing In Large Language Models
por: Verma, Prateek, et al.
Publicado: (2024)
por: Verma, Prateek, et al.
Publicado: (2024)
ThinkingViT: Matryoshka Thinking Vision Transformer for Elastic Inference
por: Hojjat, Ali, et al.
Publicado: (2025)
por: Hojjat, Ali, et al.
Publicado: (2025)
Federated Model Heterogeneous Matryoshka Representation Learning
por: Yi, Liping, et al.
Publicado: (2024)
por: Yi, Liping, et al.
Publicado: (2024)
Scalable In-context Ranking with Generative Models
por: Gupta, Nilesh, et al.
Publicado: (2025)
por: Gupta, Nilesh, et al.
Publicado: (2025)
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models
por: An, Sohyun, et al.
Publicado: (2025)
por: An, Sohyun, et al.
Publicado: (2025)
On Discrete Prompt Optimization for Diffusion Models
por: Wang, Ruochen, et al.
Publicado: (2024)
por: Wang, Ruochen, et al.
Publicado: (2024)
Adaptive Large Language Models By Layerwise Attention Shortcuts
por: Verma, Prateek, et al.
Publicado: (2024)
por: Verma, Prateek, et al.
Publicado: (2024)
Towards Quantifying the Preconditioning Effect of Adam
por: Das, Rudrajit, et al.
Publicado: (2024)
por: Das, Rudrajit, et al.
Publicado: (2024)
Matryoshka Concept Bottleneck Models
por: Chen, Ziye, et al.
Publicado: (2026)
por: Chen, Ziye, et al.
Publicado: (2026)
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns?
por: Kao, Kuei-Chun, et al.
Publicado: (2024)
por: Kao, Kuei-Chun, et al.
Publicado: (2024)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
por: Bai, Andrew, et al.
Publicado: (2025)
por: Bai, Andrew, et al.
Publicado: (2025)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
por: He, Jack, et al.
Publicado: (2024)
por: He, Jack, et al.
Publicado: (2024)
Data Attribution for Diffusion Models: Timestep-induced Bias in Influence Estimation
por: Xie, Tong, et al.
Publicado: (2024)
por: Xie, Tong, et al.
Publicado: (2024)
Training Matryoshka Mixture-of-Experts for Elastic Inference-Time Expert Utilization
por: Wang, Yaoxiang, et al.
Publicado: (2025)
por: Wang, Yaoxiang, et al.
Publicado: (2025)
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
por: Verma, Prateek
Publicado: (2023)
por: Verma, Prateek
Publicado: (2023)
OR-Bench: An Over-Refusal Benchmark for Large Language Models
por: Cui, Justin, et al.
Publicado: (2024)
por: Cui, Justin, et al.
Publicado: (2024)
Ejemplares similares
-
LLM-guided Hierarchical Search for End-to-end Reasoning Intensive Retrieval
por: Gupta, Nilesh, et al.
Publicado: (2025) -
Matryoshka Quantization
por: Nair, Pranav, et al.
Publicado: (2025) -
Large Language Models are Interpretable Learners
por: Wang, Ruochen, et al.
Publicado: (2024) -
EHI: End-to-end Learning of Hierarchical Index for Efficient Dense Retrieval
por: Kumar, Ramnath, et al.
Publicado: (2023) -
FastLane: Efficient Routed Systems for Late-Interaction Retrieval
por: Kumar, Ramnath, et al.
Publicado: (2026)