Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
Fuente:
arXiv
Salvato in:
| Autori principali: | Fashi, Parsa Ashrafi, Saxena, Utkarsh, Rezagholizadeh, Mehdi, Jafari, Aref, Haridas, Akash, Yang, Mingyu, Bhatia, Vansh, Li, Guihong, Appia, Vikram, Barsoum, Emad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
di: Haridas, Akash, et al.
Pubblicazione: (2026)
di: Haridas, Akash, et al.
Pubblicazione: (2026)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
di: Li, Guihong, et al.
Pubblicazione: (2025)
di: Li, Guihong, et al.
Pubblicazione: (2025)
Zebra-Llama: Towards Extremely Efficient Hybrid Models
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
di: Yang, Mingyu, et al.
Pubblicazione: (2025)
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
di: Dukler, Yonatan, et al.
Pubblicazione: (2025)
di: Dukler, Yonatan, et al.
Pubblicazione: (2025)
Verifier Threshold: An Efficient Test-Time Scaling Approach for Image Generation
di: Sundaresha, Vignesh, et al.
Pubblicazione: (2025)
di: Sundaresha, Vignesh, et al.
Pubblicazione: (2025)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
di: Jamialahmadi, Benyamin, et al.
Pubblicazione: (2025)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
di: Younesian, Sharareh, et al.
Pubblicazione: (2026)
di: Younesian, Sharareh, et al.
Pubblicazione: (2026)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
di: Sharma, Aman, et al.
Pubblicazione: (2025)
di: Sharma, Aman, et al.
Pubblicazione: (2025)
CASCADE: Context-Aware Relaxation for Speculative Image Decoding
di: Yildirim, Selin, et al.
Pubblicazione: (2026)
di: Yildirim, Selin, et al.
Pubblicazione: (2026)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
di: Wang, Xindi, et al.
Pubblicazione: (2024)
di: Wang, Xindi, et al.
Pubblicazione: (2024)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
di: Rajabzadeh, Hossein, et al.
Pubblicazione: (2024)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
di: Juvekar, Kush, et al.
Pubblicazione: (2025)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2023)
High-Fidelity Human Avatars from Laptop Webcams using Edge Compute
di: Haridas, Akash, et al.
Pubblicazione: (2025)
di: Haridas, Akash, et al.
Pubblicazione: (2025)
Metamaterial Bi-stable Vibration Absorbers for Railway Tracks: Experimental Study of Flexural Wave Control
di: Jafari, Ali, et al.
Pubblicazione: (2025)
di: Jafari, Ali, et al.
Pubblicazione: (2025)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
di: Wang, Suyuchen, et al.
Pubblicazione: (2024)
Self-Taught Agentic Long Context Understanding
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
di: Zhuang, Yufan, et al.
Pubblicazione: (2025)
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
di: Jafari, Aref, et al.
Pubblicazione: (2025)
di: Jafari, Aref, et al.
Pubblicazione: (2025)
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
di: Valipour, Mojtaba, et al.
Pubblicazione: (2023)
di: Valipour, Mojtaba, et al.
Pubblicazione: (2023)
DL-QAT: Weight-Decomposed Low-Rank Quantization-Aware Training for Large Language Models
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
di: Ke, Wenjin, et al.
Pubblicazione: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
di: Metel, Michael R., et al.
Pubblicazione: (2024)
di: Metel, Michael R., et al.
Pubblicazione: (2024)
KVLinC : KV Cache Quantization with Hadamard Rotation and Linear Correction
di: Saxena, Utkarsh, et al.
Pubblicazione: (2025)
di: Saxena, Utkarsh, et al.
Pubblicazione: (2025)
CourtNav: Voice-Guided, Anchor-Accurate Navigation of Long Legal Documents in Courtrooms
di: Khadloya, Sai, et al.
Pubblicazione: (2025)
di: Khadloya, Sai, et al.
Pubblicazione: (2025)
HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
di: Negi, Shubham, et al.
Pubblicazione: (2024)
di: Negi, Shubham, et al.
Pubblicazione: (2024)
Robust Steady-State-Aware Model Predictive Control for Systems with Limited Computational Resources and External Disturbances
di: Ozoumchelooei, Hassan Jafari, et al.
Pubblicazione: (2025)
di: Ozoumchelooei, Hassan Jafari, et al.
Pubblicazione: (2025)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
di: Kavehzadeh, Parsa, et al.
Pubblicazione: (2024)
Adaptive-lambda Subtracted Importance Sampled Scores in Machine Unlearning for DDPMs and VAEs
di: Dini, MohammadParsa, et al.
Pubblicazione: (2025)
di: Dini, MohammadParsa, et al.
Pubblicazione: (2025)
On the importance of Data Scale in Pretraining Arabic Language Models
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
di: Ghaddar, Abbas, et al.
Pubblicazione: (2024)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
di: Guan, Bryan, et al.
Pubblicazione: (2025)
di: Guan, Bryan, et al.
Pubblicazione: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Designer Helical Fibers and Tubes: Self‐assembling Hybrid Peptides via Leu/Ile‐Phe Zipper
di: Souvik Dutta, et al.
Pubblicazione: (2024)
di: Souvik Dutta, et al.
Pubblicazione: (2024)
Artificial Neural Networks and Guided Gene Expression Programming to Predict Wall Pressure Spectra Beneath Turbulent Boundary Layers
di: Kurhade, Nachiketa Narayan, et al.
Pubblicazione: (2023)
di: Kurhade, Nachiketa Narayan, et al.
Pubblicazione: (2023)
Hybrid Cyclopeptide‐Based Chemical Models for Allosteric Disulfide and Amyloid Self‐Assembly
di: Surya Kant Bhardwaj, et al.
Pubblicazione: (2025)
di: Surya Kant Bhardwaj, et al.
Pubblicazione: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
di: An, Zihao, et al.
Pubblicazione: (2025)
di: An, Zihao, et al.
Pubblicazione: (2025)
AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selection
di: Ray, Pretam, et al.
Pubblicazione: (2026)
di: Ray, Pretam, et al.
Pubblicazione: (2026)
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering
di: Joshi, Vinay, et al.
Pubblicazione: (2025)
di: Joshi, Vinay, et al.
Pubblicazione: (2025)
Expert Upcycling: Shifting the Compute-Efficient Frontier of Mixture-of-Experts
di: Dwivedi, Chaitanya, et al.
Pubblicazione: (2026)
di: Dwivedi, Chaitanya, et al.
Pubblicazione: (2026)
An Adaptive Intelligent Thermal-Aware Routing Protocol for Wireless Body Area Networks
di: Rahimi, Abdollah, et al.
Pubblicazione: (2025)
di: Rahimi, Abdollah, et al.
Pubblicazione: (2025)
Mock Theta Functions as Optimal Stopping Criteria for Photonic Quantum Entropy Computation
di: Ansh Sharma, Ansh, et al.
Pubblicazione: (2026)
di: Ansh Sharma, Ansh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
di: Haridas, Akash, et al.
Pubblicazione: (2026) -
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
di: Li, Guihong, et al.
Pubblicazione: (2025) -
Zebra-Llama: Towards Extremely Efficient Hybrid Models
di: Yang, Mingyu, et al.
Pubblicazione: (2025) -
FarSkip-Collective: Unhobbling Blocking Communication in Mixture of Experts Models
di: Dukler, Yonatan, et al.
Pubblicazione: (2025) -
Verifier Threshold: An Efficient Test-Time Scaling Approach for Image Generation
di: Sundaresha, Vignesh, et al.
Pubblicazione: (2025)