DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Sharma, Aman, Najafi, Saeed, Farinneya, Parsa, Jamialahmadi, Benyamin, Tahaei, Marzieh S., Fan, Yuhe, Rezagholizadeh, Mehdi, Chen, Boxing, Jafari, Aref |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
by: Jafari, Aref, et al.
Published: (2025)
by: Jafari, Aref, et al.
Published: (2025)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023)
by: Kavehzadeh, Parsa, et al.
Published: (2023)
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
by: Valipour, Mojtaba, et al.
Published: (2023)
by: Valipour, Mojtaba, et al.
Published: (2023)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
by: Kavehzadeh, Parsa, et al.
Published: (2024)
by: Kavehzadeh, Parsa, et al.
Published: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
by: Fashi, Parsa Ashrafi, et al.
Published: (2026)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026)
by: Haridas, Akash, et al.
Published: (2026)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024)
by: Huang, Chenyang, et al.
Published: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
InvertiTune: High-Quality Data Synthesis for Cost-Effective Single-Shot Text-to-Knowledge Graph Generation
by: Faez, Faezeh, et al.
Published: (2025)
by: Faez, Faezeh, et al.
Published: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
by: Guimarães, Heitor R., et al.
Published: (2024)
by: Guimarães, Heitor R., et al.
Published: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
Metamaterial Bi-stable Vibration Absorbers for Railway Tracks: Experimental Study of Flexural Wave Control
by: Jafari, Ali, et al.
Published: (2025)
by: Jafari, Ali, et al.
Published: (2025)
An Adaptive Intelligent Thermal-Aware Routing Protocol for Wireless Body Area Networks
by: Rahimi, Abdollah, et al.
Published: (2025)
by: Rahimi, Abdollah, et al.
Published: (2025)
A Comparative Study on the Bonding and Electronic Characteristics of DPPH Radical, Its Oxidized, Reduced, and Neutralized Forms Using DFT + Disp Calculations
by: Fatemeh Jafari Khorshidi, et al.
Published: (2025)
by: Fatemeh Jafari Khorshidi, et al.
Published: (2025)
Communicative Universals and the Limits of Language in the Search for Extraterrestrial Intelligence
by: Jafari, Saeed
Published: (2026)
by: Jafari, Saeed
Published: (2026)
Priority‐Based QoS Aware Routing Protocol for Wireless Body Area Networks
by: Abdollah Rahimi, et al.
Published: (2025)
by: Abdollah Rahimi, et al.
Published: (2025)
Adaptive-lambda Subtracted Importance Sampled Scores in Machine Unlearning for DDPMs and VAEs
by: Dini, MohammadParsa, et al.
Published: (2025)
by: Dini, MohammadParsa, et al.
Published: (2025)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
by: Guan, Bryan, et al.
Published: (2025)
by: Guan, Bryan, et al.
Published: (2025)
Towards Practical Tool Usage for Continually Learning LLMs
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Effective Context in Transformers: An Analysis of Fragmentation and Tokenization
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
by: Fesharaki, Amirmehdi Jafari, et al.
Published: (2026)
RIFF: Learning to Rephrase Inputs for Few-shot Fine-tuning of Language Models
by: Najafi, Saeed, et al.
Published: (2024)
by: Najafi, Saeed, et al.
Published: (2024)
Offline Preference Optimization via Maximum Marginal Likelihood Estimation
by: Najafi, Saeed, et al.
Published: (2025)
by: Najafi, Saeed, et al.
Published: (2025)
Towards Secure and Efficient Data Aggregation in Blockchain‐Driven IoT Environments: A Comprehensive and Systematic Study
by: Xujun Tong, et al.
Published: (2025)
by: Xujun Tong, et al.
Published: (2025)
The Filament Rift: $Λ$CDM's Structural Challenge Against Observation
by: Tavasoli, Saeed, et al.
Published: (2025)
by: Tavasoli, Saeed, et al.
Published: (2025)
GrAviPaSt's Lens to the Past: Unveiling the Evolution of Filamentary Structures
by: Ghafour, Parsa, et al.
Published: (2025)
by: Ghafour, Parsa, et al.
Published: (2025)
Cosmic Environment as the Primary Driver of Dwarf Satellite Statistics
by: Tavasoli, Saeed, et al.
Published: (2026)
by: Tavasoli, Saeed, et al.
Published: (2026)
Exploring Next Token Prediction For Optimizing Databases
by: Rayhan, Yeasir, et al.
Published: (2025)
by: Rayhan, Yeasir, et al.
Published: (2025)
RadarSeq: A Temporal Vision Framework for User Churn Prediction via Radar Chart Sequences
by: Najafi, Sina, et al.
Published: (2025)
by: Najafi, Sina, et al.
Published: (2025)
Accuracy and componentwise accuracy in multilinear PageRank
by: Kalyani, Mehdi Najafi, et al.
Published: (2025)
by: Kalyani, Mehdi Najafi, et al.
Published: (2025)
The Impact of Lipoprotein Apheresis on Changes in Biochemical Enzymes: A Systematic Review and Meta‐Analysis
by: Alireza Hatami, et al.
Published: (2026)
by: Alireza Hatami, et al.
Published: (2026)
Similar Items
-
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
by: Jafari, Aref, et al.
Published: (2025) -
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
by: Jamialahmadi, Benyamin, et al.
Published: (2025) -
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023) -
SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
by: Valipour, Mojtaba, et al.
Published: (2023) -
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)