SortedNet: A Scalable and Generalized Framework for Training Modular Deep Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Valipour, Mojtaba, Rezagholizadeh, Mehdi, Rajabzadeh, Hossein, Kavehzadeh, Parsa, Tahaei, Marzieh, Chen, Boxing, Ghodsi, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023)
by: Kavehzadeh, Parsa, et al.
Published: (2023)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
by: Jamialahmadi, Benyamin, et al.
Published: (2025)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
by: Kavehzadeh, Parsa, et al.
Published: (2024)
by: Kavehzadeh, Parsa, et al.
Published: (2024)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
by: Sharma, Aman, et al.
Published: (2025)
by: Sharma, Aman, et al.
Published: (2025)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
SCOUT: Toward Sub-Quadratic Attention via Segment Compression for Optimized Utility in Transformers
by: Jafari, Aref, et al.
Published: (2025)
by: Jafari, Aref, et al.
Published: (2025)
Batch-Max: Higher LLM Throughput using Larger Batch Sizes and KV Cache Compression
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
On the importance of Data Scale in Pretraining Arabic Language Models
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
ReGLA: Refining Gated Linear Attention
by: Lu, Peng, et al.
Published: (2025)
by: Lu, Peng, et al.
Published: (2025)
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Draft on the Fly: Adaptive Self-Speculative Decoding using Cosine Similarity
by: Metel, Michael R., et al.
Published: (2024)
by: Metel, Michael R., et al.
Published: (2024)
Overcoming stretching and shortening assumptions in Euler-Bernoulli theory using nonlinear Hencky beam models: applicable to partly-shortened and partly-stretched beams
by: Rezaei, Mohammad Parsa, et al.
Published: (2024)
by: Rezaei, Mohammad Parsa, et al.
Published: (2024)
OTTAWA: Optimal TransporT Adaptive Word Aligner for Hallucination and Omission Translation Errors Detection
by: Huang, Chenyang, et al.
Published: (2024)
by: Huang, Chenyang, et al.
Published: (2024)
CHARP: Conversation History AwaReness Probing for Knowledge-grounded Dialogue Systems
by: Ghaddar, Abbas, et al.
Published: (2024)
by: Ghaddar, Abbas, et al.
Published: (2024)
How Many Heads Make an SSM? A Unified Framework for Attention and State Space Models
by: Ghodsi, Ali
Published: (2025)
by: Ghodsi, Ali
Published: (2025)
Symbolic-Diffusion: Deep Learning Based Symbolic Regression with D3PM Discrete Token Diffusion
by: Tymkow, Ryan T., et al.
Published: (2025)
by: Tymkow, Ryan T., et al.
Published: (2025)
ECHO-LLaMA: Efficient Caching for High-Performance LLaMA Training
by: Dialameh, Maryam, et al.
Published: (2025)
by: Dialameh, Maryam, et al.
Published: (2025)
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
by: Haridas, Akash, et al.
Published: (2026)
by: Haridas, Akash, et al.
Published: (2026)
Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research
by: Sanaei, Ali, et al.
Published: (2025)
by: Sanaei, Ali, et al.
Published: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
by: Guimarães, Heitor R., et al.
Published: (2024)
by: Guimarães, Heitor R., et al.
Published: (2024)
CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search
by: Mo, Fengran, et al.
Published: (2024)
by: Mo, Fengran, et al.
Published: (2024)
GraphPI: Efficient Protein Inference with Graph Neural Networks
by: Ma, Zheng, et al.
Published: (2026)
by: Ma, Zheng, et al.
Published: (2026)
Learning Chemotherapy Drug Action via Universal Physics-Informed Neural Networks
by: Podina, Lena, et al.
Published: (2024)
by: Podina, Lena, et al.
Published: (2024)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
by: Wang, Xindi, et al.
Published: (2024)
by: Wang, Xindi, et al.
Published: (2024)
InvertiTune: High-Quality Data Synthesis for Cost-Effective Single-Shot Text-to-Knowledge Graph Generation
by: Faez, Faezeh, et al.
Published: (2025)
by: Faez, Faezeh, et al.
Published: (2025)
Self-similar group actions on ultragraphs and associated $C^*$-algebras
by: Larki, Hossein, et al.
Published: (2025)
by: Larki, Hossein, et al.
Published: (2025)
Minimality and effectiveness of the groupoid associated to a self-similar ultragraph
by: Larki, Hossein, et al.
Published: (2025)
by: Larki, Hossein, et al.
Published: (2025)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
by: Li, Guihong, et al.
Published: (2025)
by: Li, Guihong, et al.
Published: (2025)
The effect of fabrication procedures and thermomechanical loading on the structural properties of screw‐retained metal‐ceramic implant restorations: An in vitro study
by: Hosein Mohebbi, et al.
Published: (2024)
by: Hosein Mohebbi, et al.
Published: (2024)
Scalable Graph Self-Supervised Learning
by: Pasand, Ali Saheb, et al.
Published: (2024)
by: Pasand, Ali Saheb, et al.
Published: (2024)
Cell divisions suppress dynamical correlations in solid tissues
by: Tahaei, Ali, et al.
Published: (2026)
by: Tahaei, Ali, et al.
Published: (2026)
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation
by: Thakur, Nandan, et al.
Published: (2023)
by: Thakur, Nandan, et al.
Published: (2023)
Perception-Informed Neural Networks: Beyond Physics-Informed Neural Networks
by: Mazandarani, Mehran, et al.
Published: (2025)
by: Mazandarani, Mehran, et al.
Published: (2025)
Diagnosing epilepsy using entropy measures and embedding parameters of EEG signals
by: Valipour, Fatemeh, et al.
Published: (2024)
by: Valipour, Fatemeh, et al.
Published: (2024)
Path Analysis for Effective Fault Localization in Deep Neural Networks
by: Hashemifar, Soroush, et al.
Published: (2023)
by: Hashemifar, Soroush, et al.
Published: (2023)
Resonance RoPE: Improving Context Length Generalization of Large Language Models
by: Wang, Suyuchen, et al.
Published: (2024)
by: Wang, Suyuchen, et al.
Published: (2024)
Auto-Regressive Masked Diffusion Models
by: Karami, Mahdi, et al.
Published: (2026)
by: Karami, Mahdi, et al.
Published: (2026)
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2024)
by: Karami, Mahdi, et al.
Published: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
by: Huang, Jerry, et al.
Published: (2024)
by: Huang, Jerry, et al.
Published: (2024)
Similar Items
-
Sorted LLaMA: Unlocking the Potential of Intermediate Layers of Large Language Models for Dynamic Inference
by: Kavehzadeh, Parsa, et al.
Published: (2023) -
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
by: Jamialahmadi, Benyamin, et al.
Published: (2025) -
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
by: Rajabzadeh, Hossein, et al.
Published: (2024) -
S2D: Sorted Speculative Decoding For More Efficient Deployment of Nested Large Language Models
by: Kavehzadeh, Parsa, et al.
Published: (2024) -
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
by: Sharma, Aman, et al.
Published: (2025)