Improving Recursive Transformers with Mixture of LoRAs
Fuente:
arXiv
Saved in:
| Main Authors: | Nouriborji, Mohammadmahdi, Rohanian, Morteza, Rohanian, Omid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
by: Rohanian, Omid, et al.
Published: (2023)
by: Rohanian, Omid, et al.
Published: (2023)
Lightweight Transformers for Clinical Natural Language Processing
by: Rohanian, Omid, et al.
Published: (2023)
by: Rohanian, Omid, et al.
Published: (2023)
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
Efficient Interleaved Speech Modeling through Knowledge Distillation
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
Exploring Efficient Learning of Small BERT Networks with LoRA and DoRA
by: Frees, Daniel, et al.
Published: (2025)
by: Frees, Daniel, et al.
Published: (2025)
QuAILoRA: Quantization-Aware Initialization for LoRA
by: Lawton, Neal, et al.
Published: (2024)
by: Lawton, Neal, et al.
Published: (2024)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
Rapid Biomedical Research Classification: The Pandemic PACT Advanced Categorisation Engine
by: Rohanian, Omid, et al.
Published: (2024)
by: Rohanian, Omid, et al.
Published: (2024)
Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
by: Dahlem, Dominik, et al.
Published: (2026)
by: Dahlem, Dominik, et al.
Published: (2026)
Online AUC Optimization Based on Second-order Surrogate Loss
by: Luo, JunRu, et al.
Published: (2025)
by: Luo, JunRu, et al.
Published: (2025)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Evolutionary Computation as Natural Generative AI
by: Shi, Yaxin, et al.
Published: (2025)
by: Shi, Yaxin, et al.
Published: (2025)
Exact Sequence Interpolation with Transformers
by: Alcalde, Albert, et al.
Published: (2025)
by: Alcalde, Albert, et al.
Published: (2025)
The Meta-Prompting Protocol: Orchestrating LLMs via Adversarial Feedback Loops
by: Fu, Fanzhe
Published: (2025)
by: Fu, Fanzhe
Published: (2025)
Extracting Sentence Embeddings from Pretrained Transformer Models
by: Stankevičius, Lukas, et al.
Published: (2024)
by: Stankevičius, Lukas, et al.
Published: (2024)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
by: Imanov, Olaf Yunus Laitinen
Published: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
by: Grigaliūnas, Domas, et al.
Published: (2024)
by: Grigaliūnas, Domas, et al.
Published: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
Clinical Data Goes MEDS? Let's OWL make sense of it
by: Marfoglia, Alberto, et al.
Published: (2026)
by: Marfoglia, Alberto, et al.
Published: (2026)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
Surfing the modeling of PoS taggers in low-resource scenarios
by: Ferro, Manuel Vilares, et al.
Published: (2024)
by: Ferro, Manuel Vilares, et al.
Published: (2024)
Robust Hybrid Classical-Quantum Transfer Learning Model for Text Classification Using GPT-Neo 125M with LoRA & SMOTE Enhancement
by: Wishal, Santanam
Published: (2025)
by: Wishal, Santanam
Published: (2025)
A Language Model-Driven Semi-Supervised Ensemble Framework for Illicit Market Detection Across Deep/Dark Web and Social Platforms
by: Yazdanjue, Navid, et al.
Published: (2025)
by: Yazdanjue, Navid, et al.
Published: (2025)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
by: Wang, Zhixiang
Published: (2025)
by: Wang, Zhixiang
Published: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
by: Yudin, Nikolay
Published: (2026)
by: Yudin, Nikolay
Published: (2026)
See-Saw Generative Mechanism for Scalable Recursive Code Generation with Generative AI
by: Vsevolodovna, Ruslan Idelfonso Magaña
Published: (2024)
by: Vsevolodovna, Ruslan Idelfonso Magaña
Published: (2024)
Beyond the Surface: Uncovering Implicit Locations with LLMs for Personalized Local News
by: Katz, Gali, et al.
Published: (2025)
by: Katz, Gali, et al.
Published: (2025)
Unraveling Media Perspectives: A Comprehensive Methodology Combining Large Language Models, Topic Modeling, Sentiment Analysis, and Ontology Learning to Analyse Media Bias
by: Jähde, Orlando, et al.
Published: (2025)
by: Jähde, Orlando, et al.
Published: (2025)
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024)
by: Hostnik, Marko, et al.
Published: (2024)
Uncertainty Estimation and Quantification for LLMs: A Simple Supervised Approach
by: Liu, Linyu, et al.
Published: (2024)
by: Liu, Linyu, et al.
Published: (2024)
Dark LLMs: The Growing Threat of Unaligned AI Models
by: Fire, Michael, et al.
Published: (2025)
by: Fire, Michael, et al.
Published: (2025)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
by: Mutlu, Abdulvahap, et al.
Published: (2026)
by: Mutlu, Abdulvahap, et al.
Published: (2026)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
by: Vileikytė, Brigita, et al.
Published: (2024)
by: Vileikytė, Brigita, et al.
Published: (2024)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
by: Chen, Yihong, et al.
Published: (2022)
by: Chen, Yihong, et al.
Published: (2022)
Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models
by: Wallace, Tom, et al.
Published: (2025)
by: Wallace, Tom, et al.
Published: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
Reduced Jeffries-Matusita distance: A Novel Loss Function to Improve Generalization Performance of Deep Classification Models
by: Lashkari, Mohammad, et al.
Published: (2024)
by: Lashkari, Mohammad, et al.
Published: (2024)
Similar Items
-
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
by: Rohanian, Omid, et al.
Published: (2023) -
Lightweight Transformers for Clinical Natural Language Processing
by: Rohanian, Omid, et al.
Published: (2023) -
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025) -
Efficient Interleaved Speech Modeling through Knowledge Distillation
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025) -
Exploring Efficient Learning of Small BERT Networks with LoRA and DoRA
by: Frees, Daniel, et al.
Published: (2025)