Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
Fuente:
arXiv
Saved in:
| Main Authors: | Bae, Sangmin, Kim, Yujin, Bayat, Reza, Kim, Sungnyun, Ha, Jiyoun, Schuster, Tal, Fisch, Adam, Harutyunyan, Hrayr, Ji, Ziwei, Courville, Aaron, Yun, Se-Young |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024)
by: Bae, Sangmin, et al.
Published: (2024)
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
by: Ho, Namgyu, et al.
Published: (2024)
by: Ho, Namgyu, et al.
Published: (2024)
Diffusion-based Episodes Augmentation for Offline Multi-Agent Reinforcement Learning
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation
by: Cho, Sungwoo, et al.
Published: (2025)
by: Cho, Sungwoo, et al.
Published: (2025)
Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
by: Bae, Sangmin, et al.
Published: (2023)
by: Bae, Sangmin, et al.
Published: (2023)
VACoDe: Visual Augmented Contrastive Decoding
by: Kim, Sihyeon, et al.
Published: (2024)
by: Kim, Sihyeon, et al.
Published: (2024)
FedDr+: Stabilizing Dot-regression with Global Feature Distillation for Federated Learning
by: Kim, Seongyoon, et al.
Published: (2024)
by: Kim, Seongyoon, et al.
Published: (2024)
Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models
by: Park, Youngrok, et al.
Published: (2025)
by: Park, Youngrok, et al.
Published: (2025)
Scalable Frameworks for Real-World Audio-Visual Speech Recognition
by: Kim, Sungnyun
Published: (2025)
by: Kim, Sungnyun
Published: (2025)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Fine-tuned In-Context Learning Transformers are Excellent Tabular Data Classifiers
by: Breejen, Felix den, et al.
Published: (2024)
by: Breejen, Felix den, et al.
Published: (2024)
Carpe Diem: On the Evaluation of World Knowledge in Lifelong Language Models
by: Kim, Yujin, et al.
Published: (2023)
by: Kim, Yujin, et al.
Published: (2023)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
by: Oh, Jihwan, et al.
Published: (2026)
by: Oh, Jihwan, et al.
Published: (2026)
Guiding Reasoning in Small Language Models with LLM Assistance
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
by: Bae, Jongseong, et al.
Published: (2024)
by: Bae, Jongseong, et al.
Published: (2024)
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
In-context Learning in Presence of Spurious Correlations
by: Harutyunyan, Hrayr, et al.
Published: (2024)
by: Harutyunyan, Hrayr, et al.
Published: (2024)
RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification
by: Kim, June-Woo, et al.
Published: (2024)
by: Kim, June-Woo, et al.
Published: (2024)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Generative Recursive Reasoning
by: Baek, Junyeob, et al.
Published: (2026)
by: Baek, Junyeob, et al.
Published: (2026)
Scaffolding Recursive Divergence and Convergence in Story Ideation
by: Kim, Taewook, et al.
Published: (2025)
by: Kim, Taewook, et al.
Published: (2025)
Improving Recursive Transformers with Mixture of LoRAs
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
by: Nouriborji, Mohammadmahdi, et al.
Published: (2025)
Three Cars Approaching within 100m! Enhancing Distant Geometry by Tri-Axis Voxel Scanning for Camera-based Semantic Scene Completion
by: Bae, Jongseong, et al.
Published: (2024)
by: Bae, Jongseong, et al.
Published: (2024)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
by: Jang, Kangwook, et al.
Published: (2023)
by: Jang, Kangwook, et al.
Published: (2023)
Recursive-algebraic solution of the closed string tachyon vacuum equation
by: Kim, Manki
Published: (2026)
by: Kim, Manki
Published: (2026)
Constrained Recursive Logit for Route Choice Analysis
by: Tran, Hung, et al.
Published: (2025)
by: Tran, Hung, et al.
Published: (2025)
Scalable Learning of Intrusion Responses through Recursive Decomposition
by: Hammar, Kim, et al.
Published: (2023)
by: Hammar, Kim, et al.
Published: (2023)
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
RVLM: Recursive Vision-Language Models with Adaptive Depth
by: Mayumu, Nicanor, et al.
Published: (2026)
by: Mayumu, Nicanor, et al.
Published: (2026)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
The Secretion of Inflammatory Cytokines Triggered by TLR2 Through Calcium‐Dependent and Calcium‐Independent Pathways in Keratinocytes
by: Eun-Ok Kim, et al.
Published: (2024)
by: Eun-Ok Kim, et al.
Published: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
Scattered Mixture-of-Experts Implementation
by: Tan, Shawn, et al.
Published: (2024)
by: Tan, Shawn, et al.
Published: (2024)
Mimetic Initialization Helps State Space Models Learn to Recall
by: Trockman, Asher, et al.
Published: (2024)
by: Trockman, Asher, et al.
Published: (2024)
Similar Items
-
Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
by: Bae, Sangmin, et al.
Published: (2024) -
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2025) -
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024) -
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
by: Kim, Sungnyun, et al.
Published: (2025) -
Block Transformer: Global-to-Local Language Modeling for Fast Inference
by: Ho, Namgyu, et al.
Published: (2024)