Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Ruihan, Gao, Yuting, Wang, Lan, Li, Jianing, Chen, Weihao, Guo, Qingpei, Yang, Ming, Zhang, Shiliang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
por: Gao, Yuting, et al.
Publicado: (2025)
por: Gao, Yuting, et al.
Publicado: (2025)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
por: Xu, Ruihan, et al.
Publicado: (2026)
por: Xu, Ruihan, et al.
Publicado: (2026)
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
por: Gao, Yuting, et al.
Publicado: (2025)
por: Gao, Yuting, et al.
Publicado: (2025)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
por: Xuan, Shiyu, et al.
Publicado: (2023)
por: Xuan, Shiyu, et al.
Publicado: (2023)
LoopQ: Quantization for Recursive Transformers
por: Fang, Rui, et al.
Publicado: (2026)
por: Fang, Rui, et al.
Publicado: (2026)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
por: Ma, Ziping, et al.
Publicado: (2024)
por: Ma, Ziping, et al.
Publicado: (2024)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
por: Ma, Zehong, et al.
Publicado: (2026)
por: Ma, Zehong, et al.
Publicado: (2026)
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models
por: Cameron, Chris, et al.
Publicado: (2026)
por: Cameron, Chris, et al.
Publicado: (2026)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
por: Guo, Qingpei, et al.
Publicado: (2024)
por: Guo, Qingpei, et al.
Publicado: (2024)
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
por: Yang, Xiao-Wen, et al.
Publicado: (2025)
por: Yang, Xiao-Wen, et al.
Publicado: (2025)
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
por: Jing, Linglin, et al.
Publicado: (2025)
por: Jing, Linglin, et al.
Publicado: (2025)
TabKANet: Tabular Data Modeling with Kolmogorov-Arnold Network and Transformer
por: Gao, Weihao, et al.
Publicado: (2024)
por: Gao, Weihao, et al.
Publicado: (2024)
NN-Former: Rethinking Graph Structure in Neural Architecture Representation
por: Xu, Ruihan, et al.
Publicado: (2025)
por: Xu, Ruihan, et al.
Publicado: (2025)
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
por: Sun, Qixin, et al.
Publicado: (2025)
por: Sun, Qixin, et al.
Publicado: (2025)
Moving Forward, Looking Back
por: Hagener, Malte
Publicado: (2010)
por: Hagener, Malte
Publicado: (2010)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
por: Xu, Peiran, et al.
Publicado: (2025)
por: Xu, Peiran, et al.
Publicado: (2025)
What Makes Looped Transformers Perform Better Than Non-Recursive Ones
por: Gong, Zixuan, et al.
Publicado: (2025)
por: Gong, Zixuan, et al.
Publicado: (2025)
IG-MCTS: Human-in-the-Loop Cooperative Navigation under Incomplete Information
por: Chen, Shenghui, et al.
Publicado: (2025)
por: Chen, Shenghui, et al.
Publicado: (2025)
Computing: Looking Back and Moving Forward
por: Golec, Muhammed, et al.
Publicado: (2024)
por: Golec, Muhammed, et al.
Publicado: (2024)
LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models
por: Tian, Shi-Yu, et al.
Publicado: (2026)
por: Tian, Shi-Yu, et al.
Publicado: (2026)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
por: Sheng, Zhengyan, et al.
Publicado: (2025)
por: Sheng, Zhengyan, et al.
Publicado: (2025)
BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
por: Liu, Shu, et al.
Publicado: (2025)
por: Liu, Shu, et al.
Publicado: (2025)
Social Debiasing for Fair Multi-modal LLMs
por: Cheng, Harry, et al.
Publicado: (2024)
por: Cheng, Harry, et al.
Publicado: (2024)
RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
por: Yu, Chao, et al.
Publicado: (2025)
por: Yu, Chao, et al.
Publicado: (2025)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
por: Zheng, Ruobing, et al.
Publicado: (2026)
por: Zheng, Ruobing, et al.
Publicado: (2026)
Benchmarking Open-Source Large Language Models on Healthcare Text Classification Tasks
por: Guo, Yuting, et al.
Publicado: (2025)
por: Guo, Yuting, et al.
Publicado: (2025)
RecurrentGemma: Moving Past Transformers for Efficient Open Language Models
por: Botev, Aleksandar, et al.
Publicado: (2024)
por: Botev, Aleksandar, et al.
Publicado: (2024)
You Only Forward Once: An Efficient Compositional Judging Paradigm
por: Zhang, Tianlong, et al.
Publicado: (2025)
por: Zhang, Tianlong, et al.
Publicado: (2025)
ReCAP: Recursive Context-Aware Reasoning and Planning for Large Language Model Agents
por: Zhang, Zhenyu, et al.
Publicado: (2025)
por: Zhang, Zhenyu, et al.
Publicado: (2025)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
por: Vendrell, Victor Conchello, et al.
Publicado: (2026)
por: Vendrell, Victor Conchello, et al.
Publicado: (2026)
Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy Refinement
por: Zhang, Yunjian, et al.
Publicado: (2026)
por: Zhang, Yunjian, et al.
Publicado: (2026)
XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models
por: Dong, Yixin, et al.
Publicado: (2024)
por: Dong, Yixin, et al.
Publicado: (2024)
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
por: Fang, Chengyu, et al.
Publicado: (2026)
por: Fang, Chengyu, et al.
Publicado: (2026)
LOOPRAG: Enhancing Loop Transformation Optimization with Retrieval-Augmented Large Language Models
por: Zhi, Yijie, et al.
Publicado: (2025)
por: Zhi, Yijie, et al.
Publicado: (2025)
Large Multimodal Model Compression via Efficient Pruning and Distillation at AntGroup
por: Wang, Maolin, et al.
Publicado: (2023)
por: Wang, Maolin, et al.
Publicado: (2023)
Efficient Prompt Tuning by Multi-Space Projection and Prompt Fusion
por: Lan, Pengxiang, et al.
Publicado: (2024)
por: Lan, Pengxiang, et al.
Publicado: (2024)
Agent Skills for Large Language Models: Architecture, Acquisition, Security, and the Path Forward
por: Xu, Renjun, et al.
Publicado: (2026)
por: Xu, Renjun, et al.
Publicado: (2026)
Automatic Instruction Evolving for Large Language Models
por: Zeng, Weihao, et al.
Publicado: (2024)
por: Zeng, Weihao, et al.
Publicado: (2024)
Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agents
por: Shi, Yaorui, et al.
Publicado: (2025)
por: Shi, Yaorui, et al.
Publicado: (2025)
MeSH: Memory-as-State-Highways for Recursive Transformers
por: Yu, Chengting, et al.
Publicado: (2025)
por: Yu, Chengting, et al.
Publicado: (2025)
Ejemplares similares
-
OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
por: Gao, Yuting, et al.
Publicado: (2025) -
FlattenGPT: Depth Compression for Transformer with Layer Flattening
por: Xu, Ruihan, et al.
Publicado: (2026) -
AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
por: Gao, Yuting, et al.
Publicado: (2025) -
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
por: Xuan, Shiyu, et al.
Publicado: (2023) -
LoopQ: Quantization for Recursive Transformers
por: Fang, Rui, et al.
Publicado: (2026)