Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Dongjie, Li, Yongqi, Ma, Zhixin, Cai, Hongru, Hu, Yupeng, Wang, Wenjie, Nie, Liqiang, Li, Wenjie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
di: Chen, Chao, et al.
Pubblicazione: (2025)
di: Chen, Chao, et al.
Pubblicazione: (2025)
R$^2$ec: Towards Large Recommender Models with Reasoning
di: You, Runyang, et al.
Pubblicazione: (2025)
di: You, Runyang, et al.
Pubblicazione: (2025)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Parallel Test-Time Scaling for Latent Reasoning Models
di: You, Runyang, et al.
Pubblicazione: (2025)
di: You, Runyang, et al.
Pubblicazione: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
Distillation Enhanced Generative Retrieval
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
di: Xu, Qiancheng, et al.
Pubblicazione: (2026)
di: Xu, Qiancheng, et al.
Pubblicazione: (2026)
One Adapts to Any: Meta Reward Modeling for Personalized LLM Alignment
di: Cai, Hongru, et al.
Pubblicazione: (2026)
di: Cai, Hongru, et al.
Pubblicazione: (2026)
Large Language Models Empowered Personalized Web Agents
di: Cai, Hongru, et al.
Pubblicazione: (2024)
di: Cai, Hongru, et al.
Pubblicazione: (2024)
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
di: Lin, Dongding, et al.
Pubblicazione: (2026)
di: Lin, Dongding, et al.
Pubblicazione: (2026)
OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence
di: Yuan, Long, et al.
Pubblicazione: (2025)
di: Yuan, Long, et al.
Pubblicazione: (2025)
Discriminative Probing and Tuning for Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
di: Zhu, Boyu, et al.
Pubblicazione: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
di: Xia, Heming, et al.
Pubblicazione: (2025)
di: Xia, Heming, et al.
Pubblicazione: (2025)
Agent-as-a-Judge
di: You, Runyang, et al.
Pubblicazione: (2026)
di: You, Runyang, et al.
Pubblicazione: (2026)
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
di: Liu, Kai, et al.
Pubblicazione: (2025)
di: Liu, Kai, et al.
Pubblicazione: (2025)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
di: Wang, Jian, et al.
Pubblicazione: (2025)
di: Wang, Jian, et al.
Pubblicazione: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
di: Xiao, Yicheng, et al.
Pubblicazione: (2025)
Enhancing Tool Retrieval with Iterative Feedback from Large Language Models
di: Xu, Qiancheng, et al.
Pubblicazione: (2024)
di: Xu, Qiancheng, et al.
Pubblicazione: (2024)
UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
di: Lan, Zhibin, et al.
Pubblicazione: (2025)
di: Lan, Zhibin, et al.
Pubblicazione: (2025)
Open Multimodal Retrieval-Augmented Factual Image Generation
di: Tian, Yang, et al.
Pubblicazione: (2025)
di: Tian, Yang, et al.
Pubblicazione: (2025)
PEToolLLM: Towards Personalized Tool Learning in Large Language Models
di: Xu, Qiancheng, et al.
Pubblicazione: (2025)
di: Xu, Qiancheng, et al.
Pubblicazione: (2025)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
di: Li, Yongqi, et al.
Pubblicazione: (2025)
di: Li, Yongqi, et al.
Pubblicazione: (2025)
MaeFuse: Transferring Omni Features with Pretrained Masked Autoencoders for Infrared and Visible Image Fusion via Guided Training
di: Li, Jiayang, et al.
Pubblicazione: (2024)
di: Li, Jiayang, et al.
Pubblicazione: (2024)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
di: AI, Inclusion, et al.
Pubblicazione: (2025)
di: AI, Inclusion, et al.
Pubblicazione: (2025)
CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
di: Tian, Yang, et al.
Pubblicazione: (2025)
di: Tian, Yang, et al.
Pubblicazione: (2025)
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
di: Li, Yongxiang, et al.
Pubblicazione: (2026)
di: Li, Yongxiang, et al.
Pubblicazione: (2026)
Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
di: Li, Jiayang, et al.
Pubblicazione: (2025)
di: Li, Jiayang, et al.
Pubblicazione: (2025)
SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning
di: Qian, Zhe, et al.
Pubblicazione: (2026)
di: Qian, Zhe, et al.
Pubblicazione: (2026)
OmniCam: Unified Multimodal Video Generation via Camera Control
di: Yang, Xiaoda, et al.
Pubblicazione: (2025)
di: Yang, Xiaoda, et al.
Pubblicazione: (2025)
Interleaved-Modal Chain-of-Thought
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
OmniGen: Unified Image Generation
di: Xiao, Shitao, et al.
Pubblicazione: (2024)
di: Xiao, Shitao, et al.
Pubblicazione: (2024)
Personalized Large Language Model Assistant with Evolving Conditional Memory
di: Yuan, Ruifeng, et al.
Pubblicazione: (2023)
di: Yuan, Ruifeng, et al.
Pubblicazione: (2023)
Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning
di: Wu, Wenjie, et al.
Pubblicazione: (2025)
di: Wu, Wenjie, et al.
Pubblicazione: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
di: Wu, Chenyuan, et al.
Pubblicazione: (2025)
di: Wu, Chenyuan, et al.
Pubblicazione: (2025)
TimeOmni-VL: Unified Models for Time Series Understanding and Generation
di: Guan, Tong, et al.
Pubblicazione: (2026)
di: Guan, Tong, et al.
Pubblicazione: (2026)
Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
di: Cheng, Dongjie, et al.
Pubblicazione: (2026) -
Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
di: Chen, Chao, et al.
Pubblicazione: (2025) -
R$^2$ec: Towards Large Recommender Models with Reasoning
di: You, Runyang, et al.
Pubblicazione: (2025) -
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024) -
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
di: Qu, Leigang, et al.
Pubblicazione: (2024)