Toward Native Multimodal Modeling: A Roadmap
Fuente:
arXiv
Salvato in:
| Autori principali: | An, Siyu, Lu, Junru, Dong, Junnan, Wang, Qiufeng, Li, Yinghui, Fei, Weizhi, Yu, Zichao, Yuan, Zheng, Liu, Biao, Wang, Haopeng, Liang, Renzhao, Yang, Yixuan, Shen, Yunhang, Ke, Bo, Chen, Keyu, Luo, Linhao, Zou, Difan, Huang, Xiao, Yin, Di, Qiao, Ruizhi, Sun, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
di: Lu, Junru, et al.
Pubblicazione: (2025)
di: Lu, Junru, et al.
Pubblicazione: (2025)
Deep Tabular Research via Continual Experience-Driven Execution
di: Dong, Junnan, et al.
Pubblicazione: (2026)
di: Dong, Junnan, et al.
Pubblicazione: (2026)
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
di: Lu, Wensheng, et al.
Pubblicazione: (2025)
di: Lu, Wensheng, et al.
Pubblicazione: (2025)
Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
di: Zhang, Yujian, et al.
Pubblicazione: (2025)
di: Zhang, Yujian, et al.
Pubblicazione: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
When should I search more: Adaptive Complex Query Optimization with Reinforcement Learning
di: Wen, Wei, et al.
Pubblicazione: (2026)
di: Wen, Wei, et al.
Pubblicazione: (2026)
ProFlow: Zero-Shot Physics-Consistent Sampling via Proximal Flow Guidance
di: Yu, Zichao, et al.
Pubblicazione: (2026)
di: Yu, Zichao, et al.
Pubblicazione: (2026)
Relative Score Policy Optimization for Diffusion Language Models
di: Yu, Zichao, et al.
Pubblicazione: (2026)
di: Yu, Zichao, et al.
Pubblicazione: (2026)
Unifying Large Language Models and Knowledge Graphs: A Roadmap
di: Pan, Shirui, et al.
Pubblicazione: (2023)
di: Pan, Shirui, et al.
Pubblicazione: (2023)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
di: Wang, Shouju, et al.
Pubblicazione: (2026)
di: Wang, Shouju, et al.
Pubblicazione: (2026)
A cootie catcher/fortune teller to recruit “English Native Teachers” to teach in China
di: Yixuan Wang
Pubblicazione: (2025)
di: Yixuan Wang
Pubblicazione: (2025)
Towards Native AI in 6G Standardization: The Roadmap of Semantic Communication
di: Zhang, Ping, et al.
Pubblicazione: (2025)
di: Zhang, Ping, et al.
Pubblicazione: (2025)
Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex Reasoning
di: Dong, Junnan, et al.
Pubblicazione: (2025)
di: Dong, Junnan, et al.
Pubblicazione: (2025)
Multifunctional Epoxy/ HGM / MXene Composites Integrating Impact Resistance, Electrical Insulation, and Tailorable Terahertz Shielding
di: Ruizhi Wang, et al.
Pubblicazione: (2026)
di: Ruizhi Wang, et al.
Pubblicazione: (2026)
FinVerse: An Autonomous Agent System for Versatile Financial Analysis
di: An, Siyu, et al.
Pubblicazione: (2024)
di: An, Siyu, et al.
Pubblicazione: (2024)
Multimodal Label Relevance Ranking via Reinforcement Learning
di: Guo, Taian, et al.
Pubblicazione: (2024)
di: Guo, Taian, et al.
Pubblicazione: (2024)
STAIR: Manipulating Collaborative and Multimodal Information for E-Commerce Recommendation
di: Xu, Cong, et al.
Pubblicazione: (2024)
di: Xu, Cong, et al.
Pubblicazione: (2024)
AIPO: Learning to Reason from Active Interaction
di: Liu, Junnan, et al.
Pubblicazione: (2026)
di: Liu, Junnan, et al.
Pubblicazione: (2026)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
di: Chen, Xingwu, et al.
Pubblicazione: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
di: Zhang, Yi, et al.
Pubblicazione: (2025)
di: Zhang, Yi, et al.
Pubblicazione: (2025)
On the Memorization of Consistency Distillation for Diffusion Models
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)
Improving Group Robustness on Spurious Correlation Requires Preciser Group Inference
di: Han, Yujin, et al.
Pubblicazione: (2024)
di: Han, Yujin, et al.
Pubblicazione: (2024)
Cold-Start Recommendation towards the Era of Large Language Models (LLMs): A Comprehensive Survey and Roadmap
di: Zhang, Weizhi, et al.
Pubblicazione: (2025)
di: Zhang, Weizhi, et al.
Pubblicazione: (2025)
Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
di: Geng, Haopeng, et al.
Pubblicazione: (2024)
Aldo‐Keto Reductase Family 1 Member B10 as a Therapeutic Frontier: Advances in Inhibitor Design and Multimodal Approaches
di: Jikuan Shao, et al.
Pubblicazione: (2026)
di: Jikuan Shao, et al.
Pubblicazione: (2026)
Woodpecker: Hallucination Correction for Multimodal Large Language Models
di: Yin, Shukang, et al.
Pubblicazione: (2023)
di: Yin, Shukang, et al.
Pubblicazione: (2023)
ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs
di: Chen, Keyu, et al.
Pubblicazione: (2025)
di: Chen, Keyu, et al.
Pubblicazione: (2025)
Does higher interpretability imply better utility? A Pairwise Analysis on Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2025)
di: Wang, Xu, et al.
Pubblicazione: (2025)
Fibro‐Adipogenic Progenitors Regulate Orofacial Neuromuscular Junction Regeneration via Myostatin
di: Ruizhi Li, et al.
Pubblicazione: (2026)
di: Ruizhi Li, et al.
Pubblicazione: (2026)
Aria: An Open Multimodal Native Mixture-of-Experts Model
di: Li, Dongxu, et al.
Pubblicazione: (2024)
di: Li, Dongxu, et al.
Pubblicazione: (2024)
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
di: Ye, Zhipeng, et al.
Pubblicazione: (2026)
di: Ye, Zhipeng, et al.
Pubblicazione: (2026)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
di: Shi, Mengqi, et al.
Pubblicazione: (2026)
di: Shi, Mengqi, et al.
Pubblicazione: (2026)
FIPO: Free-form Instruction-oriented Prompt Optimization with Preference Dataset and Modular Fine-tuning Schema
di: Lu, Junru, et al.
Pubblicazione: (2024)
di: Lu, Junru, et al.
Pubblicazione: (2024)
Optimising Robotic Radical Gastrectomy: A Technical Roadmap for Bipolar Forceps‐Assisted Lymphadenectomy and Reconstruction
di: Fengyuan Li, et al.
Pubblicazione: (2026)
di: Fengyuan Li, et al.
Pubblicazione: (2026)
SIDE: Surrogate Conditional Data Extraction from Diffusion Models
di: Chen, Yunhao, et al.
Pubblicazione: (2024)
di: Chen, Yunhao, et al.
Pubblicazione: (2024)
F-Adapter: Frequency-Adaptive Parameter-Efficient Fine-Tuning in Scientific Machine Learning
di: Zhang, Hangwei, et al.
Pubblicazione: (2025)
di: Zhang, Hangwei, et al.
Pubblicazione: (2025)
Global, regional, and national burdens of Alzheimer's disease and other forms of dementia in the elderly population from 1999 to 2019: A trend analysis based on the Global Burden of Disease Study 2019
di: Mengdan Su, et al.
Pubblicazione: (2024)
di: Mengdan Su, et al.
Pubblicazione: (2024)
How Are Quantum Eigenfunctions of Hydrogen Atom Related To Its Classical Elliptic Orbits?
di: Yin, Yixuan, et al.
Pubblicazione: (2024)
di: Yin, Yixuan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models
di: Lu, Junru, et al.
Pubblicazione: (2025) -
Deep Tabular Research via Continual Experience-Driven Execution
di: Dong, Junnan, et al.
Pubblicazione: (2026) -
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
di: Lu, Wensheng, et al.
Pubblicazione: (2025) -
Adaptive Dual Reasoner: Large Reasoning Models Can Think Efficiently by Hybrid Reasoning
di: Zhang, Yujian, et al.
Pubblicazione: (2025) -
Structured Role-Aware Policy Optimization for Multimodal Reasoning
di: Jiang, Bingqing, et al.
Pubblicazione: (2026)