A unified multimodal understanding and generation model for cross-disciplinary scientific research
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Xiaomeng, Tan, Zhiyu, Zhong, Xiaohui, Yang, Mengping, Huang, Qiusheng, Chen, Lei, Wu, Libo, Li, Hao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
FuXi-RTM: A Physics-Guided Prediction Framework with Radiative Transfer Modeling
di: Huang, Qiusheng, et al.
Pubblicazione: (2025)
di: Huang, Qiusheng, et al.
Pubblicazione: (2025)
Data-driven ensemble prediction of the global ocean
di: Huang, Qiusheng, et al.
Pubblicazione: (2026)
di: Huang, Qiusheng, et al.
Pubblicazione: (2026)
AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecasts
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
Roadmap for using large language models (LLMs) to accelerate cross-disciplinary research with an example from computational biology
di: Ke, Ruian, et al.
Pubblicazione: (2025)
di: Ke, Ruian, et al.
Pubblicazione: (2025)
FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution
di: Huang, Qiusheng, et al.
Pubblicazione: (2025)
di: Huang, Qiusheng, et al.
Pubblicazione: (2025)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
di: Xing, Yang, et al.
Pubblicazione: (2026)
di: Xing, Yang, et al.
Pubblicazione: (2026)
COMET: "Cone of experience" enhanced large multimodal model for mathematical problem generation
di: Liu, Sannyuya, et al.
Pubblicazione: (2024)
di: Liu, Sannyuya, et al.
Pubblicazione: (2024)
Image Synthesis under Limited Data: A Survey and Taxonomy
di: Yang, Mengping, et al.
Pubblicazione: (2023)
di: Yang, Mengping, et al.
Pubblicazione: (2023)
UniGenX: a unified generative foundation model that couples sequence, structure and function to accelerate scientific design across proteins, molecules and materials
di: Zhang, Gongbo, et al.
Pubblicazione: (2025)
di: Zhang, Gongbo, et al.
Pubblicazione: (2025)
IWISDM: Assessing instruction following in multimodal models at scale
di: Lei, Xiaoxuan, et al.
Pubblicazione: (2024)
di: Lei, Xiaoxuan, et al.
Pubblicazione: (2024)
Generative artificial intelligence improves projections of climate extremes
di: Tie, Ruian, et al.
Pubblicazione: (2025)
di: Tie, Ruian, et al.
Pubblicazione: (2025)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
Enhanced predictions of the Madden-Julian oscillation using the FuXi-S2S machine learning model: Insights into physical mechanisms
di: Cao, Can, et al.
Pubblicazione: (2025)
di: Cao, Can, et al.
Pubblicazione: (2025)
FastRM: An efficient and automatic explainability framework for multimodal generative models
di: Stan, Gabriela Ben-Melech, et al.
Pubblicazione: (2024)
di: Stan, Gabriela Ben-Melech, et al.
Pubblicazione: (2024)
Simulating clinical interventions with a generative multimodal model of human physiology
di: Lutsker, Guy, et al.
Pubblicazione: (2026)
di: Lutsker, Guy, et al.
Pubblicazione: (2026)
The unified cross-disciplinary model of the operation of neurons
di: Végh, János
Pubblicazione: (2025)
di: Végh, János
Pubblicazione: (2025)
UF-AMA: A unified framework for cross-domain emotion recognition via adaptive multimodal alignment
di: Wang, Zheng, et al.
Pubblicazione: (2026)
di: Wang, Zheng, et al.
Pubblicazione: (2026)
Physics-based phenomenological characterization of cross-modal bias in multimodal models
di: Kim, Hyeongmo, et al.
Pubblicazione: (2026)
di: Kim, Hyeongmo, et al.
Pubblicazione: (2026)
No Free Lunch from Audio Pretraining in Bioacoustics: A Benchmark Study of Embeddings
di: Chen, Chenggang, et al.
Pubblicazione: (2025)
di: Chen, Chenggang, et al.
Pubblicazione: (2025)
EHWGesture -- A dataset for multimodal understanding of clinical gestures
di: Amprimo, Gianluca, et al.
Pubblicazione: (2025)
di: Amprimo, Gianluca, et al.
Pubblicazione: (2025)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
di: Yang, Mengping, et al.
Pubblicazione: (2026)
di: Yang, Mengping, et al.
Pubblicazione: (2026)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
di: Gemini Team, et al.
Pubblicazione: (2024)
di: Gemini Team, et al.
Pubblicazione: (2024)
Symbotunes: unified hub for symbolic music generative models
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
di: Skierś, Paweł, et al.
Pubblicazione: (2024)
AI-generated Essays: Characteristics and Implications on Automated Scoring and Academic Integrity
di: Zhong, Yang, et al.
Pubblicazione: (2024)
di: Zhong, Yang, et al.
Pubblicazione: (2024)
From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine
di: Buess, Lukas, et al.
Pubblicazione: (2025)
di: Buess, Lukas, et al.
Pubblicazione: (2025)
RareAgents: Autonomous Multi-disciplinary Team for Rare Disease Diagnosis and Treatment
di: Chen, Xuanzhong, et al.
Pubblicazione: (2024)
di: Chen, Xuanzhong, et al.
Pubblicazione: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
di: Tan, Zhiyu, et al.
Pubblicazione: (2024)
Can LLM generate interesting mathematical research problems?
di: Chen, Xiaoyang, et al.
Pubblicazione: (2026)
di: Chen, Xiaoyang, et al.
Pubblicazione: (2026)
NORA: A Harness-Engineered Autonomous Research Agent for End-to-End Spatial Data Science
di: Zhou, Bing, et al.
Pubblicazione: (2026)
di: Zhou, Bing, et al.
Pubblicazione: (2026)
Toward a unified framework for data-efficient evaluation of large language models
di: Liao, Lele, et al.
Pubblicazione: (2025)
di: Liao, Lele, et al.
Pubblicazione: (2025)
A neural network for modeling human concept formation, understanding and communication
di: Guo, Liangxuan, et al.
Pubblicazione: (2026)
di: Guo, Liangxuan, et al.
Pubblicazione: (2026)
End-to-end autonomous scientific discovery on a real optical platform
di: Yang, Shuxing, et al.
Pubblicazione: (2026)
di: Yang, Shuxing, et al.
Pubblicazione: (2026)
Explaining latent representations of generative models with large multimodal models
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
di: Zhu, Mengdan, et al.
Pubblicazione: (2024)
Rethinking the Chain-of-Thought: The Roles of In-Context Learning and Pre-trained Priors
di: Yang, Hao, et al.
Pubblicazione: (2025)
di: Yang, Hao, et al.
Pubblicazione: (2025)
Do generative video models understand physical principles?
di: Motamed, Saman, et al.
Pubblicazione: (2025)
di: Motamed, Saman, et al.
Pubblicazione: (2025)
RadHiera: Semantic Hierarchical Reinforcement Learning for Medical Report Generation
di: Du, Bodong, et al.
Pubblicazione: (2025)
di: Du, Bodong, et al.
Pubblicazione: (2025)
Can MLLMs generate human-like feedback in grading multimodal short answers?
di: Sil, Pritam, et al.
Pubblicazione: (2024)
di: Sil, Pritam, et al.
Pubblicazione: (2024)
Separable neural architectures as a primitive for unified predictive and generative intelligence
di: Batley, Reza T., et al.
Pubblicazione: (2026)
di: Batley, Reza T., et al.
Pubblicazione: (2026)
Documenti analoghi
-
SIPO: Stabilized and Improved Preference Optimization for Aligning Diffusion Models
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025) -
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025) -
FuXi-RTM: A Physics-Guided Prediction Framework with Radiative Transfer Modeling
di: Huang, Qiusheng, et al.
Pubblicazione: (2025) -
Data-driven ensemble prediction of the global ocean
di: Huang, Qiusheng, et al.
Pubblicazione: (2026) -
AviaSafe: A Physics-Informed Data-Driven Model for Aviation Safety-Critical Cloud Forecasts
di: Zhu, Zijian, et al.
Pubblicazione: (2026)