Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Xiaoying, Peng, Da, Zhang, Yipeng, Guo, Zonghao, Wu, Chengyue, Huang, Jen-Tse, Chen, Chi, Ke, Wei, Meng, Helen, Sun, Maosong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
di: Sun, Shichu, et al.
Pubblicazione: (2025)
di: Sun, Shichu, et al.
Pubblicazione: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
di: Zhang, Yipeng, et al.
Pubblicazione: (2024)
di: Zhang, Yipeng, et al.
Pubblicazione: (2024)
RAM: Towards an Ever-Improving Memory System by Learning from Communications
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
di: Li, Jiaqi, et al.
Pubblicazione: (2024)
Visual Jigsaw Post-Training Improves MLLMs
di: Wu, Penghao, et al.
Pubblicazione: (2025)
di: Wu, Penghao, et al.
Pubblicazione: (2025)
On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking
di: Wang, Fang, et al.
Pubblicazione: (2025)
di: Wang, Fang, et al.
Pubblicazione: (2025)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2024)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
di: Zhang, Xiaoying, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2025)
Risk Factors for First‐Ever Diabetes‐Related Foot Ulcer: A Systematic Review and Meta‐Analysis
di: Tao Yan, et al.
Pubblicazione: (2025)
di: Tao Yan, et al.
Pubblicazione: (2025)
Growing Visual Generative Capacity for Pre-Trained MLLMs
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
di: Mitra, Vikramjit, et al.
Pubblicazione: (2024)
di: Mitra, Vikramjit, et al.
Pubblicazione: (2024)
The "Ever Green" Interpretations of the "Library Bill of Rights"
di: Adams, Helen R.
Pubblicazione: (2010)
di: Adams, Helen R.
Pubblicazione: (2010)
Markovian Pre-Trained Transformer for Next-Item Recommendation
di: Xu, Cong, et al.
Pubblicazione: (2026)
di: Xu, Cong, et al.
Pubblicazione: (2026)
Incidence and Factors Associated With Cognitive Impairment 90 Days After First Ever Ischemic Stroke
di: Małgorzata Dec‐Ćwiek, et al.
Pubblicazione: (2025)
di: Małgorzata Dec‐Ćwiek, et al.
Pubblicazione: (2025)
The Impact of Perceived Stress on Frailty in Patients With First‐Ever Ischemic Stroke: A Cross‐Sectional Study
di: Xin Wen, et al.
Pubblicazione: (2026)
di: Xin Wen, et al.
Pubblicazione: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
di: Lin, Junming, et al.
Pubblicazione: (2024)
di: Lin, Junming, et al.
Pubblicazione: (2024)
Retrieve Then Rerank: An End‐to‐End Learning Paradigm for Biomedical Entity Linking
di: Yuling Cao, et al.
Pubblicazione: (2025)
di: Yuling Cao, et al.
Pubblicazione: (2025)
World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
di: Guan, Yanchen, et al.
Pubblicazione: (2025)
di: Guan, Yanchen, et al.
Pubblicazione: (2025)
Therapeutic Efficacy of Xanomeline on Psychotropic Drug‐Induced Cognitive Dysfunction in Rats
di: Cheng Peng, et al.
Pubblicazione: (2025)
di: Cheng Peng, et al.
Pubblicazione: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
di: Fu, Chaoyou, et al.
Pubblicazione: (2024)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
di: Yu, Tianyu, et al.
Pubblicazione: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
di: Conciatore, Dino, et al.
Pubblicazione: (2026)
di: Conciatore, Dino, et al.
Pubblicazione: (2026)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
di: Zhang, Xiaoying, et al.
Pubblicazione: (2026)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2026)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
di: Tian, Junjiao, et al.
Pubblicazione: (2024)
di: Tian, Junjiao, et al.
Pubblicazione: (2024)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
di: Liu, Ruicong, et al.
Pubblicazione: (2025)
di: Liu, Ruicong, et al.
Pubblicazione: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
Sliding Window Training -- Utilizing Historical Recommender Systems Data for Foundation Models
di: Joshi, Swanand, et al.
Pubblicazione: (2024)
di: Joshi, Swanand, et al.
Pubblicazione: (2024)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
di: Hono, Yukiya, et al.
Pubblicazione: (2023)
di: Hono, Yukiya, et al.
Pubblicazione: (2023)
Ever Faithful
di: Sartorious, David
Pubblicazione: (2018)
di: Sartorious, David
Pubblicazione: (2018)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
Mission Analysis for the First-Ever Saturn Trojan 2019 UO$_{14}$
di: Takao, Yuki
Pubblicazione: (2025)
di: Takao, Yuki
Pubblicazione: (2025)
Ever-Improving Test Suite by Leveraging Large Language Models
di: Qiu, Ketai
Pubblicazione: (2025)
di: Qiu, Ketai
Pubblicazione: (2025)
XDXD: End-to-end crystal structure determination with low resolution X-ray diffraction
di: Zhao, Jiale, et al.
Pubblicazione: (2025)
di: Zhao, Jiale, et al.
Pubblicazione: (2025)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
di: Lin, Chun-Hsien, et al.
Pubblicazione: (2024)
The Ever-Evolving Science Exam
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
di: Yao, Huanjin, et al.
Pubblicazione: (2025)
First Steps, Lasting Impact: Platform-Aware Forensics for the Next Generation of Analysts
di: Jain, Vinayak, et al.
Pubblicazione: (2026)
di: Jain, Vinayak, et al.
Pubblicazione: (2026)
The Grievance: First Step in Improved Library Government
di: Volkersz, Evert
Pubblicazione: (1969)
di: Volkersz, Evert
Pubblicazione: (1969)
Documenti analoghi
-
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
di: Huang, Jen-Tse, et al.
Pubblicazione: (2025) -
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
di: Sun, Shichu, et al.
Pubblicazione: (2025) -
Video-R1: Reinforcing Video Reasoning in MLLMs
di: Feng, Kaituo, et al.
Pubblicazione: (2025) -
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
di: Zhang, Yipeng, et al.
Pubblicazione: (2024) -
RAM: Towards an Ever-Improving Memory System by Learning from Communications
di: Li, Jiaqi, et al.
Pubblicazione: (2024)