Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Xiaoying, Peng, Da, Zhang, Yipeng, Guo, Zonghao, Wu, Chengyue, Huang, Jen-Tse, Chen, Chi, Ke, Wei, Meng, Helen, Sun, Maosong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
por: Huang, Jen-Tse, et al.
Publicado: (2025)
por: Huang, Jen-Tse, et al.
Publicado: (2025)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
por: Sun, Shichu, et al.
Publicado: (2025)
por: Sun, Shichu, et al.
Publicado: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
por: Feng, Kaituo, et al.
Publicado: (2025)
por: Feng, Kaituo, et al.
Publicado: (2025)
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
por: Zhang, Yipeng, et al.
Publicado: (2024)
por: Zhang, Yipeng, et al.
Publicado: (2024)
RAM: Towards an Ever-Improving Memory System by Learning from Communications
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
Visual Jigsaw Post-Training Improves MLLMs
por: Wu, Penghao, et al.
Publicado: (2025)
por: Wu, Penghao, et al.
Publicado: (2025)
On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking
por: Wang, Fang, et al.
Publicado: (2025)
por: Wang, Fang, et al.
Publicado: (2025)
Self-Tuning: Instructing LLMs to Effectively Acquire New Knowledge through Self-Teaching
por: Zhang, Xiaoying, et al.
Publicado: (2024)
por: Zhang, Xiaoying, et al.
Publicado: (2024)
Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback
por: Zhang, Xiaoying, et al.
Publicado: (2025)
por: Zhang, Xiaoying, et al.
Publicado: (2025)
Risk Factors for First‐Ever Diabetes‐Related Foot Ulcer: A Systematic Review and Meta‐Analysis
por: Tao Yan, et al.
Publicado: (2025)
por: Tao Yan, et al.
Publicado: (2025)
Growing Visual Generative Capacity for Pre-Trained MLLMs
por: Wang, Hanyu, et al.
Publicado: (2025)
por: Wang, Hanyu, et al.
Publicado: (2025)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
por: Mitra, Vikramjit, et al.
Publicado: (2024)
por: Mitra, Vikramjit, et al.
Publicado: (2024)
The "Ever Green" Interpretations of the "Library Bill of Rights"
por: Adams, Helen R.
Publicado: (2010)
por: Adams, Helen R.
Publicado: (2010)
Markovian Pre-Trained Transformer for Next-Item Recommendation
por: Xu, Cong, et al.
Publicado: (2026)
por: Xu, Cong, et al.
Publicado: (2026)
Incidence and Factors Associated With Cognitive Impairment 90 Days After First Ever Ischemic Stroke
por: Małgorzata Dec‐Ćwiek, et al.
Publicado: (2025)
por: Małgorzata Dec‐Ćwiek, et al.
Publicado: (2025)
The Impact of Perceived Stress on Frailty in Patients With First‐Ever Ischemic Stroke: A Cross‐Sectional Study
por: Xin Wen, et al.
Publicado: (2026)
por: Xin Wen, et al.
Publicado: (2026)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
por: Lin, Junming, et al.
Publicado: (2024)
por: Lin, Junming, et al.
Publicado: (2024)
Retrieve Then Rerank: An End‐to‐End Learning Paradigm for Biomedical Entity Linking
por: Yuling Cao, et al.
Publicado: (2025)
por: Yuling Cao, et al.
Publicado: (2025)
World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving
por: Guan, Yanchen, et al.
Publicado: (2025)
por: Guan, Yanchen, et al.
Publicado: (2025)
Therapeutic Efficacy of Xanomeline on Psychotropic Drug‐Induced Cognitive Dysfunction in Rats
por: Cheng Peng, et al.
Publicado: (2025)
por: Cheng Peng, et al.
Publicado: (2025)
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
por: Fu, Chaoyou, et al.
Publicado: (2024)
por: Fu, Chaoyou, et al.
Publicado: (2024)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
por: Yu, Tianyu, et al.
Publicado: (2025)
por: Yu, Tianyu, et al.
Publicado: (2025)
Beyond Pre-Training: The Full Lifecycle of Foundation Models on HPC Systems
por: Conciatore, Dino, et al.
Publicado: (2026)
por: Conciatore, Dino, et al.
Publicado: (2026)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
por: Dai, Dongyang, et al.
Publicado: (2025)
por: Dai, Dongyang, et al.
Publicado: (2025)
RetroAgent: From Solving to Evolving via Retrospective Dual Intrinsic Feedback
por: Zhang, Xiaoying, et al.
Publicado: (2026)
por: Zhang, Xiaoying, et al.
Publicado: (2026)
Rethinking Weight Decay for Robust Fine-Tuning of Foundation Models
por: Tian, Junjiao, et al.
Publicado: (2024)
por: Tian, Junjiao, et al.
Publicado: (2024)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
por: Liu, Ruicong, et al.
Publicado: (2025)
por: Liu, Ruicong, et al.
Publicado: (2025)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
por: Tasnim, Nazia, et al.
Publicado: (2025)
por: Tasnim, Nazia, et al.
Publicado: (2025)
Sliding Window Training -- Utilizing Historical Recommender Systems Data for Foundation Models
por: Joshi, Swanand, et al.
Publicado: (2024)
por: Joshi, Swanand, et al.
Publicado: (2024)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
por: Hono, Yukiya, et al.
Publicado: (2023)
por: Hono, Yukiya, et al.
Publicado: (2023)
Ever Faithful
por: Sartorious, David
Publicado: (2018)
por: Sartorious, David
Publicado: (2018)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
por: Hsiao, Chi-Yuan, et al.
Publicado: (2025)
por: Hsiao, Chi-Yuan, et al.
Publicado: (2025)
Mission Analysis for the First-Ever Saturn Trojan 2019 UO$_{14}$
por: Takao, Yuki
Publicado: (2025)
por: Takao, Yuki
Publicado: (2025)
Ever-Improving Test Suite by Leveraging Large Language Models
por: Qiu, Ketai
Publicado: (2025)
por: Qiu, Ketai
Publicado: (2025)
XDXD: End-to-end crystal structure determination with low resolution X-ray diffraction
por: Zhao, Jiale, et al.
Publicado: (2025)
por: Zhao, Jiale, et al.
Publicado: (2025)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
por: Lin, Chun-Hsien, et al.
Publicado: (2024)
The Ever-Evolving Science Exam
por: Wang, Junying, et al.
Publicado: (2025)
por: Wang, Junying, et al.
Publicado: (2025)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
por: Yao, Huanjin, et al.
Publicado: (2025)
por: Yao, Huanjin, et al.
Publicado: (2025)
First Steps, Lasting Impact: Platform-Aware Forensics for the Next Generation of Analysts
por: Jain, Vinayak, et al.
Publicado: (2026)
por: Jain, Vinayak, et al.
Publicado: (2026)
The Grievance: First Step in Improved Library Government
por: Volkersz, Evert
Publicado: (1969)
por: Volkersz, Evert
Publicado: (1969)
Ejemplares similares
-
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
por: Huang, Jen-Tse, et al.
Publicado: (2025) -
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
por: Sun, Shichu, et al.
Publicado: (2025) -
Video-R1: Reinforcing Video Reasoning in MLLMs
por: Feng, Kaituo, et al.
Publicado: (2025) -
LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer
por: Zhang, Yipeng, et al.
Publicado: (2024) -
RAM: Towards an Ever-Improving Memory System by Learning from Communications
por: Li, Jiaqi, et al.
Publicado: (2024)