Enhancing Cognition and Explainability of Multimodal Foundation Models with Self-Synthesized Data
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Yucheng, Li, Quanzheng, Sun, Jin, Li, Xiang, Liu, Ninghao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Improving Small Object Grounding in LVLMs
by: Yang, Tianze, et al.
Published: (2026)
by: Yang, Tianze, et al.
Published: (2026)
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
Measurement Geometry and Design for Trustworthy Generative Inverse Problems
by: Jin, Pengfei, et al.
Published: (2026)
by: Jin, Pengfei, et al.
Published: (2026)
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025)
by: Yang, Tianze, et al.
Published: (2025)
ECHOPulse: ECG controlled echocardio-grams video generation
by: Li, Yiwei, et al.
Published: (2024)
by: Li, Yiwei, et al.
Published: (2024)
Streamlined Photoacoustic Image Processing with Foundation Models: A Training-Free Solution
by: Deng, Handi, et al.
Published: (2024)
by: Deng, Handi, et al.
Published: (2024)
SeaMo: A Season-Aware Multimodal Foundation Model for Remote Sensing
by: Li, Xuyang, et al.
Published: (2024)
by: Li, Xuyang, et al.
Published: (2024)
High-Fidelity 3D Lung CT Synthesis in ARDS Swine Models Using Score-Based 3D Residual Diffusion Models
by: Yoon, Siyeop, et al.
Published: (2024)
by: Yoon, Siyeop, et al.
Published: (2024)
LRM-Zero: Training Large Reconstruction Models with Synthesized Data
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
Self-Supervised Weight Templates for Scalable Vision Model Initialization
by: Xie, Yucheng, et al.
Published: (2026)
by: Xie, Yucheng, et al.
Published: (2026)
Large Language Models and Foundation Models in Smart Agriculture: Basics, Opportunities, and Challenges
by: Li, Jiajia, et al.
Published: (2023)
by: Li, Jiajia, et al.
Published: (2023)
Intern-S1: A Scientific Multimodal Foundation Model
by: Bai, Lei, et al.
Published: (2025)
by: Bai, Lei, et al.
Published: (2025)
LSKNet: A Foundation Lightweight Backbone for Remote Sensing
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Data Attribution for Text-to-Image Models by Unlearning Synthesized Images
by: Wang, Sheng-Yu, et al.
Published: (2024)
by: Wang, Sheng-Yu, et al.
Published: (2024)
Synthesizing Realistic Data for Table Recognition
by: Hou, Qiyu, et al.
Published: (2024)
by: Hou, Qiyu, et al.
Published: (2024)
CLIMB: Data Foundations for Large Scale Multimodal Clinical Foundation Models
by: Dai, Wei, et al.
Published: (2025)
by: Dai, Wei, et al.
Published: (2025)
Synthesizer Based Efficient Self-Attention for Vision Tasks
by: Zhu, Guangyang, et al.
Published: (2022)
by: Zhu, Guangyang, et al.
Published: (2022)
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
by: Zou, Yicheng, et al.
Published: (2026)
by: Zou, Yicheng, et al.
Published: (2026)
Research on the Spatial Data Intelligent Foundation Model
by: Wang, Shaohua, et al.
Published: (2024)
by: Wang, Shaohua, et al.
Published: (2024)
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
by: Chen, Xiuwei, et al.
Published: (2025)
by: Chen, Xiuwei, et al.
Published: (2025)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026)
by: Li, Muyang, et al.
Published: (2026)
Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs
by: Jiang, Dawen, et al.
Published: (2025)
by: Jiang, Dawen, et al.
Published: (2025)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
XIMAGENET-12: An Explainable AI Benchmark Dataset for Model Robustness Evaluation
by: Li, Qiang, et al.
Published: (2023)
by: Li, Qiang, et al.
Published: (2023)
Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis
by: Agarwal, Naaisha, et al.
Published: (2026)
by: Agarwal, Naaisha, et al.
Published: (2026)
Backdoor Attack on Unpaired Medical Image-Text Foundation Models: A Pilot Study on MedCLIP
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
by: Qiu, Haibo, et al.
Published: (2025)
by: Qiu, Haibo, et al.
Published: (2025)
VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
by: Shen, Lefei, et al.
Published: (2025)
by: Shen, Lefei, et al.
Published: (2025)
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
by: Dang, Ronghao, et al.
Published: (2025)
by: Dang, Ronghao, et al.
Published: (2025)
Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
by: Zhu, Xun, et al.
Published: (2025)
by: Zhu, Xun, et al.
Published: (2025)
Beyond Adapter Retrieval: Latent Geometry-Preserving Composition via Sparse Task Projection
by: Jin, Pengfei, et al.
Published: (2024)
by: Jin, Pengfei, et al.
Published: (2024)
Foundation Model-oriented Robustness: Robust Image Model Evaluation with Pretrained Models
by: Zhang, Peiyan, et al.
Published: (2023)
by: Zhang, Peiyan, et al.
Published: (2023)
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
by: Li, Zongxia, et al.
Published: (2026)
by: Li, Zongxia, et al.
Published: (2026)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
by: Du, Zilin, et al.
Published: (2024)
by: Du, Zilin, et al.
Published: (2024)
Cross-Modality Clustering-based Self-Labeling for Multimodal Data Classification
by: Zyblewski, Paweł, et al.
Published: (2024)
by: Zyblewski, Paweł, et al.
Published: (2024)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
Enhancing Interpretability of AR-SSVEP-Based Motor Intention Recognition via CNN-BiLSTM and SHAP Analysis on EEG Data
by: Yang, Lin, et al.
Published: (2025)
by: Yang, Lin, et al.
Published: (2025)
EchoFM: Foundation Model for Generalizable Echocardiogram Analysis
by: Kim, Sekeun, et al.
Published: (2024)
by: Kim, Sekeun, et al.
Published: (2024)
Closer to Reality: Practical Semi-Supervised Federated Learning for Foundation Model Adaptation
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
Similar Items
-
Self-Improving Small Object Grounding in LVLMs
by: Yang, Tianze, et al.
Published: (2026) -
Concept-Centric Token Interpretation for Vector-Quantized Generative Models
by: Yang, Tianze, et al.
Published: (2025) -
Measurement Geometry and Design for Trustworthy Generative Inverse Problems
by: Jin, Pengfei, et al.
Published: (2026) -
Common Inpainted Objects In-N-Out of Context
by: Yang, Tianze, et al.
Published: (2025) -
ECHOPulse: ECG controlled echocardio-grams video generation
by: Li, Yiwei, et al.
Published: (2024)