Thinker: A vision-language foundation model for embodied intelligence
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Baiyu, Luo, Daqin, Yang, Junpeng, Wang, Jiyuan, Zhang, Yixuan, Shi, Hailin, Jiao, Jichao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge Devices
by: Pan, Baiyu, et al.
Published: (2024)
by: Pan, Baiyu, et al.
Published: (2024)
GenRL: Multimodal-foundation world models for generalization in embodied agents
by: Mazzaglia, Pietro, et al.
Published: (2024)
by: Mazzaglia, Pietro, et al.
Published: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches
by: Mumuni, Alhassan, et al.
Published: (2025)
by: Mumuni, Alhassan, et al.
Published: (2025)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Generalizing vision-language models to novel domains: A comprehensive survey
by: Li, Xinyao, et al.
Published: (2025)
by: Li, Xinyao, et al.
Published: (2025)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
Beyond the Hype: A dispassionate look at vision-language models in medical scenario
by: Nan, Yang, et al.
Published: (2024)
by: Nan, Yang, et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
The Sampling-Gaussian for stereo matching
by: Pan, Baiyu, et al.
Published: (2024)
by: Pan, Baiyu, et al.
Published: (2024)
S4DL: Shift-sensitive Spatial-Spectral Disentangling Learning for Hyperspectral Image Unsupervised Domain Adaptation
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Representation geometry shapes task performance in vision-language modeling for CT enterography
by: Minoccheri, Cristian, et al.
Published: (2026)
by: Minoccheri, Cristian, et al.
Published: (2026)
TasselNetV4: A vision foundation model for cross-scene, cross-scale, and cross-species plant counting
by: Hu, Xiaonan, et al.
Published: (2025)
by: Hu, Xiaonan, et al.
Published: (2025)
Improving vision-language alignment with graph spiking hybrid Networks
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
MIAR: Modality Interaction and Alignment Representation Fuison for Multimodal Emotion
by: Zhu, Jichao, et al.
Published: (2026)
by: Zhu, Jichao, et al.
Published: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Towards a vision foundation model for comprehensive assessment of Cardiac MRI
by: Jacob, Athira J, et al.
Published: (2024)
by: Jacob, Athira J, et al.
Published: (2024)
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
by: Yang, Hanqing, et al.
Published: (2026)
by: Yang, Hanqing, et al.
Published: (2026)
BRAVE: Broadening the visual encoding of vision-language models
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
by: Kar, Oğuzhan Fatih, et al.
Published: (2024)
Explainable artificial intelligence (XAI): from inherent explainability to large language models
by: Mumuni, Fuseini, et al.
Published: (2025)
by: Mumuni, Fuseini, et al.
Published: (2025)
Enabling clinical use of foundation models for computational pathology
by: Henriksen, Audun L, et al.
Published: (2026)
by: Henriksen, Audun L, et al.
Published: (2026)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
by: Dong, Owen, et al.
Published: (2026)
by: Dong, Owen, et al.
Published: (2026)
MI-VisionShot: Few-shot adaptation of vision-language models for slide-level classification of histopathological images
by: Meseguer, Pablo, et al.
Published: (2024)
by: Meseguer, Pablo, et al.
Published: (2024)
An interpretable framework using foundation models for fish sex identification
by: Miao, Zheng, et al.
Published: (2026)
by: Miao, Zheng, et al.
Published: (2026)
Compound Expression Recognition via Multi Model Ensemble
by: Yu, Jun, et al.
Published: (2024)
by: Yu, Jun, et al.
Published: (2024)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
by: Liu, Xiaohong, et al.
Published: (2024)
by: Liu, Xiaohong, et al.
Published: (2024)
Prompting with the human-touch: evaluating model-sensitivity of foundation models for musculoskeletal CT segmentation
by: Magg, Caroline, et al.
Published: (2026)
by: Magg, Caroline, et al.
Published: (2026)
Concurrent validity of computer-vision artificial intelligence player tracking software using broadcast footage
by: Crang, Zachary L., et al.
Published: (2025)
by: Crang, Zachary L., et al.
Published: (2025)
ActiveMark: on watermarking of visual foundation models via massive activations
by: Chistyakova, Anna, et al.
Published: (2025)
by: Chistyakova, Anna, et al.
Published: (2025)
Similar Items
-
Distill-then-prune: An Efficient Compression Framework for Real-time Stereo Matching Network on Edge Devices
by: Pan, Baiyu, et al.
Published: (2024) -
GenRL: Multimodal-foundation world models for generalization in embodied agents
by: Mazzaglia, Pietro, et al.
Published: (2024) -
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024) -
Large language models for artificial general intelligence (AGI): A survey of foundational principles and approaches
by: Mumuni, Alhassan, et al.
Published: (2025) -
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)