Accelerating Pre-training of Multimodal LLMs via Chain-of-Sight
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Ziyuan, Ji, Kaixiang, Gong, Biao, Qing, Zhiwu, Zhang, Qinglong, Zheng, Kecheng, Wang, Jian, Chen, Jingdong, Yang, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024)
von: Li, Honglin, et al.
Veröffentlicht: (2024)
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
von: Zeng, Weili, et al.
Veröffentlicht: (2025)
von: Zeng, Weili, et al.
Veröffentlicht: (2025)
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025)
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
von: Gong, Biao, et al.
Veröffentlicht: (2023)
von: Gong, Biao, et al.
Veröffentlicht: (2023)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
von: Du, Yipeng, et al.
Veröffentlicht: (2025)
von: Du, Yipeng, et al.
Veröffentlicht: (2025)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
von: Shi, Shuwei, et al.
Veröffentlicht: (2024)
von: Shi, Shuwei, et al.
Veröffentlicht: (2024)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
von: Yu, Hu, et al.
Veröffentlicht: (2025)
von: Yu, Hu, et al.
Veröffentlicht: (2025)
POA: Pre-training Once for Models of All Sizes
von: Zhang, Yingying, et al.
Veröffentlicht: (2024)
von: Zhang, Yingying, et al.
Veröffentlicht: (2024)
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
von: Yao, Xin, et al.
Veröffentlicht: (2025)
von: Yao, Xin, et al.
Veröffentlicht: (2025)
DreamLIP: Language-Image Pre-training with Long Captions
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
Compressing then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
von: Li, Da, et al.
Veröffentlicht: (2025)
von: Li, Da, et al.
Veröffentlicht: (2025)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
von: Chen, Cong, et al.
Veröffentlicht: (2025)
von: Chen, Cong, et al.
Veröffentlicht: (2025)
LumiSculpt: Enabling Consistent Portrait Lighting in Video Generation
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2024)
StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
von: Li, Wen, et al.
Veröffentlicht: (2024)
von: Li, Wen, et al.
Veröffentlicht: (2024)
Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyuan, et al.
Veröffentlicht: (2024)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Huang, Ziyuan, et al.
Veröffentlicht: (2025)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024)
von: Lu, Fan, et al.
Veröffentlicht: (2024)
PLIP: Language-Image Pre-training for Person Representation Learning
von: Zuo, Jialong, et al.
Veröffentlicht: (2023)
von: Zuo, Jialong, et al.
Veröffentlicht: (2023)
Multimodal Medical Image Classification via Synergistic Learning Pre-training
von: Lin, Qinghua, et al.
Veröffentlicht: (2025)
von: Lin, Qinghua, et al.
Veröffentlicht: (2025)
MotionChain: Conversational Motion Controllers via Multimodal Prompts
von: Jiang, Biao, et al.
Veröffentlicht: (2024)
von: Jiang, Biao, et al.
Veröffentlicht: (2024)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
von: Zhang, Dingwen, et al.
Veröffentlicht: (2024)
von: Zhang, Dingwen, et al.
Veröffentlicht: (2024)
TIP: Tabular-Image Pre-training for Multimodal Classification with Incomplete Data
von: Du, Siyi, et al.
Veröffentlicht: (2024)
von: Du, Siyi, et al.
Veröffentlicht: (2024)
Generic Knowledge Boosted Pre-training For Remote Sensing Images
von: Huang, Ziyue, et al.
Veröffentlicht: (2024)
von: Huang, Ziyue, et al.
Veröffentlicht: (2024)
Progressive Local Alignment for Medical Multimodal Pre-training
von: Yan, Huimin, et al.
Veröffentlicht: (2025)
von: Yan, Huimin, et al.
Veröffentlicht: (2025)
Enhancing Visual Question Answering with Multimodal LLMs via Chain-of-Question Guided Retrieval-Augmented Generation
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
von: Xu, Quanxing, et al.
Veröffentlicht: (2026)
Research on fusing topological data analysis with convolutional neural network
von: Han, Yang, et al.
Veröffentlicht: (2024)
von: Han, Yang, et al.
Veröffentlicht: (2024)
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning
von: Gong, Xuan, et al.
Veröffentlicht: (2026)
von: Gong, Xuan, et al.
Veröffentlicht: (2026)
MaskHOI: Robust 3D Hand-Object Interaction Estimation via Masked Pre-training
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
von: Xie, Yuechen, et al.
Veröffentlicht: (2025)
Q-DeepSight: Incentivizing Thinking with Images for Image Quality Assessment and Refinement
von: Li, Xudong, et al.
Veröffentlicht: (2026)
von: Li, Xudong, et al.
Veröffentlicht: (2026)
Pre-training a Density-Aware Pose Transformer for Robust LiDAR-based 3D Human Pose Estimation
von: An, Xiaoqi, et al.
Veröffentlicht: (2024)
von: An, Xiaoqi, et al.
Veröffentlicht: (2024)
Multimodal Autoregressive Pre-training of Large Vision Encoders
von: Fini, Enrico, et al.
Veröffentlicht: (2024)
von: Fini, Enrico, et al.
Veröffentlicht: (2024)
Vehicle-centric Perception via Multimodal Structured Pre-training
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
von: Wu, Wentao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024) -
Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
von: Zeng, Weili, et al.
Veröffentlicht: (2025) -
ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
von: Wang, Xiaolong, et al.
Veröffentlicht: (2025) -
Animate-X: Universal Character Image Animation with Enhanced Motion Representation
von: Tan, Shuai, et al.
Veröffentlicht: (2024) -
Mimir: Improving Video Diffusion Models for Precise Text Understanding
von: Tan, Shuai, et al.
Veröffentlicht: (2024)