Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Zhaoyang, Shen, Yantao, Shi, Kunyu, Cai, Zhaowei, Fang, Jun, Deng, Siqi, Yang, Hao, Modolo, Davide, Tu, Zhuowen, Soatto, Stefano |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Non-autoregressive Sequence-to-Sequence Vision-Language Models
di: Shi, Kunyu, et al.
Pubblicazione: (2024)
di: Shi, Kunyu, et al.
Pubblicazione: (2024)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
di: Tan, Jing, et al.
Pubblicazione: (2026)
di: Tan, Jing, et al.
Pubblicazione: (2026)
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
di: Zhao, Yue, et al.
Pubblicazione: (2022)
di: Zhao, Yue, et al.
Pubblicazione: (2022)
Open-World Dynamic Prompt and Continual Visual Representation Learning
di: Kim, Youngeun, et al.
Pubblicazione: (2024)
di: Kim, Youngeun, et al.
Pubblicazione: (2024)
Enhancing Vision-Language Pre-training with Rich Supervisions
di: Gao, Yuan, et al.
Pubblicazione: (2024)
di: Gao, Yuan, et al.
Pubblicazione: (2024)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
di: Becker, Evan, et al.
Pubblicazione: (2024)
di: Becker, Evan, et al.
Pubblicazione: (2024)
Mixed-Query Transformer: A Unified Image Segmentation Architecture
di: Wang, Pei, et al.
Pubblicazione: (2024)
di: Wang, Pei, et al.
Pubblicazione: (2024)
Visual Reasoning through Tool-supervised Reinforcement Learning
di: Dong, Qihua, et al.
Pubblicazione: (2026)
di: Dong, Qihua, et al.
Pubblicazione: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2026)
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2026)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
di: Kim, Sungnyun, et al.
Pubblicazione: (2024)
di: Kim, Sungnyun, et al.
Pubblicazione: (2024)
ViTGAN: Training GANs with Vision Transformers
di: Lee, Kwonjoon, et al.
Pubblicazione: (2021)
di: Lee, Kwonjoon, et al.
Pubblicazione: (2021)
Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability
di: Chang, Tyler A., et al.
Pubblicazione: (2023)
di: Chang, Tyler A., et al.
Pubblicazione: (2023)
Hyperbolic Learning with Synthetic Captions for Open-World Detection
di: Kong, Fanjie, et al.
Pubblicazione: (2024)
di: Kong, Fanjie, et al.
Pubblicazione: (2024)
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale
di: Gueuwou, Shester, et al.
Pubblicazione: (2024)
di: Gueuwou, Shester, et al.
Pubblicazione: (2024)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
di: Zancato, Luca, et al.
Pubblicazione: (2024)
di: Zancato, Luca, et al.
Pubblicazione: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Asymmetric Actor-Critic for Multi-turn LLM Agents
di: Jiang, Shuli, et al.
Pubblicazione: (2026)
di: Jiang, Shuli, et al.
Pubblicazione: (2026)
Descriminative-Generative Custom Tokens for Vision-Language Models
di: Perera, Pramuditha, et al.
Pubblicazione: (2025)
di: Perera, Pramuditha, et al.
Pubblicazione: (2025)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
di: Li, Xiaolong, et al.
Pubblicazione: (2024)
di: Li, Xiaolong, et al.
Pubblicazione: (2024)
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
di: Chen, Zeyuan, et al.
Pubblicazione: (2025)
di: Chen, Zeyuan, et al.
Pubblicazione: (2025)
Training Data Protection with Compositional Diffusion Models
di: Golatkar, Aditya, et al.
Pubblicazione: (2023)
di: Golatkar, Aditya, et al.
Pubblicazione: (2023)
All for One, and One for All: UrbanSyn Dataset, the third Musketeer of Synthetic Driving Scenes
di: Gómez, Jose L., et al.
Pubblicazione: (2023)
di: Gómez, Jose L., et al.
Pubblicazione: (2023)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
di: Trager, Matthew, et al.
Pubblicazione: (2023)
di: Trager, Matthew, et al.
Pubblicazione: (2023)
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
di: Liu, Xiaoze, et al.
Pubblicazione: (2026)
di: Liu, Xiaoze, et al.
Pubblicazione: (2026)
Revisiting Prompt Pretraining of Vision-Language Models
di: Chen, Zhenyuan, et al.
Pubblicazione: (2024)
di: Chen, Zhenyuan, et al.
Pubblicazione: (2024)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
di: Ram, Dhananjay, et al.
Pubblicazione: (2025)
MultiGPrompt for Multi-Task Pre-Training and Prompting on Graphs
di: Yu, Xingtong, et al.
Pubblicazione: (2023)
di: Yu, Xingtong, et al.
Pubblicazione: (2023)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
di: Zhang, Jipeng, et al.
Pubblicazione: (2025)
di: Zhang, Jipeng, et al.
Pubblicazione: (2025)
Joint Post-Training Quantization of Vision Transformers with Learned Prompt-Guided Data Generation
di: Li, Shile, et al.
Pubblicazione: (2026)
di: Li, Shile, et al.
Pubblicazione: (2026)
Learning Compact Video Representations for Efficient Long-form Video Understanding in Large Multimodal Models
di: Chen, Yuxiao, et al.
Pubblicazione: (2026)
di: Chen, Yuxiao, et al.
Pubblicazione: (2026)
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models
di: Shan, Xiaojun, et al.
Pubblicazione: (2026)
di: Shan, Xiaojun, et al.
Pubblicazione: (2026)
Cascade Prompt Learning for Vision-Language Model Adaptation
di: Wu, Ge, et al.
Pubblicazione: (2024)
di: Wu, Ge, et al.
Pubblicazione: (2024)
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
di: Tao, Chaofan, et al.
Pubblicazione: (2024)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
di: Xu, Zhenlin, et al.
Pubblicazione: (2023)
AffordanceLLM: Grounding Affordance from Vision Language Models
di: Qian, Shengyi, et al.
Pubblicazione: (2024)
di: Qian, Shengyi, et al.
Pubblicazione: (2024)
Autoregressive Image Generation with Vision Full-view Prompt
di: Cai, Miaomiao, et al.
Pubblicazione: (2025)
di: Cai, Miaomiao, et al.
Pubblicazione: (2025)
Sensivity of LLMs' Explanations to the Training Randomness:Context, Class & Task Dependencies
di: Loncour, Romain, et al.
Pubblicazione: (2026)
di: Loncour, Romain, et al.
Pubblicazione: (2026)
Diffeomorphic Template Registration for Atmospheric Turbulence Mitigation
di: Lao, Dong, et al.
Pubblicazione: (2024)
di: Lao, Dong, et al.
Pubblicazione: (2024)
Restoration by Generation with Constrained Priors
di: Ding, Zheng, et al.
Pubblicazione: (2023)
di: Ding, Zheng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Non-autoregressive Sequence-to-Sequence Vision-Language Models
di: Shi, Kunyu, et al.
Pubblicazione: (2024) -
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
di: Tan, Jing, et al.
Pubblicazione: (2026) -
ELODI: Ensemble Logit Difference Inhibition for Positive-Congruent Training
di: Zhao, Yue, et al.
Pubblicazione: (2022) -
Open-World Dynamic Prompt and Continual Visual Representation Learning
di: Kim, Youngeun, et al.
Pubblicazione: (2024) -
Enhancing Vision-Language Pre-training with Rich Supervisions
di: Gao, Yuan, et al.
Pubblicazione: (2024)