Non-autoregressive Sequence-to-Sequence Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Kunyu, Dong, Qi, Goncalves, Luis, Tu, Zhuowen, Soatto, Stefano |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2023)
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2023)
Enhancing Vision-Language Pre-training with Rich Supervisions
di: Gao, Yuan, et al.
Pubblicazione: (2024)
di: Gao, Yuan, et al.
Pubblicazione: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
di: Li, Xiaolong, et al.
Pubblicazione: (2024)
di: Li, Xiaolong, et al.
Pubblicazione: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
di: Trager, Matthew, et al.
Pubblicazione: (2023)
di: Trager, Matthew, et al.
Pubblicazione: (2023)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
di: Kaul, Prannay, et al.
Pubblicazione: (2024)
di: Kaul, Prannay, et al.
Pubblicazione: (2024)
Sub-token ViT Embedding via Stochastic Resonance Transformers
di: Lao, Dong, et al.
Pubblicazione: (2023)
di: Lao, Dong, et al.
Pubblicazione: (2023)
DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
di: Fu, Jia, et al.
Pubblicazione: (2025)
di: Fu, Jia, et al.
Pubblicazione: (2025)
Divided Attention: Unsupervised Multi-Object Discovery with Contextually Separated Slots
di: Lao, Dong, et al.
Pubblicazione: (2023)
di: Lao, Dong, et al.
Pubblicazione: (2023)
NeRF-Insert: 3D Local Editing with Multimodal Control Signals
di: Sabat, Benet Oriol, et al.
Pubblicazione: (2024)
di: Sabat, Benet Oriol, et al.
Pubblicazione: (2024)
Sharingan: Extract User Action Sequence from Desktop Recordings
di: Chen, Yanting, et al.
Pubblicazione: (2024)
di: Chen, Yanting, et al.
Pubblicazione: (2024)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
di: Segu, Mattia, et al.
Pubblicazione: (2024)
di: Segu, Mattia, et al.
Pubblicazione: (2024)
DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video Generation
di: Cheng, Hanbo, et al.
Pubblicazione: (2024)
di: Cheng, Hanbo, et al.
Pubblicazione: (2024)
Flow caching for autoregressive video generation
di: Ma, Yuexiao, et al.
Pubblicazione: (2026)
di: Ma, Yuexiao, et al.
Pubblicazione: (2026)
RadarSeq: A Temporal Vision Framework for User Churn Prediction via Radar Chart Sequences
di: Najafi, Sina, et al.
Pubblicazione: (2025)
di: Najafi, Sina, et al.
Pubblicazione: (2025)
Predicting Road Crossing Behaviour using Pose Detection and Sequence Modelling
di: Dasgupta, Subhasis, et al.
Pubblicazione: (2025)
di: Dasgupta, Subhasis, et al.
Pubblicazione: (2025)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
di: Yin, Xiangyu, et al.
Pubblicazione: (2025)
di: Yin, Xiangyu, et al.
Pubblicazione: (2025)
SeqTex: Generate Mesh Textures in Video Sequence
di: Yuan, Ze, et al.
Pubblicazione: (2025)
di: Yuan, Ze, et al.
Pubblicazione: (2025)
Extending Information Bottleneck Attribution to Video Sequences
di: Solopova, Veronika, et al.
Pubblicazione: (2025)
di: Solopova, Veronika, et al.
Pubblicazione: (2025)
Semantic Segmentation of Video Sequences with Convolutional LSTMs
di: Pfeuffer, Andreas, et al.
Pubblicazione: (2019)
di: Pfeuffer, Andreas, et al.
Pubblicazione: (2019)
Training Data Protection with Compositional Diffusion Models
di: Golatkar, Aditya, et al.
Pubblicazione: (2023)
di: Golatkar, Aditya, et al.
Pubblicazione: (2023)
Mode-as-Sequence: Translating Multimodal Motion Prediction into Unified Sequential Mode Modeling
di: Zhou, Zikang, et al.
Pubblicazione: (2026)
di: Zhou, Zikang, et al.
Pubblicazione: (2026)
VHELM: A Holistic Evaluation of Vision Language Models
di: Lee, Tony, et al.
Pubblicazione: (2024)
di: Lee, Tony, et al.
Pubblicazione: (2024)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
di: Hu, Rui, et al.
Pubblicazione: (2024)
di: Hu, Rui, et al.
Pubblicazione: (2024)
FILA: Fine-Grained Vision Language Models
di: Zhu, Shiding, et al.
Pubblicazione: (2024)
di: Zhu, Shiding, et al.
Pubblicazione: (2024)
KAN Text to Vision? The Exploration of Kolmogorov-Arnold Networks for Multi-Scale Sequence-Based Pose Animation from Sign Language Notation
di: Du, Guanyi, et al.
Pubblicazione: (2026)
di: Du, Guanyi, et al.
Pubblicazione: (2026)
Monocular Normal Estimation via Shading Sequence Estimation
di: Li, Zongrui, et al.
Pubblicazione: (2026)
di: Li, Zongrui, et al.
Pubblicazione: (2026)
Hands-On: Segmenting Individual Signs from Continuous Sequences
di: Low, JianHe, et al.
Pubblicazione: (2025)
di: Low, JianHe, et al.
Pubblicazione: (2025)
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
di: Xue, Yuan, et al.
Pubblicazione: (2024)
di: Xue, Yuan, et al.
Pubblicazione: (2024)
Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles
di: Lao, Dong, et al.
Pubblicazione: (2025)
di: Lao, Dong, et al.
Pubblicazione: (2025)
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
di: Tian, Xinyu, et al.
Pubblicazione: (2025)
di: Tian, Xinyu, et al.
Pubblicazione: (2025)
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
di: Ogezi, Michael, et al.
Pubblicazione: (2025)
di: Ogezi, Michael, et al.
Pubblicazione: (2025)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
di: Tao, Xijia, et al.
Pubblicazione: (2024)
di: Tao, Xijia, et al.
Pubblicazione: (2024)
GSPN-2: Efficient Parallel Sequence Modeling
di: Wang, Hongjun, et al.
Pubblicazione: (2025)
di: Wang, Hongjun, et al.
Pubblicazione: (2025)
ProcessPainter: Learn Painting Process from Sequence Data
di: Song, Yiren, et al.
Pubblicazione: (2024)
di: Song, Yiren, et al.
Pubblicazione: (2024)
AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
di: Li, Jieyu, et al.
Pubblicazione: (2025)
di: Li, Jieyu, et al.
Pubblicazione: (2025)
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
di: Wei, Hongyang, et al.
Pubblicazione: (2026)
Detection of Intoxicated Individuals from Facial Video Sequences via a Recurrent Fusion Model
di: Baroutian, Bita, et al.
Pubblicazione: (2025)
di: Baroutian, Bita, et al.
Pubblicazione: (2025)
Integrating Sequence and Image Modeling in Irregular Medical Time Series Through Self-Supervised Learning
di: Chen, Liuqing, et al.
Pubblicazione: (2025)
di: Chen, Liuqing, et al.
Pubblicazione: (2025)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
di: Tian, Kexin, et al.
Pubblicazione: (2025)
di: Tian, Kexin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
di: Zhang, Zhaoyang, et al.
Pubblicazione: (2023) -
Enhancing Vision-Language Pre-training with Rich Supervisions
di: Gao, Yuan, et al.
Pubblicazione: (2024) -
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
di: Li, Xiaolong, et al.
Pubblicazione: (2024) -
On the Scalability of Diffusion-based Text-to-Image Generation
di: Li, Hao, et al.
Pubblicazione: (2024) -
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
di: Trager, Matthew, et al.
Pubblicazione: (2023)