MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Karim, Rezaul, Zhao, He, Wildes, Richard P., Siam, Mennatullah |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025)
by: Cheshmi, Leila, et al.
Published: (2025)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)
by: Kowal, Matthew, et al.
Published: (2022)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
SemiVT-Surge: Semi-Supervised Video Transformer for Surgical Phase Recognition
by: Li, Yiping, et al.
Published: (2025)
by: Li, Yiping, et al.
Published: (2025)
HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition
by: Wu, Qian, et al.
Published: (2024)
by: Wu, Qian, et al.
Published: (2024)
Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and Benchmark
by: Broni-Bediako, Clifford, et al.
Published: (2024)
by: Broni-Bediako, Clifford, et al.
Published: (2024)
Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
by: Karim, A H M Rezaul, et al.
Published: (2025)
by: Karim, A H M Rezaul, et al.
Published: (2025)
T-MPEDNet: Unveiling the Synergy of Transformer-aware Multiscale Progressive Encoder-Decoder Network with Feature Recalibration for Tumor and Liver Segmentation
by: Raghaw, Chandravardhan Singh, et al.
Published: (2025)
by: Raghaw, Chandravardhan Singh, et al.
Published: (2025)
Unified Multimodal Models as Auto-Encoders
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Selective, Interpretable, and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
by: Ilic, Filip, et al.
Published: (2024)
by: Ilic, Filip, et al.
Published: (2024)
HTR-VT: Handwritten Text Recognition with Vision Transformer
by: Li, Yuting, et al.
Published: (2024)
by: Li, Yuting, et al.
Published: (2024)
Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
A Cross-Hierarchical Difference Feature Fusion Network Based on Multiscale Encoder-Decoder for Hyperspectral Change Detection
by: Sheng, Mingshuai, et al.
Published: (2025)
by: Sheng, Mingshuai, et al.
Published: (2025)
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)
by: Parajuli, Kabita, et al.
Published: (2023)
by: Parajuli, Kabita, et al.
Published: (2023)
TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection
by: Yoon, Jiae, et al.
Published: (2026)
by: Yoon, Jiae, et al.
Published: (2026)
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
by: Jia, Zi-Yi, et al.
Published: (2026)
by: Jia, Zi-Yi, et al.
Published: (2026)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
by: Li, Kaining, et al.
Published: (2025)
by: Li, Kaining, et al.
Published: (2025)
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
by: Yang, Zuhao, et al.
Published: (2026)
by: Yang, Zuhao, et al.
Published: (2026)
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
by: Yang, Zuhao, et al.
Published: (2025)
by: Yang, Zuhao, et al.
Published: (2025)
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
by: Adewale, Sikiru, et al.
Published: (2023)
by: Adewale, Sikiru, et al.
Published: (2023)
T-DEED: Temporal-Discriminability Enhancer Encoder-Decoder for Precise Event Spotting in Sports Videos
by: Xarles, Artur, et al.
Published: (2024)
by: Xarles, Artur, et al.
Published: (2024)
Discrete Wavelet Transform as a Facilitator for Expressive Latent Space Representation in Variational Autoencoders in Satellite Imagery
by: Mahara, Arpan, et al.
Published: (2025)
by: Mahara, Arpan, et al.
Published: (2025)
Fuse after Align: Improving Face-Voice Association Learning via Multimodal Encoder
by: Peng, Chong, et al.
Published: (2024)
by: Peng, Chong, et al.
Published: (2024)
Federated Modality-specific Encoders and Partially Personalized Fusion Decoder for Multimodal Brain Tumor Segmentation
by: Liu, Hong, et al.
Published: (2026)
by: Liu, Hong, et al.
Published: (2026)
CLIP-Optimized Multimodal Image Enhancement via ISP-CNN Fusion for Coal Mine IoVT under Uneven Illumination
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
by: Huang, Xiaohu, et al.
Published: (2025)
by: Huang, Xiaohu, et al.
Published: (2025)
Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
by: Liu, Xiaohong, et al.
Published: (2025)
by: Liu, Xiaohong, et al.
Published: (2025)
A Shared Encoder Approach to Multimodal Representation Learning
by: Roy, Shuvendu, et al.
Published: (2025)
by: Roy, Shuvendu, et al.
Published: (2025)
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023)
by: Yuan, Dehao, et al.
Published: (2023)
Unified Map Prior Encoder for Mapping and Planning
by: Zhang, Zongzheng, et al.
Published: (2026)
by: Zhang, Zongzheng, et al.
Published: (2026)
Video In-context Learning: Autoregressive Transformers are Zero-Shot Video Imitators
by: Zhang, Wentao, et al.
Published: (2024)
by: Zhang, Wentao, et al.
Published: (2024)
EDIT: Enhancing Vision Transformers by Mitigating Attention Sink through an Encoder-Decoder Architecture
by: Feng, Wenfeng, et al.
Published: (2025)
by: Feng, Wenfeng, et al.
Published: (2025)
Beyond the Encoder: Joint Encoder-Decoder Contrastive Pre-Training Improves Dense Prediction
by: Quetin, Sébastien, et al.
Published: (2025)
by: Quetin, Sébastien, et al.
Published: (2025)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
Towards a Better Understanding of the Computer Vision Research Community in Africa
by: Omotayo, Abdul-Hakeem, et al.
Published: (2023)
by: Omotayo, Abdul-Hakeem, et al.
Published: (2023)
Similar Items
-
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023) -
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025) -
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025) -
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)