Multimodal Fusion and Coherence Modeling for Video Topic Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Hai, Deng, Chong, Zhang, Qinglin, Liu, Jiaqing, Chen, Qian, Wang, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robust Semi-supervised Multimodal Medical Image Segmentation via Cross Modality Collaboration
by: Zhou, Xiaogen, et al.
Published: (2024)
by: Zhou, Xiaogen, et al.
Published: (2024)
Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
MedMimic: Physician-Inspired Multimodal Fusion for Early Diagnosis of Fever of Unknown Origin
by: Chen, Minrui, et al.
Published: (2025)
by: Chen, Minrui, et al.
Published: (2025)
Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
by: He, Lian, et al.
Published: (2025)
by: He, Lian, et al.
Published: (2025)
Towards a Multimodal MRI-Based Foundation Model for Multi-Level Feature Exploration in Segmentation, Molecular Subtyping, and Grading of Glioma
by: Farahani, Somayeh, et al.
Published: (2025)
by: Farahani, Somayeh, et al.
Published: (2025)
Multimodal Medical Image Binding via Shared Text Embeddings
by: Liu, Yunhao, et al.
Published: (2025)
by: Liu, Yunhao, et al.
Published: (2025)
MedSAM2: Segment Anything in 3D Medical Images and Videos
by: Ma, Jun, et al.
Published: (2025)
by: Ma, Jun, et al.
Published: (2025)
Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion
by: Zhou, Meng, et al.
Published: (2024)
by: Zhou, Meng, et al.
Published: (2024)
Brain-Adapter: Enhancing Neurological Disorder Analysis with Adapter-Tuning Multimodal Large Language Models
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
Enhancing Synthetic CT from CBCT via Multimodal Fusion and End-To-End Registration
by: Tschuchnig, Maximilian, et al.
Published: (2025)
by: Tschuchnig, Maximilian, et al.
Published: (2025)
ClinicalFMamba: Advancing Clinical Assessment using Mamba-based Multimodal Neuroimaging Fusion
by: Zhou, Meng, et al.
Published: (2025)
by: Zhou, Meng, et al.
Published: (2025)
PAD: Phase-Amplitude Decoupling Fusion for Multi-Modal Land Cover Classification
by: Zheng, Huiling, et al.
Published: (2025)
by: Zheng, Huiling, et al.
Published: (2025)
InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model
by: Guan, Fengbin, et al.
Published: (2025)
by: Guan, Fengbin, et al.
Published: (2025)
InSight: AI Mobile Screening Tool for Multiple Eye Disease Detection using Multimodal Fusion
by: Raghu, Ananya, et al.
Published: (2025)
by: Raghu, Ananya, et al.
Published: (2025)
TSUBF-Net: Trans-Spatial UNet-like Network with Bi-direction Fusion for Segmentation of Adenoid Hypertrophy in CT
by: Zhou, Rulin, et al.
Published: (2024)
by: Zhou, Rulin, et al.
Published: (2024)
U-DFA: A Unified DINOv2-Unet with Dual Fusion Attention for Multi-Dataset Medical Segmentation
by: Sajjad, Zulkaif, et al.
Published: (2025)
by: Sajjad, Zulkaif, et al.
Published: (2025)
DCD: A Semantic Segmentation Model for Fetal Ultrasound Four-Chamber View
by: Li, Donglian, et al.
Published: (2025)
by: Li, Donglian, et al.
Published: (2025)
Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts
by: Zhuang, Farica, et al.
Published: (2025)
by: Zhuang, Farica, et al.
Published: (2025)
EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence
by: Li, Yahui, et al.
Published: (2026)
by: Li, Yahui, et al.
Published: (2026)
FACE: Few-shot Adapter with Cross-view Fusion for Cross-subject EEG Emotion Recognition
by: Liu, Haiqi, et al.
Published: (2025)
by: Liu, Haiqi, et al.
Published: (2025)
HDR Image Reconstruction using an Unsupervised Fusion Model
by: Nagaswetha, Kumbha
Published: (2025)
by: Nagaswetha, Kumbha
Published: (2025)
Dual-Encoder Transformer-Based Multimodal Learning for Ischemic Stroke Lesion Segmentation Using Diffusion MRI
by: Usman, Muhammad, et al.
Published: (2025)
by: Usman, Muhammad, et al.
Published: (2025)
Stable Diffusion Segmentation for Biomedical Images with Single-step Reverse Process
by: Lin, Tianyu, et al.
Published: (2024)
by: Lin, Tianyu, et al.
Published: (2024)
RobustSAM: Segment Anything Robustly on Degraded Images
by: Chen, Wei-Ting, et al.
Published: (2024)
by: Chen, Wei-Ting, et al.
Published: (2024)
MICCAI STS 2024 Challenge: Semi-Supervised Instance-Level Tooth Segmentation in Panoramic X-ray and CBCT Images
by: Wang, Yaqi, et al.
Published: (2025)
by: Wang, Yaqi, et al.
Published: (2025)
FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
IntelliCardiac: An Intelligent Platform for Cardiac Image Segmentation and Classification
by: Tsai, Ting Yu, et al.
Published: (2025)
by: Tsai, Ting Yu, et al.
Published: (2025)
REHRSeg: Unleashing the Power of Self-Supervised Super-Resolution for Resource-Efficient 3D MRI Segmentation
by: Song, Zhiyun, et al.
Published: (2024)
by: Song, Zhiyun, et al.
Published: (2024)
R2MF-Net: A Recurrent Residual Multi-Path Fusion Network for Robust Multi-directional Spine X-ray Segmentation
by: Li, Xuecheng, et al.
Published: (2025)
by: Li, Xuecheng, et al.
Published: (2025)
A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts
by: Ge, Peixuan, et al.
Published: (2025)
by: Ge, Peixuan, et al.
Published: (2025)
Review and Recommendations for using Artificial Intelligence in Intracoronary Optical Coherence Tomography Analysis
by: Chen, Xu, et al.
Published: (2025)
by: Chen, Xu, et al.
Published: (2025)
Multimodal Fusion at Three Tiers: Physics-Driven Data Generation and Vision-Language Guidance for Brain Tumor Segmentation
by: Zhang, Mingda
Published: (2025)
by: Zhang, Mingda
Published: (2025)
An Empirical Study on the Fairness of Foundation Models for Multi-Organ Image Segmentation
by: Li, Qin, et al.
Published: (2024)
by: Li, Qin, et al.
Published: (2024)
Introducing VaDA: Novel Image Segmentation Model for Maritime Object Segmentation Using New Dataset
by: Kim, Yongjin, et al.
Published: (2024)
by: Kim, Yongjin, et al.
Published: (2024)
Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion
by: Wang, Shengze, et al.
Published: (2024)
by: Wang, Shengze, et al.
Published: (2024)
PulmoFusion: Advancing Pulmonary Health with Efficient Multi-Modal Fusion
by: Sharshar, Ahmed, et al.
Published: (2025)
by: Sharshar, Ahmed, et al.
Published: (2025)
CV-VAE: A Compatible Video VAE for Latent Generative Video Models
by: Zhao, Sijie, et al.
Published: (2024)
by: Zhao, Sijie, et al.
Published: (2024)
Potential of Multimodal Large Language Models for Data Mining of Medical Images and Free-text Reports
by: Zhang, Yutong, et al.
Published: (2024)
by: Zhang, Yutong, et al.
Published: (2024)
Modality-Projection Universal Model for Comprehensive Full-Body Medical Imaging Segmentation
by: Chen, Yixin, et al.
Published: (2024)
by: Chen, Yixin, et al.
Published: (2024)
Similar Items
-
Robust Semi-supervised Multimodal Medical Image Segmentation via Cross Modality Collaboration
by: Zhou, Xiaogen, et al.
Published: (2024) -
Evaluating the Diagnostic Classification Ability of Multimodal Large Language Models: Insights from the Osteoarthritis Initiative
by: Wang, Li, et al.
Published: (2026) -
MedMimic: Physician-Inspired Multimodal Fusion for Early Diagnosis of Fever of Unknown Origin
by: Chen, Minrui, et al.
Published: (2025) -
Task-Aware 3D Affordance Segmentation via 2D Guidance and Geometric Refinement
by: He, Lian, et al.
Published: (2025) -
Towards a Multimodal MRI-Based Foundation Model for Multi-Level Feature Exploration in Segmentation, Molecular Subtyping, and Grading of Glioma
by: Farahani, Somayeh, et al.
Published: (2025)