SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Ziang, Rittscher, Jens, Ali, Sharib |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
von: Ancarani, Elisa, et al.
Veröffentlicht: (2025)
von: Ancarani, Elisa, et al.
Veröffentlicht: (2025)
Lester: rotoscope animation through video object segmentation and tracking
von: Tous, Ruben
Veröffentlicht: (2024)
von: Tous, Ruben
Veröffentlicht: (2024)
Consistent and Invariant Generalization Learning for Short-video Misinformation Detection
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
von: Guo, Hanghui, et al.
Veröffentlicht: (2025)
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024)
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
von: Zhu, Rui, et al.
Veröffentlicht: (2024)
Does SpatioTemporal information benefit Two video summarization benchmarks?
von: Ganesh, Aashutosh, et al.
Veröffentlicht: (2024)
von: Ganesh, Aashutosh, et al.
Veröffentlicht: (2024)
Reviewing Intelligent Cinematography: AI research for camera-based video production
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2024)
von: Azzarelli, Adrian, et al.
Veröffentlicht: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
von: Zhang, Bolin, et al.
Veröffentlicht: (2026)
Semi-supervised Chinese Poem-to-Painting Generation via Cycle-consistent Adversarial Networks
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
von: Lu, Zhengyang, et al.
Veröffentlicht: (2024)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
von: Zhang, Wang, et al.
Veröffentlicht: (2024)
von: Zhang, Wang, et al.
Veröffentlicht: (2024)
Subjective evaluation of UHD video coded using VVC with LCEVC and ML-VVC
von: Ramzan, Naeem, et al.
Veröffentlicht: (2026)
von: Ramzan, Naeem, et al.
Veröffentlicht: (2026)
A multi-center analysis of deep learning methods for video polyp detection and segmentation
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
von: Ghatwary, Noha, et al.
Veröffentlicht: (2026)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
von: Li, Deng, et al.
Veröffentlicht: (2024)
von: Li, Deng, et al.
Veröffentlicht: (2024)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
von: Ma, Xueqi, et al.
Veröffentlicht: (2025)
von: Ma, Xueqi, et al.
Veröffentlicht: (2025)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
Leveraging Automatic Personalised Nutrition: Food Image Recognition Benchmark and Dataset based on Nutrition Taxonomy
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
Bayesian uncertainty-weighted loss for improved generalisability on polyp segmentation task
von: Stone, Rebecca S., et al.
Veröffentlicht: (2023)
von: Stone, Rebecca S., et al.
Veröffentlicht: (2023)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
von: Wu, Junyan, et al.
Veröffentlicht: (2025)
Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery
von: Xu, Yulin, et al.
Veröffentlicht: (2026)
von: Xu, Yulin, et al.
Veröffentlicht: (2026)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Deep learning for 3D human pose estimation and mesh recovery: A survey
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
SVFAP: Self-supervised Video Facial Affect Perceiver
von: Sun, Licai, et al.
Veröffentlicht: (2023)
von: Sun, Licai, et al.
Veröffentlicht: (2023)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
Test-time adaptation for image compression with distribution regularization
von: Chen, Kecheng, et al.
Veröffentlicht: (2024)
von: Chen, Kecheng, et al.
Veröffentlicht: (2024)
Reversing the Damage: A QP-Aware Transformer-Diffusion Approach for 8K Video Restoration under Codec Compression
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2024)
von: Dehaghi, Ali Mollaahmadi, et al.
Veröffentlicht: (2024)
Audio-visual training for improved grounding in video-text LLMs
von: Sagare, Shivprasad, et al.
Veröffentlicht: (2024)
von: Sagare, Shivprasad, et al.
Veröffentlicht: (2024)
AKiRa: Augmentation Kit on Rays for optical video generation
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
MVPbev: Multi-view Perspective Image Generation from BEV with Test-time Controllability and Generalizability
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
von: Liu, Buyu, et al.
Veröffentlicht: (2024)
AI-Driven Virtual Teacher for Enhanced Educational Efficiency: Leveraging Large Pretrain Models for Autonomous Error Analysis and Correction
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
von: Xu, Tianlong, et al.
Veröffentlicht: (2024)
ProMamba: Prompt-Mamba for polyp segmentation
von: Xie, Jianhao, et al.
Veröffentlicht: (2024)
von: Xie, Jianhao, et al.
Veröffentlicht: (2024)
T$^\text{3}$SVFND: Towards an Evolving Fake News Detector for Emergencies with Test-time Training on Short Video Platforms
von: Zhang, Liyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Liyuan, et al.
Veröffentlicht: (2025)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
Towards Holistic Language-video Representation: the language model-enhanced MSR-Video to Text Dataset
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
von: Yang, Yuchen, et al.
Veröffentlicht: (2024)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
von: Zhao, Haochen, et al.
Veröffentlicht: (2025)
von: Zhao, Haochen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
von: Ancarani, Elisa, et al.
Veröffentlicht: (2025) -
Lester: rotoscope animation through video object segmentation and tracking
von: Tous, Ruben
Veröffentlicht: (2024) -
Consistent and Invariant Generalization Learning for Short-video Misinformation Detection
von: Guo, Hanghui, et al.
Veröffentlicht: (2025) -
Self-supervised Photographic Image Layout Representation Learning
von: Zhao, Zhaoran, et al.
Veröffentlicht: (2024) -
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
von: Yin, Jianjian, et al.
Veröffentlicht: (2025)