Joint Self-Supervised Video Alignment and Action Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Ali, Ali Shah, Mahmood, Syed Ahmed, Saeed, Mubin, Konin, Andrey, Zia, M. Zeeshan, Tran, Quoc-Huy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
by: Tran, Quoc-Huy, et al.
Published: (2023)
by: Tran, Quoc-Huy, et al.
Published: (2023)
Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
by: Hyder, Syed Waleed, et al.
Published: (2023)
by: Hyder, Syed Waleed, et al.
Published: (2023)
Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
by: Tran, Quoc-Huy, et al.
Published: (2023)
by: Tran, Quoc-Huy, et al.
Published: (2023)
Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
by: Mahmood, Syed Ahmed, et al.
Published: (2025)
by: Mahmood, Syed Ahmed, et al.
Published: (2025)
Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
by: Ahmed, Umer, et al.
Published: (2026)
by: Ahmed, Umer, et al.
Published: (2026)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
by: Fateh, Fawad Javed, et al.
Published: (2026)
by: Fateh, Fawad Javed, et al.
Published: (2026)
Vim4Path: Self-Supervised Vision Mamba for Histopathology Images
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
Pseudo-label Refinement for Improving Self-Supervised Learning Systems
by: Zia-ur-Rehman, et al.
Published: (2024)
by: Zia-ur-Rehman, et al.
Published: (2024)
Language Model Guided Interpretable Video Action Reasoning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
Leveraging counterfactual concepts for debugging and improving CNN model performance
by: Tariq, Syed Ali, et al.
Published: (2025)
by: Tariq, Syed Ali, et al.
Published: (2025)
Self-Supervised Ultrasound-Video Segmentation with Feature Prediction and 3D Localised Loss
by: Ellis, Edward, et al.
Published: (2025)
by: Ellis, Edward, et al.
Published: (2025)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
Split-Fuse-Transport: Annotation-Free Saliency via Dual Clustering and Optimal Transport Alignment
by: Ramzan, Muhammad Umer, et al.
Published: (2025)
by: Ramzan, Muhammad Umer, et al.
Published: (2025)
Self-Supervised Alignment Learning for Medical Image Segmentation
by: Li, Haofeng, et al.
Published: (2024)
by: Li, Haofeng, et al.
Published: (2024)
Semi-Supervised Semantic Segmentation using Redesigned Self-Training for White Blood Cells
by: Luu, Vinh Quoc, et al.
Published: (2024)
by: Luu, Vinh Quoc, et al.
Published: (2024)
Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment
by: Xu, Angchi, et al.
Published: (2024)
by: Xu, Angchi, et al.
Published: (2024)
Towards Counterfactual and Contrastive Explainability and Transparency of DCNN Image Classifiers
by: Tariq, Syed Ali, et al.
Published: (2025)
by: Tariq, Syed Ali, et al.
Published: (2025)
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
by: Trinh, Quoc-Huy
Published: (2025)
by: Trinh, Quoc-Huy
Published: (2025)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
by: Trinh, Quoc-Huy, et al.
Published: (2024)
by: Trinh, Quoc-Huy, et al.
Published: (2024)
A Tutorial on ALOS2 SAR Utilization: Dataset Preparation, Self-Supervised Pretraining, and Semantic Segmentation
by: Imamoglu, Nevrez, et al.
Published: (2026)
by: Imamoglu, Nevrez, et al.
Published: (2026)
KDAS: Knowledge Distillation via Attention Supervision Framework for Polyp Segmentation
by: Trinh, Quoc-Huy, et al.
Published: (2023)
by: Trinh, Quoc-Huy, et al.
Published: (2023)
Reshoot-Anything: A Self-Supervised Model for In-the-Wild Video Reshooting
by: Paliwal, Avinash, et al.
Published: (2026)
by: Paliwal, Avinash, et al.
Published: (2026)
Self-Supervised Contrastive Learning for Videos using Differentiable Local Alignment
by: Oei, Keyne, et al.
Published: (2024)
by: Oei, Keyne, et al.
Published: (2024)
ShapeMatcher: Self-Supervised Joint Shape Canonicalization, Segmentation, Retrieval and Deformation
by: Di, Yan, et al.
Published: (2023)
by: Di, Yan, et al.
Published: (2023)
Hierarchical Action Learning for Weakly-Supervised Action Segmentation
by: Huang, Junxian, et al.
Published: (2026)
by: Huang, Junxian, et al.
Published: (2026)
A comprehensive overview of deep learning models for object detection from videos/images
by: Zulfqar, Sukana, et al.
Published: (2026)
by: Zulfqar, Sukana, et al.
Published: (2026)
Supervised and Contrastive Self-Supervised In-Domain Representation Learning for Dense Prediction Problems in Remote Sensing
by: Ghanbarzade, Ali, et al.
Published: (2023)
by: Ghanbarzade, Ali, et al.
Published: (2023)
Faithful Counterfactual Visual Explanations (FCVE)
by: Khan, Bismillah, et al.
Published: (2025)
by: Khan, Bismillah, et al.
Published: (2025)
Dual-Stream Alignment for Action Segmentation
by: Gammulle, Harshala, et al.
Published: (2025)
by: Gammulle, Harshala, et al.
Published: (2025)
Lightweight Transformer Framework for Weakly Supervised Semantic Segmentation
by: Torabi, Ali, et al.
Published: (2025)
by: Torabi, Ali, et al.
Published: (2025)
SelfFed: Self-Supervised Federated Learning for Data Heterogeneity and Label Scarcity in Medical Images
by: Khowaja, Sunder Ali, et al.
Published: (2023)
by: Khowaja, Sunder Ali, et al.
Published: (2023)
EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition
by: Abdelkawy, Ahmed, et al.
Published: (2024)
by: Abdelkawy, Ahmed, et al.
Published: (2024)
Component-Aware Sketch-to-Image Generation Using Self-Attention Encoding and Coordinate-Preserving Fusion
by: Zia, Ali, et al.
Published: (2026)
by: Zia, Ali, et al.
Published: (2026)
An Ensemble Learning Approach towards Waste Segmentation in Cluttered Environment
by: Jafar, Maimoona, et al.
Published: (2026)
by: Jafar, Maimoona, et al.
Published: (2026)
Resource Efficient Multi-stain Kidney Glomeruli Segmentation via Self-supervision
by: Nisar, Zeeshan, et al.
Published: (2024)
by: Nisar, Zeeshan, et al.
Published: (2024)
Test-Time Adaptation for Anomaly Segmentation via Topology-Aware Optimal Transport Chaining
by: Zia, Ali, et al.
Published: (2026)
by: Zia, Ali, et al.
Published: (2026)
Pose-Aware Weakly-Supervised Action Segmentation
by: Zhao, Seth Z., et al.
Published: (2025)
by: Zhao, Seth Z., et al.
Published: (2025)
Similar Items
-
Permutation-Aware Action Segmentation via Unsupervised Frame-to-Segment Alignment
by: Tran, Quoc-Huy, et al.
Published: (2023) -
Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion
by: Hyder, Syed Waleed, et al.
Published: (2023) -
Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
by: Tran, Quoc-Huy, et al.
Published: (2023) -
Procedure Learning via Regularized Gromov-Wasserstein Optimal Transport
by: Mahmood, Syed Ahmed, et al.
Published: (2025) -
Unsupervised Skeleton-Based Action Segmentation via Hierarchical Spatiotemporal Vector Quantization
by: Ahmed, Umer, et al.
Published: (2026)