CollideNet: Hierarchical Multi-scale Video Representation Learning with Disentanglement for Time-To-Collision Forecasting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Desai, Nishq Poorav, Etemad, Ali, Greenspan, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2024)
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2024)
Socially-Informed Reconstruction for Pedestrian Trajectory Forecasting
von: Damirchi, Haleh, et al.
Veröffentlicht: (2024)
von: Damirchi, Haleh, et al.
Veröffentlicht: (2024)
Diffusion Models with Deterministic Normalizing Flow Priors
von: Zand, Mohsen, et al.
Veröffentlicht: (2023)
von: Zand, Mohsen, et al.
Veröffentlicht: (2023)
On The Relationship Between Continual Learning and Long-Tailed Recognition
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2023)
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2023)
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2025)
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2025)
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
von: Sarkar, Pritam, et al.
Veröffentlicht: (2025)
Self-Supervised Human Activity Recognition with Localized Time-Frequency Contrastive Representation Learning
von: Taghanaki, Setareh Rahimi, et al.
Veröffentlicht: (2022)
von: Taghanaki, Setareh Rahimi, et al.
Veröffentlicht: (2022)
Pseudo-keypoint RKHS Learning for Self-supervised 6DoF Pose Estimation
von: Wu, Yangzheng, et al.
Veröffentlicht: (2023)
von: Wu, Yangzheng, et al.
Veröffentlicht: (2023)
Consistency-guided Prompt Learning for Vision-Language Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Partial Label Learning for Emotion Recognition from EEG
von: Zhang, Guangyi, et al.
Veröffentlicht: (2023)
von: Zhang, Guangyi, et al.
Veröffentlicht: (2023)
Impact of Strategic Sampling and Supervision Policies on Semi-supervised Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2022)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2022)
A Shared Encoder Approach to Multimodal Representation Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
DLTPose: 6DoF Pose Estimation From Accurate Dense Surface Point Estimates
von: Jadhav, Akash, et al.
Veröffentlicht: (2025)
von: Jadhav, Akash, et al.
Veröffentlicht: (2025)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
SelfPrompt: Confidence-Aware Semi-Supervised Tuning for Robust Vision-Language Model Adaptation
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2025)
Exploring the Boundaries of Semi-Supervised Facial Expression Recognition using In-Distribution, Out-of-Distribution, and Unconstrained Data
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Scaling Up Semi-supervised Learning with Unconstrained Unlabelled Data
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2023)
Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation
von: Yarram, Sudhir, et al.
Veröffentlicht: (2024)
von: Yarram, Sudhir, et al.
Veröffentlicht: (2024)
Unmasking Deepfakes: Masked Autoencoding Spatiotemporal Transformers for Enhanced Video Forgery Detection
von: Das, Sayantan, et al.
Veröffentlicht: (2023)
von: Das, Sayantan, et al.
Veröffentlicht: (2023)
Learning Disentangled Representations for Generalized Multi-view Clustering
von: Zou, Xin, et al.
Veröffentlicht: (2026)
von: Zou, Xin, et al.
Veröffentlicht: (2026)
Stride-Net: Fairness-Aware Disentangled Representation Learning for Chest X-Ray Diagnosis
von: Rashid, Darakshan, et al.
Veröffentlicht: (2026)
von: Rashid, Darakshan, et al.
Veröffentlicht: (2026)
Human Pose Estimation from Ambiguous Pressure Recordings with Spatio-temporal Masked Transformers
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2023)
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2023)
Multistream Gaze Estimation with Anatomical Eye Region Isolation by Synthetic to Real Transfer Learning
von: Mahmud, Zunayed, et al.
Veröffentlicht: (2022)
von: Mahmud, Zunayed, et al.
Veröffentlicht: (2022)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
von: Jin, Peng, et al.
Veröffentlicht: (2024)
von: Jin, Peng, et al.
Veröffentlicht: (2024)
Interpretable Single-View 3D Gaussian Splatting using Unsupervised Hierarchical Disentangled Representation Learning
von: Zhang, Yuyang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyang, et al.
Veröffentlicht: (2025)
TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification
von: Wan, Zishuo, et al.
Veröffentlicht: (2025)
von: Wan, Zishuo, et al.
Veröffentlicht: (2025)
Rethinking Multi-view Representation Learning via Distilled Disentangling
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
von: Ke, Guanzhou, et al.
Veröffentlicht: (2024)
Benchmarking Vision-Language Contrastive Methods for Medical Representation Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
Unsupervised Learning of Disentangled Representations from Video
von: Denton, Remi, et al.
Veröffentlicht: (2017)
von: Denton, Remi, et al.
Veröffentlicht: (2017)
CMSA-Net: Causal Multi-scale Aggregation with Adaptive Multi-source Reference for Video Polyp Segmentation
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
Video Compression with Hierarchical Temporal Neural Representation
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
von: Zhu, Jun, et al.
Veröffentlicht: (2026)
Consistency-Guided Asynchronous Contrastive Tuning for Few-Shot Class-Incremental Tuning of Foundation Models
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
SkelFormer: Markerless 3D Pose and Shape Estimation using Skeletal Transformers
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2024)
von: Davoodnia, Vandad, et al.
Veröffentlicht: (2024)
DINOv2 based Self Supervised Learning For Few Shot Medical Image Segmentation
von: Ayzenberg, Lev, et al.
Veröffentlicht: (2024)
von: Ayzenberg, Lev, et al.
Veröffentlicht: (2024)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
von: Bond, Andrew, et al.
Veröffentlicht: (2025)
von: Bond, Andrew, et al.
Veröffentlicht: (2025)
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
von: Chen, Ye, et al.
Veröffentlicht: (2026)
von: Chen, Ye, et al.
Veröffentlicht: (2026)
A Bag of Tricks for Few-Shot Class-Incremental Learning
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
von: Roy, Shuvendu, et al.
Veröffentlicht: (2024)
ADCD-Net: Robust Document Image Forgery Localization via Adaptive DCT Feature and Hierarchical Content Disentanglement
von: Wong, Kahim, et al.
Veröffentlicht: (2025)
von: Wong, Kahim, et al.
Veröffentlicht: (2025)
Learning Group Actions In Disentangled Latent Image Representations
von: Swarnali, Farhana Hossain, et al.
Veröffentlicht: (2025)
von: Swarnali, Farhana Hossain, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CycleCrash: A Dataset of Bicycle Collision Videos for Collision Prediction and Analysis
von: Desai, Nishq Poorav, et al.
Veröffentlicht: (2024) -
Socially-Informed Reconstruction for Pedestrian Trajectory Forecasting
von: Damirchi, Haleh, et al.
Veröffentlicht: (2024) -
Diffusion Models with Deterministic Normalizing Flow Priors
von: Zand, Mohsen, et al.
Veröffentlicht: (2023) -
On The Relationship Between Continual Learning and Long-Tailed Recognition
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2023) -
PRISM: Reducing Spurious Implicit Biases in Vision-Language Models with LLM-Guided Embedding Projection
von: Molahasani, Mahdiyar, et al.
Veröffentlicht: (2025)