Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Kodathala, Sai Varun, Vunnam, Rakesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SV3.3B: A Sports Video Understanding Model for Action Recognition
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025)
by: Kodathala, Sai Varun, et al.
Published: (2025)
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Fast OTSU Thresholding Using Bisection Method
by: Kodathala, Sai Varun
Published: (2025)
by: Kodathala, Sai Varun
Published: (2025)
Six Sigma For Neural Networks: Taguchi-based optimization
by: Kodathala, Sai Varun
Published: (2025)
by: Kodathala, Sai Varun
Published: (2025)
Can Large Language Models Solve Engineering Equations? A Systematic Comparison of Direct Prediction and Solver-Assisted Approaches
by: Kodathala, Sai Varun, et al.
Published: (2026)
by: Kodathala, Sai Varun, et al.
Published: (2026)
Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking
by: John, Shahla
Published: (2025)
by: John, Shahla
Published: (2025)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
DINOv3 with Test-Time Training for Medical Image Registration
by: Wang, Shansong, et al.
Published: (2025)
by: Wang, Shansong, et al.
Published: (2025)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
by: Assran, Mido, et al.
Published: (2025)
by: Assran, Mido, et al.
Published: (2025)
DINOv3 Meets YOLO26 for Weed Detection in Vegetable Crops
by: Deng, Boyang, et al.
Published: (2026)
by: Deng, Boyang, et al.
Published: (2026)
Randomized-MLP Regularization Improves Domain Adaptation and Interpretability in DINOv2
by: Ortega, Joel Valdivia, et al.
Published: (2025)
by: Ortega, Joel Valdivia, et al.
Published: (2025)
General Purpose Image Encoder DINOv2 for Medical Image Registration
by: Song, Xinrui, et al.
Published: (2024)
by: Song, Xinrui, et al.
Published: (2024)
Social-JEPA: Emergent Geometric Isomorphism
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
DINOv3 as a Frozen Encoder for CRPS-Oriented Probabilistic Rainfall Nowcasting
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
by: Filho, Luciano Araujo Dourado, et al.
Published: (2025)
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
by: Personnic, Alexandre, et al.
Published: (2025)
by: Personnic, Alexandre, et al.
Published: (2025)
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding
by: Luo, Bingjun, et al.
Published: (2026)
by: Luo, Bingjun, et al.
Published: (2026)
Evaluating General Purpose Vision Foundation Models for Medical Image Analysis: An Experimental Study of DINOv2 on Radiology Benchmarks
by: Baharoon, Mohammed, et al.
Published: (2023)
by: Baharoon, Mohammed, et al.
Published: (2023)
Efficient Neural Video Representation with Temporally Coherent Modulation
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
Synergistic Foundation Models for Semi-Supervised Fetal Cardiac Ultrasound Analysis: SAM-Med2D Boundary Refinement and DINOv3 Semantic Enhancement
by: Zhuang, Tonghao, et al.
Published: (2026)
by: Zhuang, Tonghao, et al.
Published: (2026)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition
by: Liu, Feng, et al.
Published: (2025)
by: Liu, Feng, et al.
Published: (2025)
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
by: Liu, Hongli, et al.
Published: (2026)
by: Liu, Hongli, et al.
Published: (2026)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
Revealing the Semantic Selection Gap in DINOv3 through Training-Free Few-Shot Segmentation
by: Zakir, Hussni Mohd, et al.
Published: (2026)
by: Zakir, Hussni Mohd, et al.
Published: (2026)
One-Shot Action Recognition via Multi-Scale Spatial-Temporal Skeleton Matching
by: Yang, Siyuan, et al.
Published: (2023)
by: Yang, Siyuan, et al.
Published: (2023)
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling
by: Yan, Jiebin, et al.
Published: (2025)
by: Yan, Jiebin, et al.
Published: (2025)
Benefits of Feature Extraction and Temporal Sequence Analysis for Video Frame Prediction: An Evaluation of Hybrid Deep Learning Models
by: Velázquez, Jose M. Sánchez, et al.
Published: (2025)
by: Velázquez, Jose M. Sánchez, et al.
Published: (2025)
Align before Adapt: Leveraging Entity-to-Region Alignments for Generalizable Video Action Recognition
by: Chen, Yifei, et al.
Published: (2023)
by: Chen, Yifei, et al.
Published: (2023)
Changes in Gaza: DINOv3-Powered Multi-Class Change Detection for Damage Assessment in Conflict Zones
by: Zheng, Kai, et al.
Published: (2025)
by: Zheng, Kai, et al.
Published: (2025)
Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
by: Yang, Haiyu, et al.
Published: (2026)
by: Yang, Haiyu, et al.
Published: (2026)
SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
by: Tan, Zhentao, et al.
Published: (2024)
by: Tan, Zhentao, et al.
Published: (2024)
MUSTAN: Multi-scale Temporal Context as Attention for Robust Video Foreground Segmentation
by: Pokala, Praveen Kumar, et al.
Published: (2024)
by: Pokala, Praveen Kumar, et al.
Published: (2024)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Spatial-Temporal Deep Embedding for Vehicle Trajectory Reconstruction from High-Angle Video
by: D., Tianya T. Zhang Ph., et al.
Published: (2022)
by: D., Tianya T. Zhang Ph., et al.
Published: (2022)
Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026)
by: Alansari, Mohamad, et al.
Published: (2026)
STARFlow: Spatial Temporal Feature Re-embedding with Attentive Learning for Real-world Scene Flow
by: Lu, Zhiyang, et al.
Published: (2024)
by: Lu, Zhiyang, et al.
Published: (2024)
Mitigating Domain Drift in Multi Species Segmentation with DINOv2: A Cross-Domain Evaluation in Herbicide Research Trials
by: Picon, Artzai, et al.
Published: (2025)
by: Picon, Artzai, et al.
Published: (2025)
Similar Items
-
SV3.3B: A Sports Video Understanding Model for Action Recognition
by: Kodathala, Sai Varun, et al.
Published: (2025) -
The Describe-Then-Generate Bottleneck: How VLM Descriptions Alter Image Generation Outcomes
by: Kodathala, Sai Varun, et al.
Published: (2025) -
LLMs can Compress LLMs: Adaptive Pruning by Agents
by: Kodathala, Sai Varun, et al.
Published: (2026) -
Fast OTSU Thresholding Using Bisection Method
by: Kodathala, Sai Varun
Published: (2025) -
Six Sigma For Neural Networks: Taguchi-based optimization
by: Kodathala, Sai Varun
Published: (2025)