Graph-Based Multimodal and Multi-view Alignment for Keystep Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Romero, Julia Lee, Min, Kyle, Tripathi, Subarna, Karimzadeh, Morteza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Keystep Recognition using Graph Neural Networks
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025)
Improving Keystep Recognition in Ego-Video via Dexterous Focus
von: Chavis, Zachary, et al.
Veröffentlicht: (2025)
von: Chavis, Zachary, et al.
Veröffentlicht: (2025)
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
von: Valdez, Hector A., et al.
Veröffentlicht: (2024)
Contrastive Language Video Time Pre-training
von: Liu, Hengyue, et al.
Veröffentlicht: (2024)
von: Liu, Hengyue, et al.
Veröffentlicht: (2024)
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)
VideoSAGE: Video Summarization with Graph Representation Learning
von: Chaves, Jose M. Rojas, et al.
Veröffentlicht: (2024)
von: Chaves, Jose M. Rojas, et al.
Veröffentlicht: (2024)
EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
von: Rodin, Ivan, et al.
Veröffentlicht: (2025)
TrajPred: Trajectory-Conditioned Joint Embedding Prediction for Surgical Instrument-Tissue Interaction Recognition in Vision-Language Models
von: Cheng, Jiajun, et al.
Veröffentlicht: (2026)
von: Cheng, Jiajun, et al.
Veröffentlicht: (2026)
PALADIN : Robust Neural Fingerprinting for Text-to-Image Diffusion Models
von: L, Murthy, et al.
Veröffentlicht: (2025)
von: L, Murthy, et al.
Veröffentlicht: (2025)
Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
Harnessing Object Grounding for Time-Sensitive Video Understanding
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2025)
SGANet: Semantic and Geometric Alignment for Multimodal Multi-view Anomaly Detection
von: Bai, Letian, et al.
Veröffentlicht: (2026)
von: Bai, Letian, et al.
Veröffentlicht: (2026)
VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
von: Dipta, Shubhashis Roy, et al.
Veröffentlicht: (2025)
Search2Motion: Training-Free Object-Level Motion Control via Attention-Consensus Search
von: Liu, Sainan, et al.
Veröffentlicht: (2026)
von: Liu, Sainan, et al.
Veröffentlicht: (2026)
A Genealogy of Foundation Models in Remote Sensing
von: Lane, Kevin, et al.
Veröffentlicht: (2025)
von: Lane, Kevin, et al.
Veröffentlicht: (2025)
Trunk-branch Contrastive Network with Multi-view Deformable Aggregation for Multi-view Action Recognition
von: Yang, Yingyuan, et al.
Veröffentlicht: (2025)
von: Yang, Yingyuan, et al.
Veröffentlicht: (2025)
Leveraging Foundation Models for Multimodal Graph-Based Action Recognition
von: Ziaeetabar, Fatemeh, et al.
Veröffentlicht: (2025)
von: Ziaeetabar, Fatemeh, et al.
Veröffentlicht: (2025)
ByDeWay: Boost Your multimodal LLM with DEpth prompting in a Training-Free Way
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2025)
A Proxy Consistency Loss for Grounded Fusion of Earth Observation and Location Encoders
von: Wang, Zhongying, et al.
Veröffentlicht: (2026)
von: Wang, Zhongying, et al.
Veröffentlicht: (2026)
Action Selection Learning for Multi-label Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2024)
Adaptive Learning for Multi-view Stereo Reconstruction
von: Min, Qinglu, et al.
Veröffentlicht: (2024)
von: Min, Qinglu, et al.
Veröffentlicht: (2024)
Multi-view Structural Convolution Network for Domain-Invariant Point Cloud Recognition of Autonomous Vehicles
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
von: Kim, Younggun, et al.
Veröffentlicht: (2025)
View-aware Cross-modal Distillation for Multi-view Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
Skarimva: Skeleton-based Action Recognition is a Multi-view Application
von: Bermuth, Daniel, et al.
Veröffentlicht: (2026)
von: Bermuth, Daniel, et al.
Veröffentlicht: (2026)
Multi-view Action Recognition via Directed Gromov-Wasserstein Discrepancy
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2024)
von: Nguyen, Hoang-Quan, et al.
Veröffentlicht: (2024)
Multimodal Prompt Alignment for Facial Expression Recognition
von: Ma, Fuyan, et al.
Veröffentlicht: (2025)
von: Ma, Fuyan, et al.
Veröffentlicht: (2025)
Multi-view Video-Pose Pretraining for Operating Room Surgical Activity Recognition
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
von: Hamoud, Idris, et al.
Veröffentlicht: (2025)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Investigating the Effect of Spatial Context on Multi-Task Sea Ice Segmentation
von: Vahedi, Behzad, et al.
Veröffentlicht: (2025)
von: Vahedi, Behzad, et al.
Veröffentlicht: (2025)
MonoInstance: Enhancing Monocular Priors via Multi-view Instance Alignment for Neural Rendering and Reconstruction
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
von: Mikeštíková, Anna Šárová, et al.
Veröffentlicht: (2025)
von: Mikeštíková, Anna Šárová, et al.
Veröffentlicht: (2025)
Multimodal Graph Representation Learning for Robust Surgical Workflow Recognition with Adversarial Feature Disentanglement
von: Bai, Long, et al.
Veröffentlicht: (2025)
von: Bai, Long, et al.
Veröffentlicht: (2025)
Ice-FMBench: A Foundation Model Benchmark for Sea Ice Type Segmentation
von: Taleghan, Samira Alkaee, et al.
Veröffentlicht: (2025)
von: Taleghan, Samira Alkaee, et al.
Veröffentlicht: (2025)
Multi Teacher Privileged Knowledge Distillation for Multimodal Expression Recognition
von: Aslam, Muhammad Haseeb, et al.
Veröffentlicht: (2024)
von: Aslam, Muhammad Haseeb, et al.
Veröffentlicht: (2024)
SSPA: Split-and-Synthesize Prompting with Gated Alignments for Multi-Label Image Recognition
von: Tan, Hao, et al.
Veröffentlicht: (2024)
von: Tan, Hao, et al.
Veröffentlicht: (2024)
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Trung Thanh, et al.
Veröffentlicht: (2025)
MASA: Motion-aware Masked Autoencoder with Semantic Alignment for Sign Language Recognition
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
Patch as Node: Human-Centric Graph Representation Learning for Multimodal Action Recognition
von: Liang, Zeyu, et al.
Veröffentlicht: (2025)
von: Liang, Zeyu, et al.
Veröffentlicht: (2025)
Multimodal Spatio-temporal Graph Learning for Alignment-free RGBT Video Object Detection
von: Wang, Qishun, et al.
Veröffentlicht: (2025)
von: Wang, Qishun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Keystep Recognition using Graph Neural Networks
von: Romero, Julia Lee, et al.
Veröffentlicht: (2025) -
Improving Keystep Recognition in Ego-Video via Dexterous Focus
von: Chavis, Zachary, et al.
Veröffentlicht: (2025) -
SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video
von: Valdez, Hector A., et al.
Veröffentlicht: (2024) -
Contrastive Language Video Time Pre-training
von: Liu, Hengyue, et al.
Veröffentlicht: (2024) -
Ego-VPA: Egocentric Video Understanding with Parameter-efficient Adaptation
von: Wu, Tz-Ying, et al.
Veröffentlicht: (2024)