Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Junbin, Tsai, Pei-Hsuan, Tsai, Meng-Hsun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
von: Kazemi, Amir, et al.
Veröffentlicht: (2024)
von: Kazemi, Amir, et al.
Veröffentlicht: (2024)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
von: Chereddy, Sai V R
Veröffentlicht: (2026)
von: Chereddy, Sai V R
Veröffentlicht: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
von: Korolkov, Vasilii
Veröffentlicht: (2025)
von: Korolkov, Vasilii
Veröffentlicht: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Jin, Xiaofeng, et al.
Veröffentlicht: (2025)
Efficient and Privacy-Protecting Background Removal for 2D Video Streaming using iPhone 15 Pro Max LiDAR
von: Kinnevan, Jessica, et al.
Veröffentlicht: (2025)
von: Kinnevan, Jessica, et al.
Veröffentlicht: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S.
Veröffentlicht: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
Meaning over Motion: A Semantic-First Approach to 360° Viewport Prediction
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
IMKD: Intensity-Aware Multi-Level Knowledge Distillation for Camera-Radar Fusion
von: Mishra, Shashank, et al.
Veröffentlicht: (2025)
von: Mishra, Shashank, et al.
Veröffentlicht: (2025)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
von: Mohammad, Noor Islam S., et al.
Veröffentlicht: (2025)
Motion Attribution for Video Generation
von: Wu, Xindi, et al.
Veröffentlicht: (2026)
von: Wu, Xindi, et al.
Veröffentlicht: (2026)
Transforming faces into video stories -- VideoFace2.0
von: Brkljač, Branko, et al.
Veröffentlicht: (2025)
von: Brkljač, Branko, et al.
Veröffentlicht: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Gr-IoU: Ground-Intersection over Union for Robust Multi-Object Tracking with 3D Geometric Constraints
von: Toida, Keisuke, et al.
Veröffentlicht: (2024)
von: Toida, Keisuke, et al.
Veröffentlicht: (2024)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
AUTHENTICATION: Identifying Rare Failure Modes in Autonomous Vehicle Perception Systems using Adversarially Guided Diffusion Models
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
von: Zarei, Mohammad, et al.
Veröffentlicht: (2025)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
von: Dong, Zixuan, et al.
Veröffentlicht: (2025)
von: Dong, Zixuan, et al.
Veröffentlicht: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
von: Baydemir, Poyraz
Veröffentlicht: (2025)
von: Baydemir, Poyraz
Veröffentlicht: (2025)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
von: Lu, Hongyu, et al.
Veröffentlicht: (2026)
von: Lu, Hongyu, et al.
Veröffentlicht: (2026)
Corn Ear Detection and Orientation Estimation Using Deep Learning
von: Sprague, Nathan, et al.
Veröffentlicht: (2024)
von: Sprague, Nathan, et al.
Veröffentlicht: (2024)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
Image-based Facial Rig Inversion
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
von: Zhang, Junbin, et al.
Veröffentlicht: (2026) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025) -
AIDOVECL: AI-generated Dataset of Outpainted Vehicles for Eye-level Classification and Localization
von: Kazemi, Amir, et al.
Veröffentlicht: (2024) -
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
von: Kang, Xueyang, et al.
Veröffentlicht: (2026) -
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
von: Chereddy, Sai V R
Veröffentlicht: (2026)