DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yoon, Junho, Jung, Jaemo, Kim, Hyunju, Lee, Dongman |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
by: Park, Junho, et al.
Published: (2025)
by: Park, Junho, et al.
Published: (2025)
Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
by: Xie, Haoyu, et al.
Published: (2025)
by: Xie, Haoyu, et al.
Published: (2025)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
by: Park, Konyul, et al.
Published: (2024)
by: Park, Konyul, et al.
Published: (2024)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025)
by: Hyun, Jeongseok, et al.
Published: (2025)
LACO: Adaptive Latent Communication for Collaborative Driving
by: Chen, Tianhao, et al.
Published: (2026)
by: Chen, Tianhao, et al.
Published: (2026)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
STR-Match: Matching SpatioTemporal Relevance Score for Training-Free Video Editing
by: Lee, Junsung, et al.
Published: (2025)
by: Lee, Junsung, et al.
Published: (2025)
Feature Augmentation based Test-Time Adaptation
by: Cho, Younggeol, et al.
Published: (2024)
by: Cho, Younggeol, et al.
Published: (2024)
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
by: Park, Jinho, et al.
Published: (2026)
by: Park, Jinho, et al.
Published: (2026)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
FPANet: Frequency-based Video Demoireing using Frame-level Post Alignment
by: Oh, Gyeongrok, et al.
Published: (2023)
by: Oh, Gyeongrok, et al.
Published: (2023)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs
by: Li, Baiqi, et al.
Published: (2026)
by: Li, Baiqi, et al.
Published: (2026)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding
by: Li, Yuan-Ming, et al.
Published: (2024)
by: Li, Yuan-Ming, et al.
Published: (2024)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
by: Lee, Hosu, et al.
Published: (2024)
by: Lee, Hosu, et al.
Published: (2024)
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
by: Zhang, Runze, et al.
Published: (2025)
by: Zhang, Runze, et al.
Published: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
Deformable Dynamic Convolution for Accurate yet Efficient Spatio-Temporal Traffic Prediction
by: Jin, Hyeonseok, et al.
Published: (2025)
by: Jin, Hyeonseok, et al.
Published: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
by: Xu, Qi'ao, et al.
Published: (2025)
by: Xu, Qi'ao, et al.
Published: (2025)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
by: Li, Lingen, et al.
Published: (2026)
by: Li, Lingen, et al.
Published: (2026)
VG3T: Visual Geometry Grounded Gaussian Transformer
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Learning Spatio-Temporal Feature Representations for Video-Based Gaze Estimation
by: Personnic, Alexandre, et al.
Published: (2025)
by: Personnic, Alexandre, et al.
Published: (2025)
Human-Centric Video Anomaly Detection Through Spatio-Temporal Pose Tokenization and Transformer
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
by: Cao, Tri, et al.
Published: (2026)
by: Cao, Tri, et al.
Published: (2026)
SAM-PM: Enhancing Video Camouflaged Object Detection using Spatio-Temporal Attention
by: Meeran, Muhammad Nawfal, et al.
Published: (2024)
by: Meeran, Muhammad Nawfal, et al.
Published: (2024)
CVA: Context-aware Video-text Alignment for Video Temporal Grounding
by: Moon, Sungho, et al.
Published: (2026)
by: Moon, Sungho, et al.
Published: (2026)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
by: Meng, Jiahao, et al.
Published: (2025)
by: Meng, Jiahao, et al.
Published: (2025)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
STREAM: Spatio-TempoRal Evaluation and Analysis Metric for Video Generative Models
by: Kim, Pum Jun, et al.
Published: (2024)
by: Kim, Pum Jun, et al.
Published: (2024)
ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
by: Lee, Hosu, et al.
Published: (2025)
by: Lee, Hosu, et al.
Published: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
LiteFat: Lightweight Spatio-Temporal Graph Learning for Real-Time Driver Fatigue Detection
by: Ren, Jing, et al.
Published: (2025)
by: Ren, Jing, et al.
Published: (2025)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
by: Wang, Chenhao, et al.
Published: (2026)
by: Wang, Chenhao, et al.
Published: (2026)
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
by: Kang, Hyunju, et al.
Published: (2026)
by: Kang, Hyunju, et al.
Published: (2026)
STAA: Spatio-Temporal Attention Attribution for Real-Time Interpreting Transformer-based Video Models
by: Wang, Zerui, et al.
Published: (2024)
by: Wang, Zerui, et al.
Published: (2024)
Similar Items
-
EgoWorld: Translating Exocentric View to Egocentric View using Rich Exocentric Observations
by: Park, Junho, et al.
Published: (2025) -
Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
by: Xie, Haoyu, et al.
Published: (2025) -
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
by: Park, Konyul, et al.
Published: (2024) -
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
by: Hyun, Jeongseok, et al.
Published: (2025) -
LACO: Adaptive Latent Communication for Collaborative Driving
by: Chen, Tianhao, et al.
Published: (2026)