JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Seok Hwan, Son, Taein, Seo, Soo Won, Kim, Jisong, Choi, Jun Won |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
di: Son, Taein, et al.
Pubblicazione: (2024)
di: Son, Taein, et al.
Pubblicazione: (2024)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
di: Seo, Soo Won, et al.
Pubblicazione: (2026)
di: Seo, Soo Won, et al.
Pubblicazione: (2026)
CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
di: Kim, Jisong, et al.
Pubblicazione: (2024)
di: Kim, Jisong, et al.
Pubblicazione: (2024)
RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
di: Bang, Geonho, et al.
Pubblicazione: (2024)
di: Bang, Geonho, et al.
Pubblicazione: (2024)
RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection
di: Kim, Jisong, et al.
Pubblicazione: (2023)
di: Kim, Jisong, et al.
Pubblicazione: (2023)
MR-Occ: Efficient Camera-LiDAR 3D Semantic Occupancy Prediction Using Hierarchical Multi-Resolution Voxel Representation
di: Seong, Minjae, et al.
Pubblicazione: (2024)
di: Seong, Minjae, et al.
Pubblicazione: (2024)
MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
di: Kang, Changwon, et al.
Pubblicazione: (2025)
di: Kang, Changwon, et al.
Pubblicazione: (2025)
RS-Net: Context-Aware Relation Scoring for Dynamic Scene Graph Generation
di: Jo, Hae-Won, et al.
Pubblicazione: (2025)
di: Jo, Hae-Won, et al.
Pubblicazione: (2025)
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
di: Jang, Won Shik, et al.
Pubblicazione: (2026)
di: Jang, Won Shik, et al.
Pubblicazione: (2026)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
di: Lee, SuBeen, et al.
Pubblicazione: (2025)
RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
di: Bang, Geonho, et al.
Pubblicazione: (2025)
di: Bang, Geonho, et al.
Pubblicazione: (2025)
PillarGen: Enhancing Radar Point Cloud Density and Quality via Pillar-based Point Generation Network
di: Kim, Jisong, et al.
Pubblicazione: (2024)
di: Kim, Jisong, et al.
Pubblicazione: (2024)
Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes
di: Choi, JunYong, et al.
Pubblicazione: (2025)
di: Choi, JunYong, et al.
Pubblicazione: (2025)
ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query Decoder
di: Kim, Jungho, et al.
Pubblicazione: (2024)
di: Kim, Jungho, et al.
Pubblicazione: (2024)
Mask2Map: Vectorized HD Map Construction Using Bird's Eye View Segmentation Masks
di: Choi, Sehwan, et al.
Pubblicazione: (2024)
di: Choi, Sehwan, et al.
Pubblicazione: (2024)
Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly Detection
di: Chen, Chenglizhao, et al.
Pubblicazione: (2024)
di: Chen, Chenglizhao, et al.
Pubblicazione: (2024)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
di: Bae, Kyungho, et al.
Pubblicazione: (2023)
di: Bae, Kyungho, et al.
Pubblicazione: (2023)
SARA: Scene-Aware Reconstruction Accelerator
di: Lee, Jee Won, et al.
Pubblicazione: (2026)
di: Lee, Jee Won, et al.
Pubblicazione: (2026)
Semi-Supervised Domain Adaptation Using Target-Oriented Domain Augmentation for 3D Object Detection
di: Kim, Yecheol, et al.
Pubblicazione: (2024)
di: Kim, Yecheol, et al.
Pubblicazione: (2024)
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
di: Bae, Kyungho, et al.
Pubblicazione: (2025)
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity
di: Seo, SeungWon, et al.
Pubblicazione: (2024)
di: Seo, SeungWon, et al.
Pubblicazione: (2024)
MATT-GS: Masked Attention-based 3DGS for Robot Perception and Object Detection
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
di: Seong, Hyun Seok, et al.
Pubblicazione: (2026)
di: Seong, Hyun Seok, et al.
Pubblicazione: (2026)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
di: Moon, WonJun, et al.
Pubblicazione: (2026)
di: Moon, WonJun, et al.
Pubblicazione: (2026)
CAVIS: Context-Aware Video Instance Segmentation
di: Lee, Seunghun, et al.
Pubblicazione: (2024)
di: Lee, Seunghun, et al.
Pubblicazione: (2024)
OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception
di: Koh, Junho, et al.
Pubblicazione: (2025)
di: Koh, Junho, et al.
Pubblicazione: (2025)
MAIR++: Improving Multi-view Attention Inverse Rendering with Implicit Lighting Representation
di: Choi, JunYong, et al.
Pubblicazione: (2024)
di: Choi, JunYong, et al.
Pubblicazione: (2024)
Resilient Sensor Fusion under Adverse Sensor Failures via Multi-Modal Expert Fusion
di: Park, Konyul, et al.
Pubblicazione: (2025)
di: Park, Konyul, et al.
Pubblicazione: (2025)
TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection
di: Yoon, Jiae, et al.
Pubblicazione: (2026)
di: Yoon, Jiae, et al.
Pubblicazione: (2026)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
di: Kang, Ju Yeon, et al.
Pubblicazione: (2025)
Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
di: Son, Chang-Hwan, et al.
Pubblicazione: (2025)
di: Son, Chang-Hwan, et al.
Pubblicazione: (2025)
LiteVoxel: Low-memory Intelligent Thresholding for Efficient Voxel Rasterization
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
di: Lee, Jee Won, et al.
Pubblicazione: (2025)
SNeRV: Spectra-preserving Neural Representation for Video
di: Kim, Jina, et al.
Pubblicazione: (2025)
di: Kim, Jina, et al.
Pubblicazione: (2025)
Derivative-Free Diffusion Manifold-Constrained Gradient for Unified XAI
di: Kim, Won Jun, et al.
Pubblicazione: (2024)
di: Kim, Won Jun, et al.
Pubblicazione: (2024)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
di: Zhou, Zhenghong, et al.
Pubblicazione: (2026)
di: Zhou, Zhenghong, et al.
Pubblicazione: (2026)
PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
di: Seo, Wonyong, et al.
Pubblicazione: (2026)
di: Seo, Wonyong, et al.
Pubblicazione: (2026)
DBN-Mix: Training Dual Branch Network Using Bilateral Mixup Augmentation for Long-Tailed Visual Recognition
di: Baik, Jae Soon, et al.
Pubblicazione: (2022)
di: Baik, Jae Soon, et al.
Pubblicazione: (2022)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
di: Park, Konyul, et al.
Pubblicazione: (2024)
di: Park, Konyul, et al.
Pubblicazione: (2024)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
di: Won, John, et al.
Pubblicazione: (2025)
di: Won, John, et al.
Pubblicazione: (2025)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
di: Jeon, Yerim, et al.
Pubblicazione: (2025)
di: Jeon, Yerim, et al.
Pubblicazione: (2025)
Documenti analoghi
-
JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
di: Son, Taein, et al.
Pubblicazione: (2024) -
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
di: Seo, Soo Won, et al.
Pubblicazione: (2026) -
CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
di: Kim, Jisong, et al.
Pubblicazione: (2024) -
RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
di: Bang, Geonho, et al.
Pubblicazione: (2024) -
RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection
di: Kim, Jisong, et al.
Pubblicazione: (2023)