JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Taein, Seo, Soo Won, Kim, Jisong, Lee, Seok Hwan, Choi, Jun Won |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
by: Lee, Seok Hwan, et al.
Published: (2024)
by: Lee, Seok Hwan, et al.
Published: (2024)
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026)
by: Seo, Soo Won, et al.
Published: (2026)
CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
by: Kim, Jisong, et al.
Published: (2024)
by: Kim, Jisong, et al.
Published: (2024)
RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
by: Bang, Geonho, et al.
Published: (2024)
by: Bang, Geonho, et al.
Published: (2024)
RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection
by: Kim, Jisong, et al.
Published: (2023)
by: Kim, Jisong, et al.
Published: (2023)
MR-Occ: Efficient Camera-LiDAR 3D Semantic Occupancy Prediction Using Hierarchical Multi-Resolution Voxel Representation
by: Seong, Minjae, et al.
Published: (2024)
by: Seong, Minjae, et al.
Published: (2024)
MAESTRO: Task-Relevant Optimization via Adaptive Feature Enhancement and Suppression for Multi-task 3D Perception
by: Kang, Changwon, et al.
Published: (2025)
by: Kang, Changwon, et al.
Published: (2025)
Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
by: Jang, Won Shik, et al.
Published: (2026)
by: Jang, Won Shik, et al.
Published: (2026)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
by: Bang, Geonho, et al.
Published: (2025)
by: Bang, Geonho, et al.
Published: (2025)
PillarGen: Enhancing Radar Point Cloud Density and Quality via Pillar-based Point Generation Network
by: Kim, Jisong, et al.
Published: (2024)
by: Kim, Jisong, et al.
Published: (2024)
ProtoOcc: Accurate, Efficient 3D Occupancy Prediction Using Dual Branch Encoder-Prototype Query Decoder
by: Kim, Jungho, et al.
Published: (2024)
by: Kim, Jungho, et al.
Published: (2024)
Mask2Map: Vectorized HD Map Construction Using Bird's Eye View Segmentation Masks
by: Choi, Sehwan, et al.
Published: (2024)
by: Choi, Sehwan, et al.
Published: (2024)
Semi-Supervised Domain Adaptation Using Target-Oriented Domain Augmentation for 3D Object Detection
by: Kim, Yecheol, et al.
Published: (2024)
by: Kim, Yecheol, et al.
Published: (2024)
AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts
by: Yu, Jun, et al.
Published: (2024)
by: Yu, Jun, et al.
Published: (2024)
MATT-GS: Masked Attention-based 3DGS for Robot Perception and Object Detection
by: Lee, Jee Won, et al.
Published: (2025)
by: Lee, Jee Won, et al.
Published: (2025)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
by: Won, John, et al.
Published: (2025)
by: Won, John, et al.
Published: (2025)
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026)
by: Seong, Hyun Seok, et al.
Published: (2026)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
CAVIS: Context-Aware Video Instance Segmentation
by: Lee, Seunghun, et al.
Published: (2024)
by: Lee, Seunghun, et al.
Published: (2024)
OnlineBEV: Recurrent Temporal Fusion in Bird's Eye View Representations for Multi-Camera 3D Perception
by: Koh, Junho, et al.
Published: (2025)
by: Koh, Junho, et al.
Published: (2025)
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity
by: Seo, SeungWon, et al.
Published: (2024)
by: Seo, SeungWon, et al.
Published: (2024)
MAIR++: Improving Multi-view Attention Inverse Rendering with Implicit Lighting Representation
by: Choi, JunYong, et al.
Published: (2024)
by: Choi, JunYong, et al.
Published: (2024)
Resilient Sensor Fusion under Adverse Sensor Failures via Multi-Modal Expert Fusion
by: Park, Konyul, et al.
Published: (2025)
by: Park, Konyul, et al.
Published: (2025)
It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models
by: Choi, Jaeha, et al.
Published: (2026)
by: Choi, Jaeha, et al.
Published: (2026)
Channel-wise Noise Scheduled Diffusion for Inverse Rendering in Indoor Scenes
by: Choi, JunYong, et al.
Published: (2025)
by: Choi, JunYong, et al.
Published: (2025)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
by: Kang, Ju Yeon, et al.
Published: (2025)
by: Kang, Ju Yeon, et al.
Published: (2025)
Degradation-Agnostic Statistical Facial Feature Transformation for Blind Face Restoration in Adverse Weather Conditions
by: Son, Chang-Hwan, et al.
Published: (2025)
by: Son, Chang-Hwan, et al.
Published: (2025)
LiteVoxel: Low-memory Intelligent Thresholding for Efficient Voxel Rasterization
by: Lee, Jee Won, et al.
Published: (2025)
by: Lee, Jee Won, et al.
Published: (2025)
SNeRV: Spectra-preserving Neural Representation for Video
by: Kim, Jina, et al.
Published: (2025)
by: Kim, Jina, et al.
Published: (2025)
Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
by: Seo, Wonyong, et al.
Published: (2026)
by: Seo, Wonyong, et al.
Published: (2026)
DBN-Mix: Training Dual Branch Network Using Bilateral Mixup Augmentation for Long-Tailed Visual Recognition
by: Baik, Jae Soon, et al.
Published: (2022)
by: Baik, Jae Soon, et al.
Published: (2022)
Fine-Grained Pillar Feature Encoding Via Spatio-Temporal Virtual Grid for 3D Object Detection
by: Park, Konyul, et al.
Published: (2024)
by: Park, Konyul, et al.
Published: (2024)
HAMLET: Switch your Vision-Language-Action Model into a History-Aware Policy
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
RS-Net: Context-Aware Relation Scoring for Dynamic Scene Graph Generation
by: Jo, Hae-Won, et al.
Published: (2025)
by: Jo, Hae-Won, et al.
Published: (2025)
Progressive Proxy Anchor Propagation for Unsupervised Semantic Segmentation
by: Seong, Hyun Seok, et al.
Published: (2024)
by: Seong, Hyun Seok, et al.
Published: (2024)
ProJo4D: Progressive Joint Optimization for Sparse-View Inverse Physics Estimation
by: Rho, Daniel, et al.
Published: (2025)
by: Rho, Daniel, et al.
Published: (2025)
Text Embedding Knows How to Quantize Text-Guided Diffusion Models
by: Lee, Hongjae, et al.
Published: (2025)
by: Lee, Hongjae, et al.
Published: (2025)
Similar Items
-
JARViS: Detecting Actions in Video Using Unified Actor-Scene Context Relation Modeling
by: Lee, Seok Hwan, et al.
Published: (2024) -
Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
by: Seo, Soo Won, et al.
Published: (2026) -
CRT-Fusion: Camera, Radar, Temporal Fusion Using Motion Information for 3D Object Detection
by: Kim, Jisong, et al.
Published: (2024) -
RadarDistill: Boosting Radar-based Object Detection Performance via Knowledge Distillation from LiDAR Features
by: Bang, Geonho, et al.
Published: (2024) -
RCM-Fusion: Radar-Camera Multi-Level Fusion for 3D Object Detection
by: Kim, Jisong, et al.
Published: (2023)