Described Spatial-Temporal Video Detection
Fuente:
arXiv
Guardado en:
| Autores principales: | Ji, Wei, Liu, Xiangyan, Sun, Yingfei, Deng, Jiajun, Qin, You, Nuwanna, Ammar, Qiu, Mengyao, Wei, Lina, Zimmermann, Roger |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Panoptic Scene Graph Generation with Semantics-Prototype Learning
por: Li, Li, et al.
Publicado: (2023)
por: Li, Li, et al.
Publicado: (2023)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
por: Qin, You, et al.
Publicado: (2024)
por: Qin, You, et al.
Publicado: (2024)
TAIL: Text-Audio Incremental Learning
por: Sun, Yingfei, et al.
Publicado: (2025)
por: Sun, Yingfei, et al.
Publicado: (2025)
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
por: Qiu, Yicheng, et al.
Publicado: (2026)
por: Qiu, Yicheng, et al.
Publicado: (2026)
MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance
por: Guo, Jialong, et al.
Publicado: (2025)
por: Guo, Jialong, et al.
Publicado: (2025)
VideoLoom: A Video Large Language Model for Joint Spatial-Temporal Understanding
por: Shi, Jiapeng, et al.
Publicado: (2026)
por: Shi, Jiapeng, et al.
Publicado: (2026)
StreamSTGS: Streaming Spatial and Temporal Gaussian Grids for Real-Time Free-Viewpoint Video
por: Ke, Zhihui, et al.
Publicado: (2025)
por: Ke, Zhihui, et al.
Publicado: (2025)
M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly Detection
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
STNMamba: Mamba-based Spatial-Temporal Normality Learning for Video Anomaly Detection
por: Li, Zhangxun, et al.
Publicado: (2024)
por: Li, Zhangxun, et al.
Publicado: (2024)
GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping under Flexible Language Instructions
por: Chu, Xiaomeng, et al.
Publicado: (2025)
por: Chu, Xiaomeng, et al.
Publicado: (2025)
Video Forgery Detection with Optical Flow Residuals and Spatial-Temporal Consistency
por: Xue, Xi, et al.
Publicado: (2025)
por: Xue, Xi, et al.
Publicado: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
por: Liang, Yiming, et al.
Publicado: (2026)
por: Liang, Yiming, et al.
Publicado: (2026)
MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
por: Ying, Xinru, et al.
Publicado: (2025)
por: Ying, Xinru, et al.
Publicado: (2025)
Autoregressive Video Autoencoder with Decoupled Temporal and Spatial Context
por: Shen, Cuifeng, et al.
Publicado: (2025)
por: Shen, Cuifeng, et al.
Publicado: (2025)
Enhance Multi-Scale Spatial-Temporal Coherence for Configurable Video Anomaly Detection
por: Cheng, Kai, et al.
Publicado: (2023)
por: Cheng, Kai, et al.
Publicado: (2023)
TSdetector: Temporal-Spatial Self-correction Collaborative Learning for Colonoscopy Video Detection
por: Wang, Kaini, et al.
Publicado: (2024)
por: Wang, Kaini, et al.
Publicado: (2024)
Frequency Perception Network for Camouflaged Object Detection
por: Cong, Runmin, et al.
Publicado: (2023)
por: Cong, Runmin, et al.
Publicado: (2023)
SpatialMe: Stereo Video Conversion Using Depth-Warping and Blend-Inpainting
por: Zhang, Jiale, et al.
Publicado: (2024)
por: Zhang, Jiale, et al.
Publicado: (2024)
DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving
por: Tu, Jiahang, et al.
Publicado: (2024)
por: Tu, Jiahang, et al.
Publicado: (2024)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
por: Guo, Chaohong, et al.
Publicado: (2026)
por: Guo, Chaohong, et al.
Publicado: (2026)
SOGDet: Semantic-Occupancy Guided Multi-view 3D Object Detection
por: Zhou, Qiu, et al.
Publicado: (2023)
por: Zhou, Qiu, et al.
Publicado: (2023)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Spatial-Temporal Human-Object Interaction Detection
por: Sun, Xu, et al.
Publicado: (2025)
por: Sun, Xu, et al.
Publicado: (2025)
FRESCO: Spatial-Temporal Correspondence for Zero-Shot Video Translation
por: Yang, Shuai, et al.
Publicado: (2024)
por: Yang, Shuai, et al.
Publicado: (2024)
Hybrid Architecture for Real-Time Video Anomaly Detection: Integrating Spatial and Temporal Analysis
por: Poirier, Fabien
Publicado: (2024)
por: Poirier, Fabien
Publicado: (2024)
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
por: Huang, Shaofei, et al.
Publicado: (2024)
por: Huang, Shaofei, et al.
Publicado: (2024)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
por: Wang, Feng, et al.
Publicado: (2026)
por: Wang, Feng, et al.
Publicado: (2026)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
por: Liu, Zichen, et al.
Publicado: (2025)
por: Liu, Zichen, et al.
Publicado: (2025)
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
por: Hu, Lianyu, et al.
Publicado: (2024)
por: Hu, Lianyu, et al.
Publicado: (2024)
Mobius: A High Efficient Spatial-Temporal Parallel Training Paradigm for Text-to-Video Generation Task
por: Yang, Yiran, et al.
Publicado: (2024)
por: Yang, Yiran, et al.
Publicado: (2024)
CoSTL: Comprehensive Spatial-Temporal Representation Learning for Moment Retrieval and Highlight Detection
por: Dong, Xin, et al.
Publicado: (2026)
por: Dong, Xin, et al.
Publicado: (2026)
Small Object Detection Model with Spatial Laplacian Pyramid Attention and Multi-Scale Features Enhancement in Aerial Images
por: Ji, Zhangjian, et al.
Publicado: (2026)
por: Ji, Zhangjian, et al.
Publicado: (2026)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
por: Liang, Lili, et al.
Publicado: (2024)
por: Liang, Lili, et al.
Publicado: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
por: Xiong, Yuanhao, et al.
Publicado: (2023)
por: Xiong, Yuanhao, et al.
Publicado: (2023)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
por: Chen, Zhikai, et al.
Publicado: (2024)
por: Chen, Zhikai, et al.
Publicado: (2024)
STeInFormer: Spatial-Temporal Interaction Transformer Architecture for Remote Sensing Change Detection
por: Ma, Xiaowen, et al.
Publicado: (2024)
por: Ma, Xiaowen, et al.
Publicado: (2024)
Zero-Shot Video Translation and Editing with Frame Spatial-Temporal Correspondence
por: Yang, Shuai, et al.
Publicado: (2025)
por: Yang, Shuai, et al.
Publicado: (2025)
Described Object Detection: Liberating Object Detection with Flexible Expressions
por: Xie, Chi, et al.
Publicado: (2023)
por: Xie, Chi, et al.
Publicado: (2023)
Describe Anything: Detailed Localized Image and Video Captioning
por: Lian, Long, et al.
Publicado: (2025)
por: Lian, Long, et al.
Publicado: (2025)
Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
por: Hayun, Omer Ben, et al.
Publicado: (2026)
por: Hayun, Omer Ben, et al.
Publicado: (2026)
Ejemplares similares
-
Panoptic Scene Graph Generation with Semantics-Prototype Learning
por: Li, Li, et al.
Publicado: (2023) -
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
por: Qin, You, et al.
Publicado: (2024) -
TAIL: Text-Audio Incremental Learning
por: Sun, Yingfei, et al.
Publicado: (2025) -
Efficient Spatial-Temporal Focal Adapter with SSM for Temporal Action Detection
por: Qiu, Yicheng, et al.
Publicado: (2026) -
MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal Guidance
por: Guo, Jialong, et al.
Publicado: (2025)