STMixer: A One-Stage Sparse Action Detector
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Tao, Cao, Mengqi, Gao, Ziteng, Wu, Gangshan, Wang, Limin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)
by: Cui, Yutao, et al.
Published: (2023)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023)
by: Zhang, Jiaming, et al.
Published: (2023)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
StageInteractor: Query-based Object Detector with Cross-stage Interaction
by: Teng, Yao, et al.
Published: (2023)
by: Teng, Yao, et al.
Published: (2023)
GLAD: Generative Language-Assisted Visual Tracking for Low-Semantic Templates
by: Luo, Xingyu, et al.
Published: (2026)
by: Luo, Xingyu, et al.
Published: (2026)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
by: Zhao, Zhiyu, et al.
Published: (2023)
by: Zhao, Zhiyu, et al.
Published: (2023)
PixNerd: Pixel Neural Field Diffusion
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Efficient Test-Time Prompt Tuning for Vision-Language Models
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
Pure-Pass: Fine-Grained, Adaptive Masking for Dynamic Token-Mixing Routing in Lightweight Image Super-Resolution
by: Wu, Junyu, et al.
Published: (2025)
by: Wu, Junyu, et al.
Published: (2025)
Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
by: Wu, Qirui, et al.
Published: (2025)
by: Wu, Qirui, et al.
Published: (2025)
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
by: Chao, Yuhao, et al.
Published: (2025)
by: Chao, Yuhao, et al.
Published: (2025)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
by: Xu, Boyue, et al.
Published: (2026)
by: Xu, Boyue, et al.
Published: (2026)
Adapting Short-Term Transformers for Action Detection in Untrimmed Videos
by: Yang, Min, et al.
Published: (2023)
by: Yang, Min, et al.
Published: (2023)
RGB-D Video Object Segmentation via Enhanced Multi-store Feature Memory
by: Xu, Boyue, et al.
Published: (2025)
by: Xu, Boyue, et al.
Published: (2025)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
by: Liu, Xin, et al.
Published: (2025)
by: Liu, Xin, et al.
Published: (2025)
Bootstrapping SparseFormers from Vision Foundation Models
by: Gao, Ziteng, et al.
Published: (2023)
by: Gao, Ziteng, et al.
Published: (2023)
SteadyDancer: Harmonized and Coherent Human Image Animation with First-Frame Preservation
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Learning Frequency and Memory-Aware Prompts for Multi-Modal Object Tracking
by: Xu, Boyue, et al.
Published: (2025)
by: Xu, Boyue, et al.
Published: (2025)
HyPSAM: Hybrid Prompt-driven Segment Anything Model for RGB-Thermal Salient Object Detection
by: Hou, Ruichao, et al.
Published: (2025)
by: Hou, Ruichao, et al.
Published: (2025)
SwiTrack: Tri-State Switch for Cross-Modal Object Tracking
by: Xu, Boyue, et al.
Published: (2025)
by: Xu, Boyue, et al.
Published: (2025)
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
by: Yang, Jiange, et al.
Published: (2025)
by: Yang, Jiange, et al.
Published: (2025)
Deep Learning-Based Object Detection for Autonomous Vehicles: A Comparative Study of One-Stage and Two-Stage Detectors on Basic Traffic Objects
by: Karbouj, Bsher, et al.
Published: (2026)
by: Karbouj, Bsher, et al.
Published: (2026)
SAM 2++: Tracking Anything at Any Granularity
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
MTNet: Learning modality-aware representation with transformer for RGBT tracking
by: Hou, Ruichao, et al.
Published: (2025)
by: Hou, Ruichao, et al.
Published: (2025)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
by: Li, Xingyuan, et al.
Published: (2025)
by: Li, Xingyuan, et al.
Published: (2025)
Spatial-Temporal Human-Object Interaction Detection
by: Sun, Xu, et al.
Published: (2025)
by: Sun, Xu, et al.
Published: (2025)
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
A Large-Scale Study on Video Action Dataset Condensation
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
D-AR: Diffusion via Autoregressive Models
by: Gao, Ziteng, et al.
Published: (2025)
by: Gao, Ziteng, et al.
Published: (2025)
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
Sketch and Refine: Towards Fast and Accurate Lane Detection
by: Chen, Chao, et al.
Published: (2024)
by: Chen, Chao, et al.
Published: (2024)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
by: Gao, Ruopeng, et al.
Published: (2023)
by: Gao, Ruopeng, et al.
Published: (2023)
Sparse Global Matching for Video Frame Interpolation with Large Motion
by: Liu, Chunxu, et al.
Published: (2024)
by: Liu, Chunxu, et al.
Published: (2024)
EMO-X: Efficient Multi-Person Pose and Shape Estimation in One-Stage
by: Jian, Haohang, et al.
Published: (2025)
by: Jian, Haohang, et al.
Published: (2025)
Similar Items
-
Open-Vocabulary Spatio-Temporal Action Detection
by: Wu, Tao, et al.
Published: (2024) -
SportsHHI: A Dataset for Human-Human Interaction Detection in Sports Videos
by: Wu, Tao, et al.
Published: (2024) -
Dual DETRs for Multi-Label Temporal Action Detection
by: Zhu, Yuhan, et al.
Published: (2024) -
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023) -
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
by: Zhang, Jiaming, et al.
Published: (2023)