Query matching for spatio-temporal action detection with query-based object detector
Fuente:
arXiv
Salvato in:
| Autori principali: | Hori, Shimon, Omi, Kazuki, Tamaki, Toru |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Action tube generation by person query matching for spatio-temporal action detection
di: Omi, Kazuki, et al.
Pubblicazione: (2025)
di: Omi, Kazuki, et al.
Pubblicazione: (2025)
Shift and matching queries for video semantic segmentation
di: Mizuno, Tsubasa, et al.
Pubblicazione: (2024)
di: Mizuno, Tsubasa, et al.
Pubblicazione: (2024)
Multi-model learning by sequential reading of untrimmed videos for action recognition
di: Kamiya, Kodai, et al.
Pubblicazione: (2024)
di: Kamiya, Kodai, et al.
Pubblicazione: (2024)
Can masking background and object reduce static bias for zero-shot action recognition?
di: Fukuzawa, Takumi, et al.
Pubblicazione: (2025)
di: Fukuzawa, Takumi, et al.
Pubblicazione: (2025)
Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video
di: Liao, Guiqiu, et al.
Pubblicazione: (2024)
di: Liao, Guiqiu, et al.
Pubblicazione: (2024)
MoExDA: Domain Adaptation for Edge-based Action Recognition
di: Sugimoto, Takuya, et al.
Pubblicazione: (2025)
di: Sugimoto, Takuya, et al.
Pubblicazione: (2025)
Reflective Dialogue between Teacher and Solver Agents for Video Question Answering
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
Online pre-training with long-form videos
di: Kato, Itsuki, et al.
Pubblicazione: (2024)
di: Kato, Itsuki, et al.
Pubblicazione: (2024)
Fine-grained length controllable video captioning with ordinal embeddings
di: Nitta, Tomoya, et al.
Pubblicazione: (2024)
di: Nitta, Tomoya, et al.
Pubblicazione: (2024)
BFMD: A Full-Match Badminton Dense Dataset for Dense Shot Captioning
di: Ding, Ning, et al.
Pubblicazione: (2026)
di: Ding, Ning, et al.
Pubblicazione: (2026)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
di: Ding, Ning, et al.
Pubblicazione: (2025)
di: Ding, Ning, et al.
Pubblicazione: (2025)
Disentangling Static and Dynamic Information for Reducing Static Bias in Action Recognition
di: Kobayashi, Masato, et al.
Pubblicazione: (2025)
di: Kobayashi, Masato, et al.
Pubblicazione: (2025)
ST-Gait++: Leveraging spatio-temporal convolutions for gait-based emotion recognition on videos
di: Lima, Maria Luísa, et al.
Pubblicazione: (2024)
di: Lima, Maria Luísa, et al.
Pubblicazione: (2024)
TadML: A fast temporal action detection with Mechanics-MLP
di: Deng, Bowen, et al.
Pubblicazione: (2022)
di: Deng, Bowen, et al.
Pubblicazione: (2022)
Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
di: Takama, Yusaku, et al.
Pubblicazione: (2025)
di: Takama, Yusaku, et al.
Pubblicazione: (2025)
M3DDM+: An improved video outpainting by a modified masking strategy
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
di: Murakawa, Takuya, et al.
Pubblicazione: (2026)
Detection of moving objects through turbulent media. Decomposition of Oscillatory vs Non-Oscillatory spatio-temporal vector fields
di: Gilles, Jerome, et al.
Pubblicazione: (2024)
di: Gilles, Jerome, et al.
Pubblicazione: (2024)
Two-stream joint matching method based on contrastive learning for few-shot action recognition
di: Deng, Long, et al.
Pubblicazione: (2024)
di: Deng, Long, et al.
Pubblicazione: (2024)
Covariant spatio-temporal receptive fields for spiking neural networks
di: Pedersen, Jens Egholm, et al.
Pubblicazione: (2024)
di: Pedersen, Jens Egholm, et al.
Pubblicazione: (2024)
Rethinking temporal self-similarity for repetitive action counting
di: Luo, Yanan, et al.
Pubblicazione: (2024)
di: Luo, Yanan, et al.
Pubblicazione: (2024)
Extended multi-stream temporal-attention module for skeleton-based human action recognition (HAR)
di: Mehmood, Faisal, et al.
Pubblicazione: (2024)
di: Mehmood, Faisal, et al.
Pubblicazione: (2024)
QdaVPR: A novel query-based domain-agnostic model for visual place recognition
di: Wan, Shanshan, et al.
Pubblicazione: (2026)
di: Wan, Shanshan, et al.
Pubblicazione: (2026)
Local2Global query Alignment for Video Instance Segmentation
di: Koner, Rajat, et al.
Pubblicazione: (2025)
di: Koner, Rajat, et al.
Pubblicazione: (2025)
UKDM: Underwater keypoint detection and matching using underwater image enhancement techniques
di: Diaz-Garcia, Pedro, et al.
Pubblicazione: (2025)
di: Diaz-Garcia, Pedro, et al.
Pubblicazione: (2025)
Computer Vision based group activity detection and action spotting
di: Sivalingam, Narthana, et al.
Pubblicazione: (2025)
di: Sivalingam, Narthana, et al.
Pubblicazione: (2025)
Large-scale unsupervised spatio-temporal semantic analysis of vast regions from satellite images sequences
di: Echegoyen, Carlos, et al.
Pubblicazione: (2022)
di: Echegoyen, Carlos, et al.
Pubblicazione: (2022)
Improving Video Question Answering through query-based frame selection
di: Patil, Himanshu, et al.
Pubblicazione: (2026)
di: Patil, Himanshu, et al.
Pubblicazione: (2026)
Symmetrical Joint Learning Support-query Prototypes for Few-shot Segmentation
di: Li, Qun, et al.
Pubblicazione: (2024)
di: Li, Qun, et al.
Pubblicazione: (2024)
Station2Radar: query conditioned gaussian splatting for precipitation field
di: Kim, Doyi, et al.
Pubblicazione: (2026)
di: Kim, Doyi, et al.
Pubblicazione: (2026)
INFANiTE: Implicit Neural representation for high-resolution Fetal brain spatio-temporal Atlas learNing from clinical Thick-slicE MRI
di: Hu, Xiaotian, et al.
Pubblicazione: (2026)
di: Hu, Xiaotian, et al.
Pubblicazione: (2026)
Single-pixel 3D imaging based on fusion temporal data of single photon detector and millimeter-wave radar
di: Lai, Tingqin, et al.
Pubblicazione: (2023)
di: Lai, Tingqin, et al.
Pubblicazione: (2023)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
di: Li, Kexin, et al.
Pubblicazione: (2024)
di: Li, Kexin, et al.
Pubblicazione: (2024)
MSTAR: Box-free Multi-query Scene Text Retrieval with Attention Recycling
di: Yin, Liang, et al.
Pubblicazione: (2025)
di: Yin, Liang, et al.
Pubblicazione: (2025)
Interlaced dynamic XCT reconstruction with spatio-temporal implicit neural representations
di: Boulanger, Mathias, et al.
Pubblicazione: (2025)
di: Boulanger, Mathias, et al.
Pubblicazione: (2025)
Talking Points: Describing and Localizing Pixels
di: Rusanovsky, Matan, et al.
Pubblicazione: (2025)
di: Rusanovsky, Matan, et al.
Pubblicazione: (2025)
IPAdapter-Instruct: Resolving Ambiguity in Image-based Conditioning using Instruct Prompts
di: Rowles, Ciara, et al.
Pubblicazione: (2024)
di: Rowles, Ciara, et al.
Pubblicazione: (2024)
OmViD: Omni-supervised active learning for video action detection
di: Rana, Aayush, et al.
Pubblicazione: (2025)
di: Rana, Aayush, et al.
Pubblicazione: (2025)
Few-shot target-driven instance detection based on open-vocabulary object detection models
di: Crulis, Ben, et al.
Pubblicazione: (2024)
di: Crulis, Ben, et al.
Pubblicazione: (2024)
Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval
di: Lan, Bangxiang, et al.
Pubblicazione: (2025)
di: Lan, Bangxiang, et al.
Pubblicazione: (2025)
Unsupervised learning based object detection using Contrastive Learning
di: Kumar, Chandan, et al.
Pubblicazione: (2024)
di: Kumar, Chandan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Action tube generation by person query matching for spatio-temporal action detection
di: Omi, Kazuki, et al.
Pubblicazione: (2025) -
Shift and matching queries for video semantic segmentation
di: Mizuno, Tsubasa, et al.
Pubblicazione: (2024) -
Multi-model learning by sequential reading of untrimmed videos for action recognition
di: Kamiya, Kodai, et al.
Pubblicazione: (2024) -
Can masking background and object reduce static bias for zero-shot action recognition?
di: Fukuzawa, Takumi, et al.
Pubblicazione: (2025) -
Disentangling spatio-temporal knowledge for weakly supervised object detection and segmentation in surgical video
di: Liao, Guiqiu, et al.
Pubblicazione: (2024)