Region-aware Image-based Human Action Retrieval with Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Hongsong, Zhao, Jianhua, Gui, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Point-Supervised Skeleton-Based Human Action Segmentation
by: Wang, Hongsong, et al.
Published: (2026)
by: Wang, Hongsong, et al.
Published: (2026)
Towards Universal Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
Attribution as Retrieval: Model-Agnostic AI-Generated Image Attribution
by: Wang, Hongsong, et al.
Published: (2026)
by: Wang, Hongsong, et al.
Published: (2026)
Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026)
by: Kuang, Jidong, et al.
Published: (2026)
Heterogeneous Skeleton-Based Action Representation Learning
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
OZ-TAL: Online Zero-Shot Temporal Action Localization
by: Han, Chaolei, et al.
Published: (2026)
by: Han, Chaolei, et al.
Published: (2026)
Zero-Shot Skeleton-based Action Recognition with Dual Visual-Text Alignment
by: Kuang, Jidong, et al.
Published: (2024)
by: Kuang, Jidong, et al.
Published: (2024)
Multimodal Skeleton-Based Action Representation Learning via Decomposition and Composition
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
by: Pu, Xinyu, et al.
Published: (2025)
by: Pu, Xinyu, et al.
Published: (2025)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
Efficient Diffusion-Based 3D Human Pose Estimation with Hierarchical Temporal Pruning
by: Bi, Yuquan, et al.
Published: (2025)
by: Bi, Yuquan, et al.
Published: (2025)
LOTA: Bit-Planes Guided AI-Generated Image Detection
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Data-Free Class-Incremental Gesture Recognition with Prototype-Guided Pseudo Feature Replay
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
Coordinate-Based Dual-Constrained Autoregressive Motion Generation
by: Ding, Kang, et al.
Published: (2026)
by: Ding, Kang, et al.
Published: (2026)
Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar Reconstruction
by: Su, Yuze, et al.
Published: (2026)
by: Su, Yuze, et al.
Published: (2026)
Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly Detection
by: Wang, Hongsong, et al.
Published: (2024)
by: Wang, Hongsong, et al.
Published: (2024)
Not All Agents Matter: From Global Attention Dilution to Risk-Prioritized Game Planning
by: Ding, Kang, et al.
Published: (2026)
by: Ding, Kang, et al.
Published: (2026)
Foundation Model for Skeleton-Based Human Action Understanding
by: Wang, Hongsong, et al.
Published: (2025)
by: Wang, Hongsong, et al.
Published: (2025)
FreqMixFormerV2: Lightweight Frequency-aware Mixed Transformer for Human Skeleton Action Recognition
by: Wu, Wenhan, et al.
Published: (2024)
by: Wu, Wenhan, et al.
Published: (2024)
LIPT: Latency-aware Image Processing Transformer
by: Qiao, Junbo, et al.
Published: (2024)
by: Qiao, Junbo, et al.
Published: (2024)
LORTSAR: Low-Rank Transformer for Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2024)
by: Oraki, Soroush, et al.
Published: (2024)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
Learning A Physical-aware Diffusion Model Based on Transformer for Underwater Image Enhancement
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
Efficient Temporal Action Segmentation via Boundary-aware Query Voting
by: Wang, Peiyao, et al.
Published: (2024)
by: Wang, Peiyao, et al.
Published: (2024)
GTPT: Group-based Token Pruning Transformer for Efficient Human Pose Estimation
by: Wang, Haonan, et al.
Published: (2024)
by: Wang, Haonan, et al.
Published: (2024)
Exploiting Regional Information Transformer for Single Image Deraining
by: Li, Baiang, et al.
Published: (2024)
by: Li, Baiang, et al.
Published: (2024)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval
by: Dubey, Ayush, et al.
Published: (2024)
by: Dubey, Ayush, et al.
Published: (2024)
Topology-aware Human Avatars with Semantically-guided Gaussian Splatting
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
Human-Centric Transformer for Domain Adaptive Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval
by: Gao, Peng, et al.
Published: (2025)
by: Gao, Peng, et al.
Published: (2025)
Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
BitC-3DGS: High-Capacity 3D Gaussian Splatting Watermarking via Bit Compression
by: Bi, Yuquan, et al.
Published: (2026)
by: Bi, Yuquan, et al.
Published: (2026)
HRGR: Enhancing Image Manipulation Detection via Hierarchical Region-aware Graph Reasoning
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
by: Do, Jeonghyeok, et al.
Published: (2024)
by: Do, Jeonghyeok, et al.
Published: (2024)
INTENT: Invariance and Discrimination-aware Noise Mitigation for Robust Composed Image Retrieval
by: Chen, Zhiwei, et al.
Published: (2026)
by: Chen, Zhiwei, et al.
Published: (2026)
A Comprehensive Survey on Underwater Image Enhancement Based on Deep Learning
by: Cong, Xiaofeng, et al.
Published: (2024)
by: Cong, Xiaofeng, et al.
Published: (2024)
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
by: Zhou, Jun, et al.
Published: (2025)
by: Zhou, Jun, et al.
Published: (2025)
CascadeFormer: A Family of Two-stage Cascading Transformers for Skeleton-based Human Action Recognition
by: Peng, Yusen, et al.
Published: (2025)
by: Peng, Yusen, et al.
Published: (2025)
SITAR: Semi-supervised Image Transformer for Action Recognition
by: Iqbal, Owais, et al.
Published: (2024)
by: Iqbal, Owais, et al.
Published: (2024)
Similar Items
-
Point-Supervised Skeleton-Based Human Action Segmentation
by: Wang, Hongsong, et al.
Published: (2026) -
Towards Universal Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026) -
Attribution as Retrieval: Model-Agnostic AI-Generated Image Attribution
by: Wang, Hongsong, et al.
Published: (2026) -
Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition
by: Kuang, Jidong, et al.
Published: (2026) -
Heterogeneous Skeleton-Based Action Representation Learning
by: Wang, Hongsong, et al.
Published: (2025)