Efficient Action Counting with Dynamic Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Xiaoxuan, Li, Zishi, Shang, Qiuyan, Zhu, Wentao, Ci, Hai, Qiao, Yu, Wang, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
by: Ye, Hang, et al.
Published: (2024)
by: Ye, Hang, et al.
Published: (2024)
Real-time Holistic Robot Pose Estimation with Unknown States
by: Ban, Shikun, et al.
Published: (2024)
by: Ban, Shikun, et al.
Published: (2024)
Seeing My Future: Predicting Situated Interaction Behavior in Virtual Reality
by: Xu, Yuan, et al.
Published: (2025)
by: Xu, Yuan, et al.
Published: (2025)
3D Human Mesh Estimation from Virtual Markers
by: Ma, Xiaoxuan, et al.
Published: (2023)
by: Ma, Xiaoxuan, et al.
Published: (2023)
Electromagnetic Inverse Scattering from a Single Transmitter
by: Cheng, Yizhe, et al.
Published: (2025)
by: Cheng, Yizhe, et al.
Published: (2025)
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
by: Wu, Rujie, et al.
Published: (2026)
by: Wu, Rujie, et al.
Published: (2026)
SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive Tokens
by: Su, Chi, et al.
Published: (2024)
by: Su, Chi, et al.
Published: (2024)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
by: Pang, Lexi, et al.
Published: (2025)
by: Pang, Lexi, et al.
Published: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
LLM-powered Query Expansion for Enhancing Boundary Prediction in Language-driven Action Localization
by: Shang, Zirui, et al.
Published: (2025)
by: Shang, Zirui, et al.
Published: (2025)
Efficient Temporal Action Segmentation via Boundary-aware Query Voting
by: Wang, Peiyao, et al.
Published: (2024)
by: Wang, Peiyao, et al.
Published: (2024)
Low-Resolution Action Recognition for Tiny Actions Challenge
by: Chen, Boyu, et al.
Published: (2022)
by: Chen, Boyu, et al.
Published: (2022)
Querying Autonomous Vehicle Point Clouds: Enhanced by 3D Object Counting with CounterNet
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using Octree Queries
by: Lu, Yuhang, et al.
Published: (2023)
by: Lu, Yuhang, et al.
Published: (2023)
Repetitive Action Counting with Hybrid Temporal Relation Modeling
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
by: Zhong, Fangwei, et al.
Published: (2024)
by: Zhong, Fangwei, et al.
Published: (2024)
DreamDA: Generative Data Augmentation with Diffusion Models
by: Fu, Yunxiang, et al.
Published: (2024)
by: Fu, Yunxiang, et al.
Published: (2024)
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
by: Yao, Ziyu, et al.
Published: (2025)
by: Yao, Ziyu, et al.
Published: (2025)
Impossible Videos
by: Bai, Zechen, et al.
Published: (2025)
by: Bai, Zechen, et al.
Published: (2025)
OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels
by: Lou, Meng, et al.
Published: (2025)
by: Lou, Meng, et al.
Published: (2025)
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
by: Lou, Meng, et al.
Published: (2026)
by: Lou, Meng, et al.
Published: (2026)
Generalized-Scale Object Counting with Gradual Query Aggregation
by: Pelhan, Jer, et al.
Published: (2025)
by: Pelhan, Jer, et al.
Published: (2025)
Visually-grounded Humanoid Agents
by: Ye, Hang, et al.
Published: (2026)
by: Ye, Hang, et al.
Published: (2026)
FDDet: Frequency-Decoupling for Boundary Refinement in Temporal Action Detection
by: Zhu, Xinnan, et al.
Published: (2025)
by: Zhu, Xinnan, et al.
Published: (2025)
AlphaChimp: Tracking and Behavior Recognition of Chimpanzees
by: Ma, Xiaoxuan, et al.
Published: (2024)
by: Ma, Xiaoxuan, et al.
Published: (2024)
B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
Causal Intervention for Subject-Deconfounded Facial Action Unit Recognition
by: Chen, Yingjie, et al.
Published: (2022)
by: Chen, Yingjie, et al.
Published: (2022)
Scaling Up Dynamic Human-Scene Interaction Modeling
by: Jiang, Nan, et al.
Published: (2024)
by: Jiang, Nan, et al.
Published: (2024)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception
by: Jin, Sheng, et al.
Published: (2023)
by: Jin, Sheng, et al.
Published: (2023)
Count Anything
by: Lei, Mengqi, et al.
Published: (2026)
by: Lei, Mengqi, et al.
Published: (2026)
LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic Counting
by: Wu, Hefeng, et al.
Published: (2023)
by: Wu, Hefeng, et al.
Published: (2023)
FCA-RAC: First Cycle Annotated Repetitive Action Counting
by: Lu, Jiada, et al.
Published: (2024)
by: Lu, Jiada, et al.
Published: (2024)
UNetMamba: An Efficient UNet-Like Mamba for Semantic Segmentation of High-Resolution Remote Sensing Images
by: Zhu, Enze, et al.
Published: (2024)
by: Zhu, Enze, et al.
Published: (2024)
GigaWorld-Policy: An Efficient Action-Centered World--Action Model
by: Ye, Angen, et al.
Published: (2026)
by: Ye, Angen, et al.
Published: (2026)
Steganalysis on Digital Watermarking: Is Your Defense Truly Impervious?
by: Yang, Pei, et al.
Published: (2024)
by: Yang, Pei, et al.
Published: (2024)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
by: Song, Yiren, et al.
Published: (2024)
by: Song, Yiren, et al.
Published: (2024)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
by: Ci, Hai, et al.
Published: (2024)
by: Ci, Hai, et al.
Published: (2024)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
Similar Items
-
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
by: Ye, Hang, et al.
Published: (2024) -
Real-time Holistic Robot Pose Estimation with Unknown States
by: Ban, Shikun, et al.
Published: (2024) -
Seeing My Future: Predicting Situated Interaction Behavior in Virtual Reality
by: Xu, Yuan, et al.
Published: (2025) -
3D Human Mesh Estimation from Virtual Markers
by: Ma, Xiaoxuan, et al.
Published: (2023) -
Electromagnetic Inverse Scattering from a Single Transmitter
by: Cheng, Yizhe, et al.
Published: (2025)