Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yulin, Zhang, Haoji, Yue, Yang, Song, Shiji, Deng, Chao, Feng, Junlan, Huang, Gao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
von: Yue, Yang, et al.
Veröffentlicht: (2025)
von: Yue, Yang, et al.
Veröffentlicht: (2025)
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
von: Yue, Yang, et al.
Veröffentlicht: (2025)
von: Yue, Yang, et al.
Veröffentlicht: (2025)
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
von: Wang, Yulin, et al.
Veröffentlicht: (2020)
von: Wang, Yulin, et al.
Veröffentlicht: (2020)
Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
von: Zheng, Ziwei, et al.
Veröffentlicht: (2024)
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
von: Wang, Yulin, et al.
Veröffentlicht: (2024)
von: Wang, Yulin, et al.
Veröffentlicht: (2024)
AdaGen: Learning Adaptive Policy for Image Synthesis
von: Ni, Zanlin, et al.
Veröffentlicht: (2026)
von: Ni, Zanlin, et al.
Veröffentlicht: (2026)
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
von: Bai, Jianhong, et al.
Veröffentlicht: (2024)
Depth Map Denoising Network and Lightweight Fusion Network for Enhanced 3D Face Recognition
von: Xu, Ruizhuo, et al.
Veröffentlicht: (2024)
von: Xu, Ruizhuo, et al.
Veröffentlicht: (2024)
AdaFPP: Adapt-Focused Bi-Propagating Prototype Learning for Panoramic Activity Recognition
von: Cao, Meiqi, et al.
Veröffentlicht: (2024)
von: Cao, Meiqi, et al.
Veröffentlicht: (2024)
AdaTooler-V: Adaptive Tool-Use for Images and Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
von: Yang, Le, et al.
Veröffentlicht: (2024)
von: Yang, Le, et al.
Veröffentlicht: (2024)
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
Motion Focus Recognition in Fast-Moving Egocentric Video
von: Hong, Si-En, et al.
Veröffentlicht: (2026)
von: Hong, Si-En, et al.
Veröffentlicht: (2026)
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding
von: Li, Handong, et al.
Veröffentlicht: (2026)
von: Li, Handong, et al.
Veröffentlicht: (2026)
AdaOcc: Adaptive-Resolution Occupancy Prediction
von: Chen, Chao, et al.
Veröffentlicht: (2024)
von: Chen, Chao, et al.
Veröffentlicht: (2024)
Enhancing Space-time Video Super-resolution via Spatial-temporal Feature Interaction
von: Yue, Zijie, et al.
Veröffentlicht: (2022)
von: Yue, Zijie, et al.
Veröffentlicht: (2022)
AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
von: Pang, Youxin, et al.
Veröffentlicht: (2025)
von: Pang, Youxin, et al.
Veröffentlicht: (2025)
The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
von: Wu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhuoyuan, et al.
Veröffentlicht: (2025)
ENAT: Rethinking Spatial-temporal Interactions in Token-based Image Synthesis
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
von: Ni, Zanlin, et al.
Veröffentlicht: (2024)
VA-Adapter: Adapting Ultrasound Foundation Model to Echocardiography Probe Guidance
von: Wang, Teng, et al.
Veröffentlicht: (2025)
von: Wang, Teng, et al.
Veröffentlicht: (2025)
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Everything to the Synthetic: Diffusion-driven Test-time Adaptation via Synthetic-Domain Alignment
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
von: Guo, Jiayi, et al.
Veröffentlicht: (2024)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
von: Tan, Haoyue, et al.
Veröffentlicht: (2026)
UniTransfer: Video Concept Transfer via Progressive Spatial and Timestep Decomposition
von: Lei, Guojun, et al.
Veröffentlicht: (2025)
von: Lei, Guojun, et al.
Veröffentlicht: (2025)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
von: Hu, Xiantao, et al.
Veröffentlicht: (2024)
von: Hu, Xiantao, et al.
Veröffentlicht: (2024)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
Improving Keystep Recognition in Ego-Video via Dexterous Focus
von: Chavis, Zachary, et al.
Veröffentlicht: (2025)
von: Chavis, Zachary, et al.
Veröffentlicht: (2025)
StepNet: Spatial-temporal Part-aware Network for Isolated Sign Language Recognition
von: Shen, Xiaolong, et al.
Veröffentlicht: (2022)
von: Shen, Xiaolong, et al.
Veröffentlicht: (2022)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2025)
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2025)
UniComp: Rethinking Video Compression Through Informational Uniqueness
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
von: Yuan, Chao, et al.
Veröffentlicht: (2025)
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models
von: Jiang, Hong, et al.
Veröffentlicht: (2026)
von: Jiang, Hong, et al.
Veröffentlicht: (2026)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
von: Wu, Song, et al.
Veröffentlicht: (2026)
von: Wu, Song, et al.
Veröffentlicht: (2026)
UniTTA: Unified Benchmark and Versatile Framework Towards Realistic Test-Time Adaptation
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
von: Du, Chaoqun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
von: Yang, Xiao, et al.
Veröffentlicht: (2026) -
Probabilistic Contrastive Learning for Long-Tailed Visual Recognition
von: Du, Chaoqun, et al.
Veröffentlicht: (2024) -
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
von: Yue, Yang, et al.
Veröffentlicht: (2025) -
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
von: Yue, Yang, et al.
Veröffentlicht: (2025) -
Meta-Semi: A Meta-learning Approach for Semi-supervised Learning
von: Wang, Yulin, et al.
Veröffentlicht: (2020)