Salvato in:
| Autori principali: | Shaikh, Muhammad Bilal, Islam, Syed Mohammed Shamsul, Chai, Douglas, Akhtar, Naveed |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2405.15813 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deep Learning Approaches for Human Action Recognition in Video Data
di: Xie, Yufei
Pubblicazione: (2024)
di: Xie, Yufei
Pubblicazione: (2024)
Next-Generation License Plate Detection and Recognition System using YOLOv8
di: Amin, Arslan, et al.
Pubblicazione: (2025)
di: Amin, Arslan, et al.
Pubblicazione: (2025)
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
di: Liu, Sicheng, et al.
Pubblicazione: (2024)
di: Liu, Sicheng, et al.
Pubblicazione: (2024)
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)
Multi-Scale Spatial-Temporal Self-Attention Graph Convolutional Networks for Skeleton-based Action Recognition
di: Nakamura, Ikuo
Pubblicazione: (2024)
di: Nakamura, Ikuo
Pubblicazione: (2024)
Quantifying and Inducing Shape Bias in CNNs via Max-Pool Dilation
di: Sawada, Takito, et al.
Pubblicazione: (2026)
di: Sawada, Takito, et al.
Pubblicazione: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
Motion-Guided Semantic Alignment with Negative Prompts for Zero-Shot Video Action Recognition
di: Wang, Yiming, et al.
Pubblicazione: (2026)
di: Wang, Yiming, et al.
Pubblicazione: (2026)
A Challenging Benchmark of Anime Style Recognition
di: Li, Haotang, et al.
Pubblicazione: (2022)
di: Li, Haotang, et al.
Pubblicazione: (2022)
Pointing-Based Object Recognition
di: Hajdúch, Lukáš, et al.
Pubblicazione: (2026)
di: Hajdúch, Lukáš, et al.
Pubblicazione: (2026)
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
di: Hassan, Imtiaz Ul, et al.
Pubblicazione: (2026)
di: Hassan, Imtiaz Ul, et al.
Pubblicazione: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
Multimodal Action Quality Assessment
di: Zeng, Ling-An, et al.
Pubblicazione: (2024)
di: Zeng, Ling-An, et al.
Pubblicazione: (2024)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
di: Zheng, Peng, et al.
Pubblicazione: (2024)
di: Zheng, Peng, et al.
Pubblicazione: (2024)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
di: Han, Yudong, et al.
Pubblicazione: (2026)
di: Han, Yudong, et al.
Pubblicazione: (2026)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
di: Zhong, Zesen, et al.
Pubblicazione: (2025)
Multi-modal Sensor Fusion for Auto Driving Perception: A Survey
di: Huang, Keli, et al.
Pubblicazione: (2022)
di: Huang, Keli, et al.
Pubblicazione: (2022)
SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking
di: Mukhtar, Muhammad Taha, et al.
Pubblicazione: (2026)
di: Mukhtar, Muhammad Taha, et al.
Pubblicazione: (2026)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
di: Castrillón-Santana, Modesto, et al.
Pubblicazione: (2025)
YotoR-You Only Transform One Representation
di: Villa, José Ignacio Díaz, et al.
Pubblicazione: (2024)
di: Villa, José Ignacio Díaz, et al.
Pubblicazione: (2024)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
di: Hu, Xuran, et al.
Pubblicazione: (2026)
di: Hu, Xuran, et al.
Pubblicazione: (2026)
Action Anticipation from SoccerNet Football Video Broadcasts
di: Dalal, Mohamad, et al.
Pubblicazione: (2025)
di: Dalal, Mohamad, et al.
Pubblicazione: (2025)
UTAL-GNN: Unsupervised Temporal Action Localization using Graph Neural Networks
di: Badatya, Bikash Kumar, et al.
Pubblicazione: (2025)
di: Badatya, Bikash Kumar, et al.
Pubblicazione: (2025)
Lost in Context: The Influence of Context on Feature Attribution Methods for Object Recognition
di: Adhikari, Sayanta, et al.
Pubblicazione: (2024)
di: Adhikari, Sayanta, et al.
Pubblicazione: (2024)
Joint Learning of Depth, Pose, and Local Radiance Field for Large Scale Monocular 3D Reconstruction
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
Towards Hard and Soft Shadow Removal via Dual-Branch Separation Network and Vision Transformer
di: Liang, Jiajia
Pubblicazione: (2025)
di: Liang, Jiajia
Pubblicazione: (2025)
Pedestrian Detection in Low-Light Conditions: A Comprehensive Survey
di: Ghari, Bahareh, et al.
Pubblicazione: (2024)
di: Ghari, Bahareh, et al.
Pubblicazione: (2024)
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
di: Berjawi, Jad, et al.
Pubblicazione: (2025)
di: Berjawi, Jad, et al.
Pubblicazione: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
di: Rajapaksha, Uchitha, et al.
Pubblicazione: (2024)
The Influence of Iconicity in Transfer Learning for Sign Language Recognition
di: Artiaga, Keren, et al.
Pubblicazione: (2026)
di: Artiaga, Keren, et al.
Pubblicazione: (2026)
WatchHAR: Real-time On-device Human Activity Recognition System for Smartwatches
di: Yeon, Taeyoung, et al.
Pubblicazione: (2025)
di: Yeon, Taeyoung, et al.
Pubblicazione: (2025)
Domain-Adaptive Pretraining Improves Primate Behavior Recognition
di: Mueller, Felix B., et al.
Pubblicazione: (2025)
di: Mueller, Felix B., et al.
Pubblicazione: (2025)
From Latent to Engine Manifolds: Analyzing ImageBind's Multimodal Embedding Space
di: Hamara, Andrew, et al.
Pubblicazione: (2024)
di: Hamara, Andrew, et al.
Pubblicazione: (2024)
A Recipe for Geometry-Aware 3D Mesh Transformers
di: Farazi, Mohammad, et al.
Pubblicazione: (2024)
di: Farazi, Mohammad, et al.
Pubblicazione: (2024)
A Comprehensive Review of YOLO Architectures in Computer Vision: From YOLOv1 to YOLOv8 and YOLO-NAS
di: Terven, Juan, et al.
Pubblicazione: (2023)
di: Terven, Juan, et al.
Pubblicazione: (2023)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
di: Su, Qile, et al.
Pubblicazione: (2025)
di: Su, Qile, et al.
Pubblicazione: (2025)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
di: Tang, Guowei
Pubblicazione: (2026)
di: Tang, Guowei
Pubblicazione: (2026)
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
di: Fang, Tiancheng, et al.
Pubblicazione: (2026)
di: Fang, Tiancheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Deep Learning Approaches for Human Action Recognition in Video Data
di: Xie, Yufei
Pubblicazione: (2024) -
Next-Generation License Plate Detection and Recognition System using YOLOv8
di: Amin, Arslan, et al.
Pubblicazione: (2025) -
Distinguishing Visually Similar Actions: Prompt-Guided Semantic Prototype Modulation for Few-Shot Action Recognition
di: Li, Xiaoyang, et al.
Pubblicazione: (2025) -
SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
di: Liu, Sicheng, et al.
Pubblicazione: (2024) -
Context-Aware Network Based on Multi-scale Spatio-temporal Attention for Action Recognition in Videos
di: Li, Xiaoyang, et al.
Pubblicazione: (2025)