Salvato in:
| Autori principali: | Bianchi, Edoardo, Liotta, Antonio |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.08665 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback
di: Bianchi, Edoardo, et al.
Pubblicazione: (2026)
di: Bianchi, Edoardo, et al.
Pubblicazione: (2026)
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
CountFormer: Multi-View Crowd Counting Transformer
di: Mo, Hong, et al.
Pubblicazione: (2024)
di: Mo, Hong, et al.
Pubblicazione: (2024)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
Gate-Shift-Pose: Enhancing Action Recognition in Sports with Skeleton Information
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025)
ViewFormer: Exploring Spatiotemporal Modeling for Multi-View 3D Occupancy Perception via View-Guided Transformers
di: Li, Jinke, et al.
Pubblicazione: (2024)
di: Li, Jinke, et al.
Pubblicazione: (2024)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
di: Azad, Shehreen, et al.
Pubblicazione: (2025)
di: Azad, Shehreen, et al.
Pubblicazione: (2025)
Omni-Video: Democratizing Unified Video Understanding and Generation
di: Tan, Zhiyu, et al.
Pubblicazione: (2025)
di: Tan, Zhiyu, et al.
Pubblicazione: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
di: Wei, Cong, et al.
Pubblicazione: (2025)
di: Wei, Cong, et al.
Pubblicazione: (2025)
STGFormer: Spatio-Temporal GraphFormer for 3D Human Pose Estimation in Video
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
Unified Panoramic Geometry Estimation via Multi-View Foundation Models
di: Bozic, Vukasin, et al.
Pubblicazione: (2026)
di: Bozic, Vukasin, et al.
Pubblicazione: (2026)
Calisthenics Skills Temporal Video Segmentation
di: Finocchiaro, Antonio, et al.
Pubblicazione: (2025)
di: Finocchiaro, Antonio, et al.
Pubblicazione: (2025)
Dense Dynamic Scene Reconstruction and Camera Pose Estimation from Multi-View Videos
di: Sun, Shuo, et al.
Pubblicazione: (2026)
di: Sun, Shuo, et al.
Pubblicazione: (2026)
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives
di: Peirone, Simone Alberto, et al.
Pubblicazione: (2024)
di: Peirone, Simone Alberto, et al.
Pubblicazione: (2024)
WidthFormer: Toward Efficient Transformer-based BEV View Transformation
di: Yang, Chenhongyi, et al.
Pubblicazione: (2024)
di: Yang, Chenhongyi, et al.
Pubblicazione: (2024)
Proto-Former: Unified Facial Landmark Detection by Prototype Transformer
di: Hu, Shengkai, et al.
Pubblicazione: (2025)
di: Hu, Shengkai, et al.
Pubblicazione: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
di: Li, Kunchang, et al.
Pubblicazione: (2022)
di: Li, Kunchang, et al.
Pubblicazione: (2022)
Video Understanding: From Geometry and Semantics to Unified Models
di: An, Zhaochong, et al.
Pubblicazione: (2026)
di: An, Zhaochong, et al.
Pubblicazione: (2026)
UniPTS: A Unified Framework for Proficient Post-Training Sparsity
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
PicoEyes: Unified Gaze Estimation Framework for Mixed Reality with a Large-Scale Multi-View Dataset
di: Duan, Fuxin, et al.
Pubblicazione: (2026)
di: Duan, Fuxin, et al.
Pubblicazione: (2026)
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
di: Chen, Feng, et al.
Pubblicazione: (2025)
di: Chen, Feng, et al.
Pubblicazione: (2025)
CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
di: Tanoue, Hayato, et al.
Pubblicazione: (2025)
di: Tanoue, Hayato, et al.
Pubblicazione: (2025)
DriveTok: 3D Driving Scene Tokenization for Unified Multi-View Reconstruction and Understanding
di: Zhuo, Dong, et al.
Pubblicazione: (2026)
di: Zhuo, Dong, et al.
Pubblicazione: (2026)
FootFormer: Estimating Stability from Visual Input
di: Kraiger, Keaton, et al.
Pubblicazione: (2025)
di: Kraiger, Keaton, et al.
Pubblicazione: (2025)
VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
di: Eskandar, George, et al.
Pubblicazione: (2026)
di: Eskandar, George, et al.
Pubblicazione: (2026)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
di: Shi, Mengqi, et al.
Pubblicazione: (2026)
di: Shi, Mengqi, et al.
Pubblicazione: (2026)
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
di: Huang, Wen, et al.
Pubblicazione: (2025)
di: Huang, Wen, et al.
Pubblicazione: (2025)
KASportsFormer: Kinematic Anatomy Enhanced Transformer for 3D Human Pose Estimation on Short Sports Scene Video
di: Yin, Zhuoer, et al.
Pubblicazione: (2025)
di: Yin, Zhuoer, et al.
Pubblicazione: (2025)
BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation
di: Pan, Yulu, et al.
Pubblicazione: (2025)
di: Pan, Yulu, et al.
Pubblicazione: (2025)
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
di: Lu, Hannan, et al.
Pubblicazione: (2024)
di: Lu, Hannan, et al.
Pubblicazione: (2024)
BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation
di: Liu, Zhijian, et al.
Pubblicazione: (2022)
di: Liu, Zhijian, et al.
Pubblicazione: (2022)
A Unified Framework for Human-centric Point Cloud Video Understanding
di: Xu, Yiteng, et al.
Pubblicazione: (2024)
di: Xu, Yiteng, et al.
Pubblicazione: (2024)
MultiWorld: Scalable Multi-Agent Multi-View Video World Models
di: Wu, Haoyu, et al.
Pubblicazione: (2026)
di: Wu, Haoyu, et al.
Pubblicazione: (2026)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
di: Zhou, Xingcheng, et al.
Pubblicazione: (2025)
di: Zhou, Xingcheng, et al.
Pubblicazione: (2025)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2025)
di: Chatterjee, Dibyadip, et al.
Pubblicazione: (2025)
MultiFormer: A Multi-Person Pose Estimation System Based on CSI and Attention Mechanism
di: Qu, Yanyi, et al.
Pubblicazione: (2025)
di: Qu, Yanyi, et al.
Pubblicazione: (2025)
Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
di: Sun, Xiangyu, et al.
Pubblicazione: (2025)
di: Sun, Xiangyu, et al.
Pubblicazione: (2025)
BezierFormer: A Unified Architecture for 2D and 3D Lane Detection
di: Dong, Zhiwei, et al.
Pubblicazione: (2024)
di: Dong, Zhiwei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025) -
Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback
di: Bianchi, Edoardo, et al.
Pubblicazione: (2026) -
ProfVLM: A lightweight video-language model for multi-view proficiency estimation
di: Bianchi, Edoardo, et al.
Pubblicazione: (2025) -
CountFormer: Multi-View Crowd Counting Transformer
di: Mo, Hong, et al.
Pubblicazione: (2024) -
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
di: Li, Xiangtai, et al.
Pubblicazione: (2023)