Y-MAP-Net: Real-time depth, normals, segmentation, multi-label captioning and 2D human pose in RGB images
Fuente:
arXiv
Saved in:
| Main Authors: | Qammaz, Ammar, Vasilikopoulos, Nikolaos, Oikonomidis, Iason, Argyros, Antonis A. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024)
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
by: Spanakis, Michail, et al.
Published: (2026)
by: Spanakis, Michail, et al.
Published: (2026)
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025)
by: Karvounas, Giorgos, et al.
Published: (2025)
Combining Facial Videos and Biosignals for Stress Estimation During Driving
by: Valergaki, Paraskevi, et al.
Published: (2026)
by: Valergaki, Paraskevi, et al.
Published: (2026)
ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
by: Savathrakis, Giorgos, et al.
Published: (2024)
by: Savathrakis, Giorgos, et al.
Published: (2024)
Vision-Based Mistake Analysis in Procedural Activities: A Review of Advances and Challenges
by: Bacharidis, Konstantinos, et al.
Published: (2025)
by: Bacharidis, Konstantinos, et al.
Published: (2025)
Reflections on Diversity: A Real-time Virtual Mirror for Inclusive 3D Face Transformations
by: Valergaki, Paraskevi, et al.
Published: (2025)
by: Valergaki, Paraskevi, et al.
Published: (2025)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
by: Albadarneh, Israa A., et al.
Published: (2025)
by: Albadarneh, Israa A., et al.
Published: (2025)
Leveraging image captions for selective whole slide image annotation
by: Qiu, Jingna, et al.
Published: (2024)
by: Qiu, Jingna, et al.
Published: (2024)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
DETRPose: Real-time end-to-end transformer model for multi-person pose estimation
by: Janampa, Sebastian, et al.
Published: (2025)
by: Janampa, Sebastian, et al.
Published: (2025)
Recognizing Unseen States of Unknown Objects by Leveraging Knowledge Graphs
by: Gouidis, Filipos, et al.
Published: (2023)
by: Gouidis, Filipos, et al.
Published: (2023)
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
Maybe you don't need a U-Net: convolutional feature upsampling for materials micrograph segmentation
by: Docherty, Ronan, et al.
Published: (2025)
by: Docherty, Ronan, et al.
Published: (2025)
Vision Transformers increase efficiency of 3D cardiac CT multi-label segmentation
by: Jollans, Lee, et al.
Published: (2023)
by: Jollans, Lee, et al.
Published: (2023)
Anticipating Object State Changes in Long Procedural Videos
by: Manousaki, Victoria, et al.
Published: (2024)
by: Manousaki, Victoria, et al.
Published: (2024)
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
by: Benavent-Lledo, Manuel, et al.
Published: (2026)
Image captioning in different languages
by: van Miltenburg, Emiel
Published: (2024)
by: van Miltenburg, Emiel
Published: (2024)
DmADs-Net: Dense multiscale attention and depth-supervised network for medical image segmentation
by: Fu, Zhaojin, et al.
Published: (2024)
by: Fu, Zhaojin, et al.
Published: (2024)
Evaluating authenticity and quality of image captions via sentiment and semantic analyses
by: Krotov, Aleksei, et al.
Published: (2024)
by: Krotov, Aleksei, et al.
Published: (2024)
ProPLIKS: Probablistic 3D human body pose estimation
by: Shetty, Karthik, et al.
Published: (2024)
by: Shetty, Karthik, et al.
Published: (2024)
Flexible graph convolutional network for 3D human pose estimation
by: Shahjahan, Abu Taib Mohammed, et al.
Published: (2024)
by: Shahjahan, Abu Taib Mohammed, et al.
Published: (2024)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
LiCAR: pseudo-RGB LiDAR image for CAR segmentation
by: Páez-Ubieta, Ignacio de Loyola, et al.
Published: (2025)
by: Páez-Ubieta, Ignacio de Loyola, et al.
Published: (2025)
Adaptive graph Kolmogorov-Arnold network for 3D human pose estimation
by: Shahjahan, Abu Taib Mohammed, et al.
Published: (2025)
by: Shahjahan, Abu Taib Mohammed, et al.
Published: (2025)
Multi-hop graph transformer network for 3D human pose estimation
by: Islam, Zaedul, et al.
Published: (2024)
by: Islam, Zaedul, et al.
Published: (2024)
IndustryShapes: An RGB-D Benchmark dataset for 6D object pose estimation of industrial assembly components and tools
by: Sapoutzoglou, Panagiotis, et al.
Published: (2026)
by: Sapoutzoglou, Panagiotis, et al.
Published: (2026)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024)
by: Du, Zhicheng, et al.
Published: (2024)
Action Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
by: Benavent-Lledo, Manuel, et al.
Published: (2025)
Soft labelling for semantic segmentation: Bringing coherence to label down-sampling
by: Alcover-Couso, Roberto, et al.
Published: (2023)
by: Alcover-Couso, Roberto, et al.
Published: (2023)
MLN-net: A multi-source medical image segmentation method for clustered microcalcifications using multiple layer normalization
by: Wang, Ke, et al.
Published: (2023)
by: Wang, Ke, et al.
Published: (2023)
Fusing Domain-Specific Content from Large Language Models into Knowledge Graphs for Enhanced Zero Shot Object State Classification
by: Gouidis, Filippos, et al.
Published: (2024)
by: Gouidis, Filippos, et al.
Published: (2024)
A comprehensive framework for occluded human pose estimation
by: Xu, Linhao, et al.
Published: (2023)
by: Xu, Linhao, et al.
Published: (2023)
MDU-Net: Multi-scale Densely Connected U-Net for biomedical image segmentation
by: Zhang, Jiawei, et al.
Published: (2018)
by: Zhang, Jiawei, et al.
Published: (2018)
Multi-Modal interpretable automatic video captioning
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
Semantic search for 100M+ galaxy images using AI-generated captions
by: Koblischke, Nolan, et al.
Published: (2025)
by: Koblischke, Nolan, et al.
Published: (2025)
CromSS: Cross-modal pre-training with noisy labels for remote sensing image segmentation
by: Liu, Chenying, et al.
Published: (2024)
by: Liu, Chenying, et al.
Published: (2024)
Tiling artifacts and trade-offs of feature normalization in the segmentation of large biological images
by: Buglakova, Elena, et al.
Published: (2025)
by: Buglakova, Elena, et al.
Published: (2025)
Deep learning for 3D human pose estimation and mesh recovery: A survey
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
by: Benavent-Lledo, Manuel, et al.
Published: (2024)
Similar Items
-
D-PoSE: Depth as an Intermediate Representation for 3D Human Pose and Shape Estimation
by: Vasilikopoulos, Nikolaos, et al.
Published: (2024) -
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
by: Spanakis, Michail, et al.
Published: (2026) -
Enhancing Monocular 3D Hand Reconstruction with Learned Texture Priors
by: Karvounas, Giorgos, et al.
Published: (2025) -
Combining Facial Videos and Biosignals for Stress Estimation During Driving
by: Valergaki, Paraskevi, et al.
Published: (2026) -
ENACT: Entropy-based Clustering of Attention Input for Reducing the Computational Needs of Object Detection Transformers
by: Savathrakis, Giorgos, et al.
Published: (2024)