Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Qianyi, Deb, Oishi, Patel, Amir, Rupprecht, Christian, Torr, Philip, Trigoni, Niki, Markham, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025)
by: Deb, Oishi, et al.
Published: (2025)
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
by: Vankadari, Madhu, et al.
Published: (2024)
by: Vankadari, Madhu, et al.
Published: (2024)
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
by: Shin, Sangyun, et al.
Published: (2023)
by: Shin, Sangyun, et al.
Published: (2023)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
Pre-training Feature Guided Diffusion Model for Speech Enhancement
by: Yang, Yiyuan, et al.
Published: (2024)
by: Yang, Yiyuan, et al.
Published: (2024)
Data Factory with Minimal Human Effort Using VLMs
by: Ye, Jiaojiao, et al.
Published: (2025)
by: Ye, Jiaojiao, et al.
Published: (2025)
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
by: Wang, Jianyuan, et al.
Published: (2023)
by: Wang, Jianyuan, et al.
Published: (2023)
Mitigating Cognitive Bias in RLHF by Altering Rationality
by: Horter, Tiffany, et al.
Published: (2026)
by: Horter, Tiffany, et al.
Published: (2026)
Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments
by: Xu, Shitong, et al.
Published: (2025)
by: Xu, Shitong, et al.
Published: (2025)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
by: Zhou, Kaichen, et al.
Published: (2020)
by: Zhou, Kaichen, et al.
Published: (2020)
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
DynPoint: Dynamic Neural Point For View Synthesis
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
WildDepth: A Multimodal Dataset for 3D Wildlife Perception and Depth Estimation
by: Aamir, Muhammad, et al.
Published: (2026)
by: Aamir, Muhammad, et al.
Published: (2026)
Scene-Conditional 3D Object Stylization and Composition
by: Zhou, Jinghao, et al.
Published: (2023)
by: Zhou, Jinghao, et al.
Published: (2023)
Multi-Modal Monocular Endoscopic Depth and Pose Estimation with Edge-Guided Self-Supervision
by: Ju, Xinwei, et al.
Published: (2026)
by: Ju, Xinwei, et al.
Published: (2026)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
by: Agarwal, Rachit, et al.
Published: (2026)
by: Agarwal, Rachit, et al.
Published: (2026)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
New keypoint-based approach for recognising British Sign Language (BSL) from sequences
by: Deb, Oishi, et al.
Published: (2024)
by: Deb, Oishi, et al.
Published: (2024)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
DiffPose-Animal: A Language-Conditioned Diffusion Framework for Animal Pose Estimation
by: Xiong, Tianyu, et al.
Published: (2025)
by: Xiong, Tianyu, et al.
Published: (2025)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
COOPERA: Continual Open-Ended Human-Robot Assistance
by: Ma, Chenyang, et al.
Published: (2025)
by: Ma, Chenyang, et al.
Published: (2025)
Mushroom Segmentation and 3D Pose Estimation from Point Clouds using Fully Convolutional Geometric Features and Implicit Pose Encoding
by: Retsinas, George, et al.
Published: (2024)
by: Retsinas, George, et al.
Published: (2024)
MMP: Towards Robust Multi-Modal Learning with Masked Modality Projection
by: Nezakati, Niki, et al.
Published: (2024)
by: Nezakati, Niki, et al.
Published: (2024)
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
by: Kaye, Ben, et al.
Published: (2024)
by: Kaye, Ben, et al.
Published: (2024)
Probabilistic Prompt Distribution Learning for Animal Pose Estimation
by: Rao, Jiyong, et al.
Published: (2025)
by: Rao, Jiyong, et al.
Published: (2025)
STEP: Simultaneous Tracking and Estimation of Pose for Animals and Humans
by: Verma, Shashikant, et al.
Published: (2025)
by: Verma, Shashikant, et al.
Published: (2025)
SPEAR: Receiver-to-Receiver Acoustic Neural Warping Field
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
Invisible Stitch: Generating Smooth 3D Scenes with Depth Inpainting
by: Engstler, Paul, et al.
Published: (2024)
by: Engstler, Paul, et al.
Published: (2024)
DeepInteraction++: Multi-Modality Interaction for Autonomous Driving
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
by: Han, Leezy, et al.
Published: (2026)
by: Han, Leezy, et al.
Published: (2026)
AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos
by: Wimbauer, Felix, et al.
Published: (2025)
by: Wimbauer, Felix, et al.
Published: (2025)
Farm3D: Learning Articulated 3D Animals by Distilling 2D Diffusion
by: Jakab, Tomas, et al.
Published: (2023)
by: Jakab, Tomas, et al.
Published: (2023)
Similar Items
-
Articulate3D: Zero-Shot Text-Driven 3D Object Posing
by: Deb, Oishi, et al.
Published: (2025) -
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
by: Vankadari, Madhu, et al.
Published: (2024) -
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024) -
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024) -
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023)