Sapiens: Foundation for Human Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Khirodkar, Rawal, Bagautdinov, Timur, Martinez, Julieta, Zhaoen, Su, James, Austin, Selednik, Peter, Anderson, Stuart, Saito, Shunsuke |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sapiens2
by: Khirodkar, Rawal, et al.
Published: (2026)
by: Khirodkar, Rawal, et al.
Published: (2026)
Pippo: High-Resolution Multi-View Humans from a Single Image
by: Kant, Yash, et al.
Published: (2025)
by: Kant, Yash, et al.
Published: (2025)
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
by: Wu, Yiqian, et al.
Published: (2026)
by: Wu, Yiqian, et al.
Published: (2026)
TurboPortrait3D: Single-step diffusion-based fast portrait novel-view synthesis
by: Kim, Emily, et al.
Published: (2025)
by: Kim, Emily, et al.
Published: (2025)
Drivable 3D Gaussian Avatars
by: Zielonka, Wojciech, et al.
Published: (2023)
by: Zielonka, Wojciech, et al.
Published: (2023)
DuoMo: Dual Motion Diffusion for World-Space Human Reconstruction
by: Wang, Yufu, et al.
Published: (2026)
by: Wang, Yufu, et al.
Published: (2026)
SapiensID: Foundation for Human Recognition
by: Kim, Minchul, et al.
Published: (2025)
by: Kim, Minchul, et al.
Published: (2025)
ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling
by: Park, Jinhyung, et al.
Published: (2025)
by: Park, Jinhyung, et al.
Published: (2025)
Generalizable Neural Human Renderer
by: Masuda, Mana, et al.
Published: (2024)
by: Masuda, Mana, et al.
Published: (2024)
Harmony4D: A Video Dataset for In-The-Wild Close Human Interactions
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
URAvatar: Universal Relightable Gaussian Codec Avatars
by: Li, Junxuan, et al.
Published: (2024)
by: Li, Junxuan, et al.
Published: (2024)
CamLit: Unified Video Diffusion with Explicit Camera and Lighting Control
by: Kuang, Zhiyi, et al.
Published: (2026)
by: Kuang, Zhiyi, et al.
Published: (2026)
GeoRelight: Learning Joint Geometrical Relighting and Reconstruction with Flexible Multi-Modal Diffusion Transformers
by: Xue, Yuxuan, et al.
Published: (2026)
by: Xue, Yuxuan, et al.
Published: (2026)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
by: Ng, Evonne, et al.
Published: (2024)
by: Ng, Evonne, et al.
Published: (2024)
Multi-Person 3D Pose Estimation from Multi-View Uncalibrated Depth Cameras
by: Li, Yu-Jhe, et al.
Published: (2024)
by: Li, Yu-Jhe, et al.
Published: (2024)
Real-Time Simulated Avatar from Head-Mounted Sensors
by: Luo, Zhengyi, et al.
Published: (2024)
by: Luo, Zhengyi, et al.
Published: (2024)
Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models
by: Colin, Julien, et al.
Published: (2026)
by: Colin, Julien, et al.
Published: (2026)
Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone Images
by: Garbin, Emanuel, et al.
Published: (2025)
by: Garbin, Emanuel, et al.
Published: (2025)
MotivNet: Evolving Meta-Sapiens into an Emotionally Intelligent Foundation Model
by: Medicharla, Rahul, et al.
Published: (2025)
by: Medicharla, Rahul, et al.
Published: (2025)
Quickly Tuning Foundation Models for Image Segmentation
by: Das, Breenda, et al.
Published: (2025)
by: Das, Breenda, et al.
Published: (2025)
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
by: Tang, Jiapeng, et al.
Published: (2025)
by: Tang, Jiapeng, et al.
Published: (2025)
Expressive Whole-Body 3D Gaussian Avatar
by: Moon, Gyeongsik, et al.
Published: (2024)
by: Moon, Gyeongsik, et al.
Published: (2024)
VFA: Vision Frequency Analysis of Foundation Models and Human
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2024)
by: Darvishi-Bayazi, Mohammad-Javad, et al.
Published: (2024)
NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning
by: Rawal, Ishaan, et al.
Published: (2026)
by: Rawal, Ishaan, et al.
Published: (2026)
SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
by: Xiang, Shuai, et al.
Published: (2026)
by: Xiang, Shuai, et al.
Published: (2026)
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
by: Teuber, Carolin, et al.
Published: (2026)
by: Teuber, Carolin, et al.
Published: (2026)
Relightable Full-Body Gaussian Codec Avatars
by: Wang, Shaofei, et al.
Published: (2025)
by: Wang, Shaofei, et al.
Published: (2025)
Vision-Language Enhanced Foundation Model for Semi-supervised Medical Image Segmentation
by: Guo, Jiaqi, et al.
Published: (2025)
by: Guo, Jiaqi, et al.
Published: (2025)
Agtech Framework for Cranberry-Ripening Analysis Using Vision Foundation Models
by: Johnson, Faith, et al.
Published: (2024)
by: Johnson, Faith, et al.
Published: (2024)
On the Use of Hierarchical Vision Foundation Models for Low-Cost Human Mesh Recovery and Pose Estimation
by: Tarashima, Shuhei, et al.
Published: (2025)
by: Tarashima, Shuhei, et al.
Published: (2025)
Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
Generative Modeling of Shape-Dependent Self-Contact Human Poses
by: Ohkawa, Takehiko, et al.
Published: (2025)
by: Ohkawa, Takehiko, et al.
Published: (2025)
Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining
by: Li, Junxuan, et al.
Published: (2026)
by: Li, Junxuan, et al.
Published: (2026)
Learning to Fuse: Modality-Aware Adaptive Scheduling for Robust Multimodal Foundation Models
by: Bennett, Liam, et al.
Published: (2025)
by: Bennett, Liam, et al.
Published: (2025)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head Avatars
by: Kirschstein, Tobias, et al.
Published: (2025)
by: Kirschstein, Tobias, et al.
Published: (2025)
AIpparel: A Multimodal Foundation Model for Digital Garments
by: Nakayama, Kiyohiro, et al.
Published: (2024)
by: Nakayama, Kiyohiro, et al.
Published: (2024)
Score Distillation via Reparametrized DDIM
by: Lukoianov, Artem, et al.
Published: (2024)
by: Lukoianov, Artem, et al.
Published: (2024)
Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
by: Zhang, Xiaoran, et al.
Published: (2025)
by: Zhang, Xiaoran, et al.
Published: (2025)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
by: Zheng, Yaoyan, et al.
Published: (2025)
by: Zheng, Yaoyan, et al.
Published: (2025)
Similar Items
-
Sapiens2
by: Khirodkar, Rawal, et al.
Published: (2026) -
Pippo: High-Resolution Multi-View Humans from a Single Image
by: Kant, Yash, et al.
Published: (2025) -
GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild Videos
by: Wu, Yiqian, et al.
Published: (2026) -
TurboPortrait3D: Single-step diffusion-based fast portrait novel-view synthesis
by: Kim, Emily, et al.
Published: (2025) -
Drivable 3D Gaussian Avatars
by: Zielonka, Wojciech, et al.
Published: (2023)