Saved in:
| Main Authors: | Xu, Chao, Zhang, Suyu, Liu, Yang, Sun, Baigui, Chen, Weihong, Xu, Bo, Liu, Qi, Wang, Juncheng, Wang, Shujun, Luo, Shan, Peters, Jan, Vasilakos, Athanasios V., Zafeiriou, Stefanos, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.11362 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025)
by: Liu, Dayong, et al.
Published: (2025)
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
by: Potamias, Rolandos Alexandros, et al.
Published: (2025)
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
by: Babiloni, Francesca, et al.
Published: (2024)
by: Babiloni, Francesca, et al.
Published: (2024)
Improving face generation quality and prompt following with synthetic captions
by: Tarasiou, Michail, et al.
Published: (2024)
by: Tarasiou, Michail, et al.
Published: (2024)
Deep Face Restoration: A Survey
by: Wang, Tao, et al.
Published: (2022)
by: Wang, Tao, et al.
Published: (2022)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
by: Wang, Juncheng, et al.
Published: (2026)
by: Wang, Juncheng, et al.
Published: (2026)
TransFace++: Rethinking the Face Recognition Paradigm with a Focus on Accuracy, Efficiency, and Security
by: Dan, Jun, et al.
Published: (2023)
by: Dan, Jun, et al.
Published: (2023)
Spatio-temporal Prompting Network for Robust Video Feature Extraction
by: Sun, Guanxiong, et al.
Published: (2024)
by: Sun, Guanxiong, et al.
Published: (2024)
MaDiS: Taming Masked Diffusion Language Models for Sign Language Generation
by: Zuo, Ronglai, et al.
Published: (2026)
by: Zuo, Ronglai, et al.
Published: (2026)
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024)
by: Ververas, Evangelos, et al.
Published: (2024)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
by: Zuo, Ronglai, et al.
Published: (2024)
by: Zuo, Ronglai, et al.
Published: (2024)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation
by: Qian, Yijie, et al.
Published: (2025)
by: Qian, Yijie, et al.
Published: (2025)
Dex2HOI: Dexterous Bimanual Two-Object Interaction Generation
by: Pratikaki, Chrysa, et al.
Published: (2026)
by: Pratikaki, Chrysa, et al.
Published: (2026)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
by: Chen, Zerui, et al.
Published: (2026)
by: Chen, Zerui, et al.
Published: (2026)
Neural Sign Actors: A diffusion model for 3D sign language production from text
by: Baltatzis, Vasileios, et al.
Published: (2023)
by: Baltatzis, Vasileios, et al.
Published: (2023)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Arc2Face: A Foundation Model for ID-Consistent Human Faces
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2024)
Geo-ID: Test-Time Geometric Consensus for Cross-View Consistent Intrinsics
by: Dirik, Alara, et al.
Published: (2026)
by: Dirik, Alara, et al.
Published: (2026)
360SFUDA++: Towards Source-free UDA for Panoramic Segmentation by Learning Reliable Category Prototypes
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Semantics, Distortion, and Style Matter: Towards Source-free UDA for Panoramic Segmentation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
TopoFR: A Closer Look at Topology Alignment on Face Recognition
by: Dan, Jun, et al.
Published: (2024)
by: Dan, Jun, et al.
Published: (2024)
Unsupervised Visible-Infrared ReID via Pseudo-label Correction and Modality-level Alignment
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
ID-Consistent, Precise Expression Generation with Blendshape-Guided Diffusion
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
by: Papantoniou, Foivos Paraperas, et al.
Published: (2025)
MA-FSAR: Multimodal Adaptation of CLIP for Few-Shot Action Recognition
by: Xing, Jiazheng, et al.
Published: (2023)
by: Xing, Jiazheng, et al.
Published: (2023)
Synchronized Video-to-Audio Generation via Mel Quantization-Continuum Decomposition
by: Wang, Juncheng, et al.
Published: (2025)
by: Wang, Juncheng, et al.
Published: (2025)
Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera
by: Yu, Zhengdi, et al.
Published: (2024)
by: Yu, Zhengdi, et al.
Published: (2024)
DermaFlux: Synthetic Skin Lesion Generation with Rectified Flows for Enhanced Image Classification
by: Galanakis, Stathis, et al.
Published: (2026)
by: Galanakis, Stathis, et al.
Published: (2026)
UV-free Texture Generation with Denoising and Geodesic Heat Diffusions
by: Foti, Simone, et al.
Published: (2024)
by: Foti, Simone, et al.
Published: (2024)
Distribution Matching for Multi-Task Learning of Classification Tasks: a Large-Scale Study on Faces & Beyond
by: Kollias, Dimitrios, et al.
Published: (2024)
by: Kollias, Dimitrios, et al.
Published: (2024)
360$^\circ$ High-Resolution Depth Estimation via Uncertainty-aware Structural Knowledge Transfer
by: Cao, Zidong, et al.
Published: (2023)
by: Cao, Zidong, et al.
Published: (2023)
Design2Cloth: 3D Cloth Generation from 2D Masks
by: Zheng, Jiali, et al.
Published: (2024)
by: Zheng, Jiali, et al.
Published: (2024)
Context-self contrastive pretraining for crop type semantic segmentation
by: Tarasiou, Michail, et al.
Published: (2021)
by: Tarasiou, Michail, et al.
Published: (2021)
On Committor Functions in Milestoning
by: Ji, Xiaojun, et al.
Published: (2023)
by: Ji, Xiaojun, et al.
Published: (2023)
PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling
by: Dirik, Alara, et al.
Published: (2025)
by: Dirik, Alara, et al.
Published: (2025)
ReasonX: MLLM-Guided Intrinsic Image Decomposition
by: Dirik, Alara, et al.
Published: (2025)
by: Dirik, Alara, et al.
Published: (2025)
Beyond Accuracy: Robustness, Interpretability and Expressiveness of EEG Foundation Models
by: Širca, Urban, et al.
Published: (2026)
by: Širca, Urban, et al.
Published: (2026)
Large Learning Rates Simultaneously Achieve Robustness to Spurious Correlations and Compressibility
by: Barsbey, Melih, et al.
Published: (2025)
by: Barsbey, Melih, et al.
Published: (2025)
FitDiff: Robust monocular 3D facial shape and reflectance estimation using Diffusion Models
by: Galanakis, Stathis, et al.
Published: (2023)
by: Galanakis, Stathis, et al.
Published: (2023)
Similar Items
-
Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
by: Liu, Dayong, et al.
Published: (2025) -
ImHead: A Large-scale Implicit Morphable Model for Localized Head Modeling
by: Potamias, Rolandos Alexandros, et al.
Published: (2025) -
ID-to-3D: Expressive ID-guided 3D Heads via Score Distillation Sampling
by: Babiloni, Francesca, et al.
Published: (2024) -
Improving face generation quality and prompt following with synthetic captions
by: Tarasiou, Michail, et al.
Published: (2024) -
Deep Face Restoration: A Survey
by: Wang, Tao, et al.
Published: (2022)