GPTFace: Generative Pre-training of Facial-Linguistic Transformer by Span Masking and Weakly Correlated Text-image Data
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yudong, Li, Hao, Hou, Xianxu, Shen, Linlin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
by: Wang, Xiaoqin, et al.
Published: (2025)
by: Wang, Xiaoqin, et al.
Published: (2025)
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
by: Wang, Xiaoqin, et al.
Published: (2025)
by: Wang, Xiaoqin, et al.
Published: (2025)
Deep Feature Consistent Variational Autoencoder
by: Hou, Xianxu, et al.
Published: (2016)
by: Hou, Xianxu, et al.
Published: (2016)
Supervision-by-Hallucination-and-Transfer: A Weakly-Supervised Approach for Robust and Precise Facial Landmark Detection
by: Wan, Jun, et al.
Published: (2026)
by: Wan, Jun, et al.
Published: (2026)
Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
LightQANet: Quantized and Adaptive Feature Learning for Low-Light Image Enhancement
by: Wu, Xu, et al.
Published: (2025)
by: Wu, Xu, et al.
Published: (2025)
MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
by: Cao, Yushe, et al.
Published: (2025)
by: Cao, Yushe, et al.
Published: (2025)
Generative Pre-training for Subjective Tasks: A Diffusion Transformer-Based Framework for Facial Beauty Prediction
by: Boukhari, Djamel Eddine, et al.
Published: (2025)
by: Boukhari, Djamel Eddine, et al.
Published: (2025)
Low-Light Enhancement Effect on Classification and Detection: An Empirical Study
by: Wu, Xu, et al.
Published: (2024)
by: Wu, Xu, et al.
Published: (2024)
MLIP: Medical Language-Image Pre-training with Masked Local Representation Learning
by: Liu, Jiarun, et al.
Published: (2024)
by: Liu, Jiarun, et al.
Published: (2024)
WeakSupCon: Weakly Supervised Contrastive Learning for Encoder Pre-training
by: Zhang, Bodong, et al.
Published: (2025)
by: Zhang, Bodong, et al.
Published: (2025)
DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
by: Gao, Ziqi, et al.
Published: (2025)
by: Gao, Ziqi, et al.
Published: (2025)
Visually Guided Generative Text-Layout Pre-training for Document Intelligence
by: Mao, Zhiming, et al.
Published: (2024)
by: Mao, Zhiming, et al.
Published: (2024)
Emerging Property of Masked Token for Effective Pre-training
by: Choi, Hyesong, et al.
Published: (2024)
by: Choi, Hyesong, et al.
Published: (2024)
Efficient Vision-Language Pre-training by Cluster Masking
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
Dataset Ownership Verification for Pre-trained Masked Models
by: Xie, Yuechen, et al.
Published: (2025)
by: Xie, Yuechen, et al.
Published: (2025)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
by: Li, Wenyu, et al.
Published: (2025)
by: Li, Wenyu, et al.
Published: (2025)
Rethinking UMM Visual Generation: Masked Modeling for Efficient Image-Only Pre-training
by: Sun, Peng, et al.
Published: (2026)
by: Sun, Peng, et al.
Published: (2026)
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
by: Yan, Xin, et al.
Published: (2023)
by: Yan, Xin, et al.
Published: (2023)
Bridging Synthetic and Real Worlds for Pre-training Scene Text Detectors
by: Guan, Tongkun, et al.
Published: (2023)
by: Guan, Tongkun, et al.
Published: (2023)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
Masked Pre-training Enables Universal Zero-shot Denoiser
by: Ma, Xiaoxiao, et al.
Published: (2024)
by: Ma, Xiaoxiao, et al.
Published: (2024)
Soften the Mask: Adaptive Temporal Soft Mask for Efficient Dynamic Facial Expression Recognition
by: Li, Meng-zhu, et al.
Published: (2025)
by: Li, Meng-zhu, et al.
Published: (2025)
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
by: Bai, Jinbin, et al.
Published: (2024)
by: Bai, Jinbin, et al.
Published: (2024)
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
by: Zhang, Yifei, et al.
Published: (2025)
by: Zhang, Yifei, et al.
Published: (2025)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
by: Xuan, Shiyu, et al.
Published: (2025)
by: Xuan, Shiyu, et al.
Published: (2025)
RobustFormer: Noise-Robust Pre-training for images and videos
by: Bastola, Ashish, et al.
Published: (2024)
by: Bastola, Ashish, et al.
Published: (2024)
ImplantFormer: Vision Transformer based Implant Position Regression Using Dental CBCT Data
by: Yang, Xinquan, et al.
Published: (2022)
by: Yang, Xinquan, et al.
Published: (2022)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data
by: He, Xilin, et al.
Published: (2024)
by: He, Xilin, et al.
Published: (2024)
Pre-training on Synthetic Driving Data for Trajectory Prediction
by: Li, Yiheng, et al.
Published: (2023)
by: Li, Yiheng, et al.
Published: (2023)
FineXtrol: Controllable Motion Generation via Fine-Grained Text
by: Shen, Keming, et al.
Published: (2025)
by: Shen, Keming, et al.
Published: (2025)
MaskDiffusion: Exploiting Pre-trained Diffusion Models for Semantic Segmentation
by: Kawano, Yasufumi, et al.
Published: (2024)
by: Kawano, Yasufumi, et al.
Published: (2024)
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
by: Lv, Meng, et al.
Published: (2026)
by: Lv, Meng, et al.
Published: (2026)
MaskedCLIP: Bridging the Masked and CLIP Space for Semi-Supervised Medical Vision-Language Pre-training
by: Zhu, Lei, et al.
Published: (2025)
by: Zhu, Lei, et al.
Published: (2025)
MaskHOI: Robust 3D Hand-Object Interaction Estimation via Masked Pre-training
by: Xie, Yuechen, et al.
Published: (2025)
by: Xie, Yuechen, et al.
Published: (2025)
Similar Items
-
DisFaceRep: Representation Disentanglement for Co-occurring Facial Components in Weakly Supervised Face Parsing
by: Wang, Xiaoqin, et al.
Published: (2025) -
FaceBench: A Multi-View Multi-Level Facial Attribute VQA Dataset for Benchmarking Face Perception MLLMs
by: Wang, Xiaoqin, et al.
Published: (2025) -
Deep Feature Consistent Variational Autoencoder
by: Hou, Xianxu, et al.
Published: (2016) -
Supervision-by-Hallucination-and-Transfer: A Weakly-Supervised Approach for Robust and Precise Facial Landmark Detection
by: Wan, Jun, et al.
Published: (2026) -
Data-efficient Event Camera Pre-training via Disentangled Masked Modeling
by: Huang, Zhenpeng, et al.
Published: (2024)