PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Yin, Chen, Zhichao, Xiao, Zeyu, Zhao, Yongle, An, Xiang, Yang, Kaicheng, Ran, Zimin, Guo, Jia, Feng, Ziyong, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
High-Fidelity Facial Albedo Estimation via Texture Quantization
by: Ran, Zimin, et al.
Published: (2024)
by: Ran, Zimin, et al.
Published: (2024)
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
by: Cao, Haofan, et al.
Published: (2026)
by: Cao, Haofan, et al.
Published: (2026)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
by: Chen, Zhichao, et al.
Published: (2026)
by: Chen, Zhichao, et al.
Published: (2026)
PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
by: Ping, Bowen, et al.
Published: (2025)
by: Ping, Bowen, et al.
Published: (2025)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
by: Cui, Siying, et al.
Published: (2024)
by: Cui, Siying, et al.
Published: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
by: Xie, Yin, et al.
Published: (2024)
by: Xie, Yin, et al.
Published: (2024)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
by: Dai, Dongyang, et al.
Published: (2025)
by: Dai, Dongyang, et al.
Published: (2025)
LMAD: Integrated End-to-End Vision-Language Model for Explainable Autonomous Driving
by: Song, Nan, et al.
Published: (2025)
by: Song, Nan, et al.
Published: (2025)
Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash Function
by: Yin, Shuo, et al.
Published: (2025)
by: Yin, Shuo, et al.
Published: (2025)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
by: Ma, Zehong, et al.
Published: (2025)
by: Ma, Zehong, et al.
Published: (2025)
End-to-End Speech Recognition with Pre-trained Masked Language Model
by: Higuchi, Yosuke, et al.
Published: (2024)
by: Higuchi, Yosuke, et al.
Published: (2024)
Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
ESC-MVQ: End-to-End Semantic Communication With Multi-Codebook Vector Quantization
by: Shin, Junyong, et al.
Published: (2025)
by: Shin, Junyong, et al.
Published: (2025)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
PixelBrax: Learning Continuous Control from Pixels End-to-End on the GPU
by: McInroe, Trevor, et al.
Published: (2025)
by: McInroe, Trevor, et al.
Published: (2025)
Masks and Manuscripts: Advancing Medical Pre-training with End-to-End Masking and Narrative Structuring
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
The End of Manual Decoding: Towards Truly End-to-End Language Models
by: Wang, Zhichao, et al.
Published: (2025)
by: Wang, Zhichao, et al.
Published: (2025)
Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving
by: Li, Tengpeng, et al.
Published: (2025)
by: Li, Tengpeng, et al.
Published: (2025)
AlignDrive: Aligned Lateral-Longitudinal Planning for End-to-End Autonomous Driving
by: Wu, Yanhao, et al.
Published: (2026)
by: Wu, Yanhao, et al.
Published: (2026)
TLC-Plan: A Two-Level Codebook Based Network for End-to-End Vector Floorplan Generation
by: Xiong, Biao, et al.
Published: (2026)
by: Xiong, Biao, et al.
Published: (2026)
TLC‐Plan: A Two‐Level Codebook Based Network for End‐to‐End Vector Floorplan Generation
by: Biao Xiong, et al.
Published: (2026)
by: Biao Xiong, et al.
Published: (2026)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
Enhancing Fully Formatted End-to-End Speech Recognition with Knowledge Distillation via Multi-Codebook Vector Quantization
by: You, Jian, et al.
Published: (2025)
by: You, Jian, et al.
Published: (2025)
PRIX: Learning to Plan from Raw Pixels for End-to-End Autonomous Driving
by: Wozniak, Maciej K., et al.
Published: (2025)
by: Wozniak, Maciej K., et al.
Published: (2025)
End-to-End Rate-Distortion Optimized 3D Gaussian Representation
by: Wang, Henan, et al.
Published: (2024)
by: Wang, Henan, et al.
Published: (2024)
Codebook-enabled Generative End-to-end Semantic Communication Powered by Transformer
by: Ye, Peigen, et al.
Published: (2024)
by: Ye, Peigen, et al.
Published: (2024)
DINO Pre-training for Vision-based End-to-end Autonomous Driving
by: Juneja, Shubham, et al.
Published: (2024)
by: Juneja, Shubham, et al.
Published: (2024)
Prototyping an End-to-End Multi-Modal Tiny-CNN for Cardiovascular Sensor Patches
by: Ibrahim, Mustafa Fuad Rifet, et al.
Published: (2025)
by: Ibrahim, Mustafa Fuad Rifet, et al.
Published: (2025)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset
by: Shen, Hengyu, et al.
Published: (2026)
by: Shen, Hengyu, et al.
Published: (2026)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
by: Ye, Junyi, et al.
Published: (2024)
by: Ye, Junyi, et al.
Published: (2024)
Similar Items
-
High-Fidelity Facial Albedo Estimation via Texture Quantization
by: Ran, Zimin, et al.
Published: (2024) -
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025) -
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024) -
PaCo-VLA: Passivity-Shielded Compliance Prior for Contact-Rich Vision-Language-Action Manipulation
by: Cao, Haofan, et al.
Published: (2026) -
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
by: Chen, Zhichao, et al.
Published: (2026)