CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Le, Anh-Duy, Pham, Van-Linh, Vo, Thanh-Nam, Mai, Xuan Toan, Tran, Tuan-Anh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2024)
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2024)
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025)
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)
Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2026)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2026)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2024)
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2024)
AsCL: An Asymmetry-sensitive Contrastive Learning Method for Image-Text Retrieval with Cross-Modal Fusion
von: Gong, Ziyu, et al.
Veröffentlicht: (2024)
von: Gong, Ziyu, et al.
Veröffentlicht: (2024)
Proposing Smart System for Detecting and Monitoring Vehicle Using Multiobject Multicamera Tracking
von: Phat Nguyen Huu, et al.
Veröffentlicht: (2024)
von: Phat Nguyen Huu, et al.
Veröffentlicht: (2024)
MarsSQE: Stereo Quality Enhancement for Martian Images Using Bi-level Cross-view Attention
von: Xu, Mai, et al.
Veröffentlicht: (2024)
von: Xu, Mai, et al.
Veröffentlicht: (2024)
Modality-Aware Contrastive and Uncertainty-Regularized Emotion Recognition
von: Zhuang, Yan, et al.
Veröffentlicht: (2026)
von: Zhuang, Yan, et al.
Veröffentlicht: (2026)
Enhancing Few-Shot Classification without Forgetting through Multi-Level Contrastive Constraints
von: Chen, Bingzhi, et al.
Veröffentlicht: (2024)
von: Chen, Bingzhi, et al.
Veröffentlicht: (2024)
Modeling the Impacts of Swipe Delay on User Quality of Experience in Short Video Streaming
von: Nguyen, Duc V., et al.
Veröffentlicht: (2026)
von: Nguyen, Duc V., et al.
Veröffentlicht: (2026)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
von: Yeh, Catherine, et al.
Veröffentlicht: (2026)
Adaptive Radial Projection on Fourier Magnitude Spectrum for Document Image Skew Estimation
von: Pham, Luan, et al.
Veröffentlicht: (2026)
von: Pham, Luan, et al.
Veröffentlicht: (2026)
MFQE 2.0: A New Approach for Multi-frame Quality Enhancement on Compressed Video
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
von: Xing, Qunliang, et al.
Veröffentlicht: (2019)
Is One-Shot In-Context Learning Helpful for Data Selection in Task-Specific Fine-Tuning of Multimodal LLMs?
von: An, Xiao, et al.
Veröffentlicht: (2026)
von: An, Xiao, et al.
Veröffentlicht: (2026)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
von: Nguyen, Duc, et al.
Veröffentlicht: (2024)
von: Nguyen, Duc, et al.
Veröffentlicht: (2024)
Continuous Patch Stitching for Block-wise Image Compression
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
Short-Form Video Viewing Behavior Analysis and Multi-Step Viewing Time Prediction
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
von: Yen, Vu Thi Hai, et al.
Veröffentlicht: (2026)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
von: Gupta, Parul, et al.
Veröffentlicht: (2025)
von: Gupta, Parul, et al.
Veröffentlicht: (2025)
Quality-Aware Dynamic Resolution Adaptation Framework for Adaptive Video Streaming
von: Premkumar, Amritha, et al.
Veröffentlicht: (2024)
von: Premkumar, Amritha, et al.
Veröffentlicht: (2024)
Real-time 3D Light-field Viewing with Eye-tracking on Conventional Displays
von: Pham, Trung Hieu, et al.
Veröffentlicht: (2025)
von: Pham, Trung Hieu, et al.
Veröffentlicht: (2025)
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
von: Yang, An, et al.
Veröffentlicht: (2025)
von: Yang, An, et al.
Veröffentlicht: (2025)
DA-PTQ: Drift-Aware Post-Training Quantization for Efficient Vision-Language-Action Models
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
von: Xu, Siyuan, et al.
Veröffentlicht: (2026)
Optimal Quality and Efficiency in Adaptive Live Streaming with JND-Aware Low latency Encoding
von: Menon, Vignesh V, et al.
Veröffentlicht: (2024)
von: Menon, Vignesh V, et al.
Veröffentlicht: (2024)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
von: Vo-Thanh, Hoang-Son, et al.
Veröffentlicht: (2024)
TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Style-Preserving Lip Sync via Audio-Aware Style Reference
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
von: Zhong, Weizhi, et al.
Veröffentlicht: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
CASUAL: Conditional Support Alignment for Domain Adaptation with Label Shift
von: Nguyen, Anh T, et al.
Veröffentlicht: (2023)
von: Nguyen, Anh T, et al.
Veröffentlicht: (2023)
A Hybrid Vision Transformer Approach for Mathematical Expression Recognition
von: Le, Anh Duy, et al.
Veröffentlicht: (2026)
von: Le, Anh Duy, et al.
Veröffentlicht: (2026)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Harmonizing Pixels and Melodies: Maestro-Guided Film Score Generation and Composition Style Transfer
von: Qi, F., et al.
Veröffentlicht: (2024)
von: Qi, F., et al.
Veröffentlicht: (2024)
SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2025)
von: Nguyen, Truong Thanh Hung, et al.
Veröffentlicht: (2025)
OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
von: Tran, Quang-Linh, et al.
Veröffentlicht: (2025)
von: Tran, Quang-Linh, et al.
Veröffentlicht: (2025)
Voxel-GS: Quantized Scaffold Gaussian Splatting Compression with Run-Length Coding
von: Fu, Chunyang, et al.
Veröffentlicht: (2025)
von: Fu, Chunyang, et al.
Veröffentlicht: (2025)
MU-MAE: Multimodal Masked Autoencoders-Based One-Shot Learning
von: Liu, Rex, et al.
Veröffentlicht: (2024)
von: Liu, Rex, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Ada2I: Enhancing Modality Balance for Multimodal Conversational Emotion Recognition
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2024) -
E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025) -
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025) -
Fact-Checking at Scale: Multimodal AI for Authenticity and Context Verification in Online Media
von: Phan, Van-Hoang, et al.
Veröffentlicht: (2025) -
Conversation Understanding using Relational Temporal Graph Neural Networks with Auxiliary Cross-Modality Interaction
von: Nguyen, Cam-Van Thi, et al.
Veröffentlicht: (2023)