GRIF-DM: Generation of Rich Impression Fonts using Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Lei, Yang, Fei, Wang, Kai, Souibgui, Mohamed Ali, Gomez, Lluis, Fornés, Alicia, Valveny, Ernest, Karatzas, Dimosthenis |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024)
by: Kang, Lei, et al.
Published: (2024)
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
by: Souibgui, Mohamed Ali, et al.
Published: (2025)
ComicsPAP: understanding comic strips by picking the correct panel
by: Vivoli, Emanuele, et al.
Published: (2025)
by: Vivoli, Emanuele, et al.
Published: (2025)
Learning Quantifiable Visual Explanations Without Ground-Truth
by: Singh, Amritpal, et al.
Published: (2026)
by: Singh, Amritpal, et al.
Published: (2026)
Reading in the Dark: Low-light Scene Text Recognition
by: Fu, Xuanshuo, et al.
Published: (2026)
by: Fu, Xuanshuo, et al.
Published: (2026)
A Fast Hierarchical Method for Multi-script and Arbitrary Oriented Scene Text Extraction
by: Gomez, Lluis, et al.
Published: (2014)
by: Gomez, Lluis, et al.
Published: (2014)
Image-text matching for large-scale book collections
by: Llabrés, Artemis, et al.
Published: (2024)
by: Llabrés, Artemis, et al.
Published: (2024)
Multimodal Transformer for Comics Text-Cloze
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
One missing piece in Vision and Language: A Survey on Comics Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
xLSTM-ECG: Multi-label ECG Classification via Feature Fusion with xLSTM
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Retrieval Augmented Verification for Zero-Shot Detection of Multimodal Disinformation
by: Dey, Arka Ujjal, et al.
Published: (2024)
by: Dey, Arka Ujjal, et al.
Published: (2024)
Privacy-Aware Document Visual Question Answering
by: Tito, Rubèn, et al.
Published: (2023)
by: Tito, Rubèn, et al.
Published: (2023)
CSSL-MHTR: Continual Self-Supervised Learning for Scalable Multi-script Handwritten Text Recognition
by: Dhiaf, Marwa, et al.
Published: (2023)
by: Dhiaf, Marwa, et al.
Published: (2023)
LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis
by: Kang, Lei, et al.
Published: (2025)
by: Kang, Lei, et al.
Published: (2025)
Federated Document Visual Question Answering: A Pilot Study
by: Nguyen, Khanh, et al.
Published: (2024)
by: Nguyen, Khanh, et al.
Published: (2024)
Impression-CLIP: Contrastive Shape-Impression Embedding for Fonts
by: Kubota, Yugo, et al.
Published: (2024)
by: Kubota, Yugo, et al.
Published: (2024)
AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering
by: Li, Zongmin, et al.
Published: (2026)
by: Li, Zongmin, et al.
Published: (2026)
Enhancing Document VQA Models via Retrieval-Augmented Generation
by: López, Eric, et al.
Published: (2025)
by: López, Eric, et al.
Published: (2025)
Font Impression Estimation in the Wild
by: Kitajima, Kazuki, et al.
Published: (2024)
by: Kitajima, Kazuki, et al.
Published: (2024)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
by: Mishra, Pritam, et al.
Published: (2025)
by: Mishra, Pritam, et al.
Published: (2025)
CoMix: A Comprehensive Benchmark for Multi-Task Comic Understanding
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
by: Mishra, Pritam, et al.
Published: (2026)
by: Mishra, Pritam, et al.
Published: (2026)
ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering
by: Lassoued, Aymen, et al.
Published: (2026)
by: Lassoued, Aymen, et al.
Published: (2026)
Hierarchical Co-Embedding of Font Shapes and Impression Tags
by: Kubota, Yugo, et al.
Published: (2026)
by: Kubota, Yugo, et al.
Published: (2026)
ComiCap: A VLMs pipeline for dense captioning of Comic Panels
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Counterfeit Answers: Adversarial Forgery against OCR-Free Document Visual Question Answering
by: Pintore, Marco, et al.
Published: (2025)
by: Pintore, Marco, et al.
Published: (2025)
Embedding Font Impression Word Tags Based on Co-occurrence
by: Kubota, Yugo, et al.
Published: (2025)
by: Kubota, Yugo, et al.
Published: (2025)
CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books
by: Ortega, Marc Serra, et al.
Published: (2025)
by: Ortega, Marc Serra, et al.
Published: (2025)
GAN-based Content-Conditioned Generation of Handwritten Musical Symbols
by: Asbert, Gerard, et al.
Published: (2025)
by: Asbert, Gerard, et al.
Published: (2025)
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
by: Lagos, Maximiliano Hormazábal, et al.
Published: (2025)
Learning to Decipher from Pixels -- A Case Study of Copiale
by: Kang, Lei, et al.
Published: (2026)
by: Kang, Lei, et al.
Published: (2026)
Reading Between the Lanes: Text VideoQA on the Road
by: Tom, George, et al.
Published: (2023)
by: Tom, George, et al.
Published: (2023)
NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA
by: Tobaben, Marlon, et al.
Published: (2024)
by: Tobaben, Marlon, et al.
Published: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Subjective Face Transform using Human First Impressions
by: Roygaga, Chaitanya, et al.
Published: (2023)
by: Roygaga, Chaitanya, et al.
Published: (2023)
FontStudio: Shape-Adaptive Diffusion Model for Coherent and Consistent Font Effect Generation
by: Mu, Xinzhi, et al.
Published: (2024)
by: Mu, Xinzhi, et al.
Published: (2024)
VecFusion: Vector Font Generation with Diffusion
by: Thamizharasan, Vikas, et al.
Published: (2023)
by: Thamizharasan, Vikas, et al.
Published: (2023)
Similar Items
-
Machine Unlearning for Document Classification
by: Kang, Lei, et al.
Published: (2024) -
Preserving Privacy Without Compromising Accuracy: Machine Unlearning for Handwritten Text Recognition
by: Kang, Lei, et al.
Published: (2025) -
Multi-Page Document Visual Question Answering using Self-Attention Scoring Mechanism
by: Kang, Lei, et al.
Published: (2024) -
A Benchmark for Symbolic Reasoning from Pixel Sequences: Grid-Level Visual Completion and Correction
by: Kang, Lei, et al.
Published: (2025) -
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
by: Souibgui, Mohamed Ali, et al.
Published: (2025)