Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Haonan, Luo, Yuxuan, Lian, Zhouhui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
di: Luo, Yuxuan, et al.
Pubblicazione: (2025)
di: Luo, Yuxuan, et al.
Pubblicazione: (2025)
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
di: Tang, Hao, et al.
Pubblicazione: (2025)
di: Tang, Hao, et al.
Pubblicazione: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
HFH-Font: Few-shot Chinese Font Synthesis with Higher Quality, Faster Speed, and Higher Resolution
di: Li, Hua, et al.
Pubblicazione: (2024)
di: Li, Hua, et al.
Pubblicazione: (2024)
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)
Visual Autoregressive Modeling for Instruction-Guided Image Editing
di: Mao, Qingyang, et al.
Pubblicazione: (2025)
di: Mao, Qingyang, et al.
Pubblicazione: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
Across-Game Engagement Modelling via Few-Shot Learning
di: Pinitas, Kosmas, et al.
Pubblicazione: (2024)
di: Pinitas, Kosmas, et al.
Pubblicazione: (2024)
Context-aware TFL: A Universal Context-aware Contrastive Learning Framework for Temporal Forgery Localization
di: Yin, Qilin, et al.
Pubblicazione: (2025)
di: Yin, Qilin, et al.
Pubblicazione: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
di: Farina, Matteo, et al.
Pubblicazione: (2025)
di: Farina, Matteo, et al.
Pubblicazione: (2025)
MU-MAE: Multimodal Masked Autoencoders-Based One-Shot Learning
di: Liu, Rex, et al.
Pubblicazione: (2024)
di: Liu, Rex, et al.
Pubblicazione: (2024)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
di: Shukor, Mustafa, et al.
Pubblicazione: (2023)
di: Shukor, Mustafa, et al.
Pubblicazione: (2023)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
di: Roy, Prasun, et al.
Pubblicazione: (2019)
di: Roy, Prasun, et al.
Pubblicazione: (2019)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
di: Das, Alloy, et al.
Pubblicazione: (2023)
di: Das, Alloy, et al.
Pubblicazione: (2023)
TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views
di: Hung, Hsiang-Hui, et al.
Pubblicazione: (2025)
di: Hung, Hsiang-Hui, et al.
Pubblicazione: (2025)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
di: Zhao, Pengcheng, et al.
Pubblicazione: (2024)
di: Zhao, Pengcheng, et al.
Pubblicazione: (2024)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
di: Zheng, Guangting, et al.
Pubblicazione: (2025)
di: Zheng, Guangting, et al.
Pubblicazione: (2025)
A Simple Task-aware Contrastive Local Descriptor Selection Strategy for Few-shot Learning between inter class and intra class
di: Qiao, Qian, et al.
Pubblicazione: (2024)
di: Qiao, Qian, et al.
Pubblicazione: (2024)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
di: Li, Yingxuan, et al.
Pubblicazione: (2024)
di: Li, Yingxuan, et al.
Pubblicazione: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
di: Li, Yongqi, et al.
Pubblicazione: (2024)
di: Li, Yongqi, et al.
Pubblicazione: (2024)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
di: Xu, Shuolin, et al.
Pubblicazione: (2025)
di: Xu, Shuolin, et al.
Pubblicazione: (2025)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
di: Cai, Lingling, et al.
Pubblicazione: (2024)
di: Cai, Lingling, et al.
Pubblicazione: (2024)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
di: Zhao, Xinyu, et al.
Pubblicazione: (2026)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
di: Hao, Jing, et al.
Pubblicazione: (2025)
di: Hao, Jing, et al.
Pubblicazione: (2025)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
di: Shi, Haoyuan, et al.
Pubblicazione: (2026)
di: Shi, Haoyuan, et al.
Pubblicazione: (2026)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
di: Yuan, Hangjie, et al.
Pubblicazione: (2025)
di: Yuan, Hangjie, et al.
Pubblicazione: (2025)
Continuous Patch Stitching for Block-wise Image Compression
di: Zhang, Zifu, et al.
Pubblicazione: (2025)
di: Zhang, Zifu, et al.
Pubblicazione: (2025)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
di: Wu, Yuheng, et al.
Pubblicazione: (2026)
di: Wu, Yuheng, et al.
Pubblicazione: (2026)
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
di: Imajuku, Yuki, et al.
Pubblicazione: (2024)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
di: Lionar, Stefan, et al.
Pubblicazione: (2025)
di: Lionar, Stefan, et al.
Pubblicazione: (2025)
Creatively Upscaling Images with Global-Regional Priors
di: Qian, Yurui, et al.
Pubblicazione: (2025)
di: Qian, Yurui, et al.
Pubblicazione: (2025)
URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection
di: Wang, Zhenyu, et al.
Pubblicazione: (2026)
di: Wang, Zhenyu, et al.
Pubblicazione: (2026)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
di: Wang, Zihang, et al.
Pubblicazione: (2026)
di: Wang, Zihang, et al.
Pubblicazione: (2026)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
di: Jiang, Jingjing, et al.
Pubblicazione: (2025)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
di: Li, Zhangbin, et al.
Pubblicazione: (2024)
di: Li, Zhangbin, et al.
Pubblicazione: (2024)
ArchGPT: Understanding the World's Architectures with Large Multimodal Models
di: Wang, Yuze, et al.
Pubblicazione: (2025)
di: Wang, Yuze, et al.
Pubblicazione: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
di: Dai, Yusheng, et al.
Pubblicazione: (2026)
di: Dai, Yusheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
di: Luo, Yuxuan, et al.
Pubblicazione: (2025) -
Connecting Giants: Synergistic Knowledge Transfer of Large Multimodal Models for Few-Shot Learning
di: Tang, Hao, et al.
Pubblicazione: (2025) -
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025) -
HFH-Font: Few-shot Chinese Font Synthesis with Higher Quality, Faster Speed, and Higher Resolution
di: Li, Hua, et al.
Pubblicazione: (2024) -
CONSTANT: Towards High-Quality One-Shot Handwriting Generation with Patch Contrastive Enhancement and Style-Aware Quantization
di: Le, Anh-Duy, et al.
Pubblicazione: (2026)