Efficient Personalized Text-to-image Generation by Leveraging Textual Subspace
Fuente:
arXiv
Guardado en:
| Autores principales: | Du, Shian, Cheng, Xiaotian, Qian, Qi, Wei, Henglu, Xu, Yi, Ji, Xiangyang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Solution for Single Object Tracking Task of Perception Test Challenge 2024
por: Zhong, Zhiqiang, et al.
Publicado: (2024)
por: Zhong, Zhiqiang, et al.
Publicado: (2024)
SynFog: A Photo-realistic Synthetic Fog Dataset based on End-to-end Imaging Simulation for Advancing Real-World Defogging in Autonomous Driving
por: Xie, Yiming, et al.
Publicado: (2024)
por: Xie, Yiming, et al.
Publicado: (2024)
ParCo: Part-Coordinating Text-to-Motion Synthesis
por: Zou, Qiran, et al.
Publicado: (2024)
por: Zou, Qiran, et al.
Publicado: (2024)
Beyond the Textual: Generating Coherent Visual Options for MCQs
por: Wang, Wanqiang, et al.
Publicado: (2025)
por: Wang, Wanqiang, et al.
Publicado: (2025)
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023)
por: Gordon, Brian, et al.
Publicado: (2023)
$λ$-ECLIPSE: Multi-Concept Personalized Text-to-Image Diffusion Models by Leveraging CLIP Latent Space
por: Patel, Maitreya, et al.
Publicado: (2024)
por: Patel, Maitreya, et al.
Publicado: (2024)
Tell Me What's Next: Textual Foresight for Generic UI Representations
por: Burns, Andrea, et al.
Publicado: (2024)
por: Burns, Andrea, et al.
Publicado: (2024)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
por: Wu, Yiming, et al.
Publicado: (2024)
por: Wu, Yiming, et al.
Publicado: (2024)
Solution for Point Tracking Task of ICCV 1st Perception Test Challenge 2023
por: Pan, Hongpeng, et al.
Publicado: (2024)
por: Pan, Hongpeng, et al.
Publicado: (2024)
VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
por: Islam, Md. Adnanul, et al.
Publicado: (2025)
por: Islam, Md. Adnanul, et al.
Publicado: (2025)
Optimizing Prompts for Text-to-Image Generation
por: Hao, Yaru, et al.
Publicado: (2022)
por: Hao, Yaru, et al.
Publicado: (2022)
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning
por: Wang, Yifan, et al.
Publicado: (2026)
por: Wang, Yifan, et al.
Publicado: (2026)
Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models
por: Xu, Yexing, et al.
Publicado: (2026)
por: Xu, Yexing, et al.
Publicado: (2026)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
por: Jing, Liqiang, et al.
Publicado: (2025)
por: Jing, Liqiang, et al.
Publicado: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
por: Cheng, Sheng, et al.
Publicado: (2024)
por: Cheng, Sheng, et al.
Publicado: (2024)
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
por: Du, Shian, et al.
Publicado: (2025)
por: Du, Shian, et al.
Publicado: (2025)
CLEAR: Character Unlearning in Textual and Visual Modalities
por: Dontsov, Alexey, et al.
Publicado: (2024)
por: Dontsov, Alexey, et al.
Publicado: (2024)
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
por: Gao, Xin, et al.
Publicado: (2026)
por: Gao, Xin, et al.
Publicado: (2026)
Directional Textual Inversion for Personalized Text-to-Image Generation
por: Kim, Kunhee, et al.
Publicado: (2025)
por: Kim, Kunhee, et al.
Publicado: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
por: Qi, Ji, et al.
Publicado: (2025)
por: Qi, Ji, et al.
Publicado: (2025)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
por: Ye, Junyi, et al.
Publicado: (2024)
por: Ye, Junyi, et al.
Publicado: (2024)
TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation
por: Ozaki, Shintaro, et al.
Publicado: (2025)
por: Ozaki, Shintaro, et al.
Publicado: (2025)
ID-EA: Identity-driven Text Enhancement and Adaptation with Textual Inversion for Personalized Text-to-Image Generation
por: Jin, Hyun-Jun, et al.
Publicado: (2025)
por: Jin, Hyun-Jun, et al.
Publicado: (2025)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
por: Lan, Jing, et al.
Publicado: (2026)
por: Lan, Jing, et al.
Publicado: (2026)
RoVRM: A Robust Visual Reward Model Optimized via Auxiliary Textual Preference Data
por: Wang, Chenglong, et al.
Publicado: (2024)
por: Wang, Chenglong, et al.
Publicado: (2024)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
por: Xu, Jiaqi, et al.
Publicado: (2023)
por: Xu, Jiaqi, et al.
Publicado: (2023)
PEA-Diffusion: Parameter-Efficient Adapter with Knowledge Distillation in non-English Text-to-Image Generation
por: Ma, Jian, et al.
Publicado: (2023)
por: Ma, Jian, et al.
Publicado: (2023)
DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
por: Zhang, Hongzhi, et al.
Publicado: (2025)
por: Zhang, Hongzhi, et al.
Publicado: (2025)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
por: Yuan, Shenghai, et al.
Publicado: (2024)
por: Yuan, Shenghai, et al.
Publicado: (2024)
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
por: Fallah, Forouzan, et al.
Publicado: (2025)
por: Fallah, Forouzan, et al.
Publicado: (2025)
Adaptive Subspace Projection for Generative Personalization
por: Nguyen, Van-Anh, et al.
Publicado: (2026)
por: Nguyen, Van-Anh, et al.
Publicado: (2026)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
por: Guo, Ziyu, et al.
Publicado: (2025)
por: Guo, Ziyu, et al.
Publicado: (2025)
Image-Text-Image Knowledge Transfer for Lifelong Person Re-Identification with Hybrid Clothing States
por: Wang, Qizao, et al.
Publicado: (2024)
por: Wang, Qizao, et al.
Publicado: (2024)
A Similarity Paradigm Through Textual Regularization Without Forgetting
por: Cui, Fangming, et al.
Publicado: (2025)
por: Cui, Fangming, et al.
Publicado: (2025)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
por: Wang, Chenglin, et al.
Publicado: (2026)
por: Wang, Chenglin, et al.
Publicado: (2026)
FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
por: Bourigault, Emmanuelle, et al.
Publicado: (2025)
por: Bourigault, Emmanuelle, et al.
Publicado: (2025)
VERA: Identifying and Leveraging Visual Evidence Retrieval Heads in Long-Context Understanding
por: Pei, Rongcan, et al.
Publicado: (2026)
por: Pei, Rongcan, et al.
Publicado: (2026)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
por: Luo, Yuxuan, et al.
Publicado: (2025)
por: Luo, Yuxuan, et al.
Publicado: (2025)
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
por: Tu, Rong-Cheng, et al.
Publicado: (2024)
Temporal Reasoning Transfer from Text to Video
por: Li, Lei, et al.
Publicado: (2024)
por: Li, Lei, et al.
Publicado: (2024)
Ejemplares similares
-
The Solution for Single Object Tracking Task of Perception Test Challenge 2024
por: Zhong, Zhiqiang, et al.
Publicado: (2024) -
SynFog: A Photo-realistic Synthetic Fog Dataset based on End-to-end Imaging Simulation for Advancing Real-World Defogging in Autonomous Driving
por: Xie, Yiming, et al.
Publicado: (2024) -
ParCo: Part-Coordinating Text-to-Motion Synthesis
por: Zou, Qiran, et al.
Publicado: (2024) -
Beyond the Textual: Generating Coherent Visual Options for MCQs
por: Wang, Wanqiang, et al.
Publicado: (2025) -
Mismatch Quest: Visual and Textual Feedback for Image-Text Misalignment
por: Gordon, Brian, et al.
Publicado: (2023)