Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Jamil, Sofia, Reddy, Bollampalli Areen, Kumar, Raghvendra, Saha, Sriparna, Joseph, K J, Goswami, Koustava |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
por: Jamil, Sofia, et al.
Publicado: (2025)
por: Jamil, Sofia, et al.
Publicado: (2025)
Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
por: Jamil, Sofia, et al.
Publicado: (2025)
por: Jamil, Sofia, et al.
Publicado: (2025)
Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
por: Jamil, Sofia, et al.
Publicado: (2025)
por: Jamil, Sofia, et al.
Publicado: (2025)
GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance
por: Jamil, Sofia, et al.
Publicado: (2025)
por: Jamil, Sofia, et al.
Publicado: (2025)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
por: Wang, Shanshan, et al.
Publicado: (2026)
por: Wang, Shanshan, et al.
Publicado: (2026)
Step-by-step Layered Design Generation
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
por: Khan, Faizan Farooq, et al.
Publicado: (2025)
PixelDiT: Pixel Diffusion Transformers for Image Generation
por: Yu, Yongsheng, et al.
Publicado: (2025)
por: Yu, Yongsheng, et al.
Publicado: (2025)
SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation
por: Nag, Sayan, et al.
Publicado: (2024)
por: Nag, Sayan, et al.
Publicado: (2024)
When Background Matters: Breaking Medical Vision Language Models by Transferable Attack
por: Ghosh, Akash, et al.
Publicado: (2026)
por: Ghosh, Akash, et al.
Publicado: (2026)
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
por: Xu, Gangwei, et al.
Publicado: (2025)
por: Xu, Gangwei, et al.
Publicado: (2025)
Ask Me Again Differently: GRAS for Measuring Bias in Vision Language Models on Gender, Race, Age, and Skin Tone
por: Malik, Shaivi, et al.
Publicado: (2025)
por: Malik, Shaivi, et al.
Publicado: (2025)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
por: Jiang, Liyao, et al.
Publicado: (2024)
por: Jiang, Liyao, et al.
Publicado: (2024)
PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation
por: Jing, Zonglei, et al.
Publicado: (2025)
por: Jing, Zonglei, et al.
Publicado: (2025)
Agentic Design Review System
por: Nag, Sayan, et al.
Publicado: (2025)
por: Nag, Sayan, et al.
Publicado: (2025)
PromptRR: Diffusion Models as Prompt Generators for Single Image Reflection Removal
por: Wang, Tao, et al.
Publicado: (2024)
por: Wang, Tao, et al.
Publicado: (2024)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
por: Kuchibhotla, Hari Chandana, et al.
Publicado: (2024)
por: Kuchibhotla, Hari Chandana, et al.
Publicado: (2024)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
por: NVIDIA, et al.
Publicado: (2024)
por: NVIDIA, et al.
Publicado: (2024)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
por: Sahoo, Pranab, et al.
Publicado: (2024)
por: Sahoo, Pranab, et al.
Publicado: (2024)
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
por: Kumar, Raghvendra, et al.
Publicado: (2026)
por: Kumar, Raghvendra, et al.
Publicado: (2026)
Concept-to-Pixel: Prompt-Free Universal Medical Image Segmentation
por: Chen, Haoyun, et al.
Publicado: (2026)
por: Chen, Haoyun, et al.
Publicado: (2026)
Poetry2Image: An Iterative Correction Framework for Images Generated from Chinese Classical Poetry
por: Jiang, Jing, et al.
Publicado: (2024)
por: Jiang, Jing, et al.
Publicado: (2024)
Pixels, Patterns, but No Poetry: To See The World like Humans
por: Gao, Hongcheng, et al.
Publicado: (2025)
por: Gao, Hongcheng, et al.
Publicado: (2025)
AntifakePrompt: Prompt-Tuned Vision-Language Models are Fake Image Detectors
por: Chang, You-Ming, et al.
Publicado: (2023)
por: Chang, You-Ming, et al.
Publicado: (2023)
Simulating Post-Neoadjuvant Chemotherapy Breast Cancer MRI via Diffusion Model with Prompt Tuning
por: Kim, Jonghun, et al.
Publicado: (2025)
por: Kim, Jonghun, et al.
Publicado: (2025)
PixelRush: Ultra-Fast, Training-Free High-Resolution Image Generation via One-step Diffusion
por: Lai, Hong-Phuc, et al.
Publicado: (2026)
por: Lai, Hong-Phuc, et al.
Publicado: (2026)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
por: Cao, Yukang, et al.
Publicado: (2025)
por: Cao, Yukang, et al.
Publicado: (2025)
From Pampas to Pixels: Fine-Tuning Diffusion Models for Gaúcho Heritage
por: Amadeus, Marcellus, et al.
Publicado: (2024)
por: Amadeus, Marcellus, et al.
Publicado: (2024)
PixIE: Prompted Pixel-Space Low-Light Image Enhancement
por: Lin, Ruirui, et al.
Publicado: (2026)
por: Lin, Ruirui, et al.
Publicado: (2026)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
por: Baade, Alan, et al.
Publicado: (2026)
por: Baade, Alan, et al.
Publicado: (2026)
FedMRL: Data Heterogeneity Aware Federated Multi-agent Deep Reinforcement Learning for Medical Imaging
por: Sahoo, Pranab, et al.
Publicado: (2024)
por: Sahoo, Pranab, et al.
Publicado: (2024)
URSimulator: Human-Perception-Driven Prompt Tuning for Enhanced Virtual Urban Renewal via Diffusion Models
por: Hu, Chuanbo, et al.
Publicado: (2024)
por: Hu, Chuanbo, et al.
Publicado: (2024)
PixelFlow: Pixel-Space Generative Models with Flow
por: Chen, Shoufa, et al.
Publicado: (2025)
por: Chen, Shoufa, et al.
Publicado: (2025)
Multimodal Large Language Models for Medical Report Generation via Customized Prompt Tuning
por: Li, Chunlei, et al.
Publicado: (2025)
por: Li, Chunlei, et al.
Publicado: (2025)
Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models
por: Shih, Chun-Yen, et al.
Publicado: (2024)
por: Shih, Chun-Yen, et al.
Publicado: (2024)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
por: Zhang, David Junhao, et al.
Publicado: (2023)
por: Zhang, David Junhao, et al.
Publicado: (2023)
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation
por: Ma, Zehong, et al.
Publicado: (2025)
por: Ma, Zehong, et al.
Publicado: (2025)
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos
por: Maity, Krishanu, et al.
Publicado: (2024)
por: Maity, Krishanu, et al.
Publicado: (2024)
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning
por: Zhang, Lei, et al.
Publicado: (2026)
por: Zhang, Lei, et al.
Publicado: (2026)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
por: Ghosh, Akash, et al.
Publicado: (2024)
por: Ghosh, Akash, et al.
Publicado: (2024)
Fine-Tuning Text-To-Image Diffusion Models for Class-Wise Spurious Feature Generation
por: MaungMaung, AprilPyone, et al.
Publicado: (2024)
por: MaungMaung, AprilPyone, et al.
Publicado: (2024)
Ejemplares similares
-
PoemTale Diffusion: Minimising Information Loss in Poem to Image Generation with Multi-Stage Prompt Refinement
por: Jamil, Sofia, et al.
Publicado: (2025) -
Do It Yourself (DIY): Modifying Images for Poems in a Zero-Shot Setting Using Weighted Prompt Manipulation
por: Jamil, Sofia, et al.
Publicado: (2025) -
Crossing Borders: A Multimodal Challenge for Indian Poetry Translation and Image Generation
por: Jamil, Sofia, et al.
Publicado: (2025) -
GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance
por: Jamil, Sofia, et al.
Publicado: (2025) -
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
por: Wang, Shanshan, et al.
Publicado: (2026)