Dual-Process Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Luo, Grace, Granskog, Jonathan, Holynski, Aleksander, Darrell, Trevor |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vision-Language Models Create Cross-Modal Task Representations
por: Luo, Grace, et al.
Publicado: (2024)
por: Luo, Grace, et al.
Publicado: (2024)
Readout Guidance: Learning Control from Diffusion Features
por: Luo, Grace, et al.
Publicado: (2023)
por: Luo, Grace, et al.
Publicado: (2023)
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
por: Luo, Grace, et al.
Publicado: (2023)
por: Luo, Grace, et al.
Publicado: (2023)
Disentangled 3D Scene Generation with Layout Learning
por: Epstein, Dave, et al.
Publicado: (2024)
por: Epstein, Dave, et al.
Publicado: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
por: Mitra, Chancharik, et al.
Publicado: (2023)
por: Mitra, Chancharik, et al.
Publicado: (2023)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
por: Girrbach, Leander, et al.
Publicado: (2025)
por: Girrbach, Leander, et al.
Publicado: (2025)
Describing Differences in Image Sets with Natural Language
por: Dunlap, Lisa, et al.
Publicado: (2023)
por: Dunlap, Lisa, et al.
Publicado: (2023)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
por: Shang, Chuyi, et al.
Publicado: (2024)
por: Shang, Chuyi, et al.
Publicado: (2024)
TULIP: Towards Unified Language-Image Pretraining
por: Tang, Zineng, et al.
Publicado: (2025)
por: Tang, Zineng, et al.
Publicado: (2025)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
por: Huang, Brandon, et al.
Publicado: (2024)
por: Huang, Brandon, et al.
Publicado: (2024)
Measuring Diversity in Co-creative Image Generation
por: Ibarrola, Francisco, et al.
Publicado: (2024)
por: Ibarrola, Francisco, et al.
Publicado: (2024)
Constantly Improving Image Models Need Constantly Improving Benchmarks
por: Ge, Jiaxin, et al.
Publicado: (2025)
por: Ge, Jiaxin, et al.
Publicado: (2025)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
por: Lee, Heekyung, et al.
Publicado: (2025)
por: Lee, Heekyung, et al.
Publicado: (2025)
Analyzing The Language of Visual Tokens
por: Chan, David M., et al.
Publicado: (2024)
por: Chan, David M., et al.
Publicado: (2024)
Generative Image Dynamics
por: Li, Zhengqi, et al.
Publicado: (2023)
por: Li, Zhengqi, et al.
Publicado: (2023)
Segment Anything without Supervision
por: Wang, XuDong, et al.
Publicado: (2024)
por: Wang, XuDong, et al.
Publicado: (2024)
ALOHa: A New Measure for Hallucination in Captioning Models
por: Petryk, Suzanne, et al.
Publicado: (2024)
por: Petryk, Suzanne, et al.
Publicado: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
por: Wang, Xudong, et al.
Publicado: (2024)
por: Wang, Xudong, et al.
Publicado: (2024)
FairImagen: Post-Processing for Bias Mitigation in Text-to-Image Models
por: Fu, Zihao, et al.
Publicado: (2025)
por: Fu, Zihao, et al.
Publicado: (2025)
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
por: Jin, Haian, et al.
Publicado: (2026)
por: Jin, Haian, et al.
Publicado: (2026)
Rethinking Score Distillation as a Bridge Between Image Distributions
por: McAllister, David, et al.
Publicado: (2024)
por: McAllister, David, et al.
Publicado: (2024)
DELAN: Dual-Level Alignment for Vision-and-Language Navigation by Cross-Modal Contrastive Learning
por: Du, Mengfei, et al.
Publicado: (2024)
por: Du, Mengfei, et al.
Publicado: (2024)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
por: Lei, Jiayi, et al.
Publicado: (2025)
por: Lei, Jiayi, et al.
Publicado: (2025)
Language-Pretraining-Induced Bias: A Strong Foundation for General Vision Tasks
por: Luo, Yaxin, et al.
Publicado: (2026)
por: Luo, Yaxin, et al.
Publicado: (2026)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
por: Yu, Eric Yang, et al.
Publicado: (2024)
por: Yu, Eric Yang, et al.
Publicado: (2024)
Text-to-Image Cross-Modal Generation: A Systematic Review
por: Żelaszczyk, Maciej, et al.
Publicado: (2024)
por: Żelaszczyk, Maciej, et al.
Publicado: (2024)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
por: Han, Xiaochuang, et al.
Publicado: (2024)
por: Han, Xiaochuang, et al.
Publicado: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
por: Li, Boyi, et al.
Publicado: (2024)
por: Li, Boyi, et al.
Publicado: (2024)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
por: Yin, Shukang, et al.
Publicado: (2024)
por: Yin, Shukang, et al.
Publicado: (2024)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
por: Meng, Debin, et al.
Publicado: (2025)
por: Meng, Debin, et al.
Publicado: (2025)
Beyond Thumbs Up/Down: Untangling Challenges of Fine-Grained Feedback for Text-to-Image Generation
por: Collins, Katherine M., et al.
Publicado: (2024)
por: Collins, Katherine M., et al.
Publicado: (2024)
MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
por: Chen, Zhaorun, et al.
Publicado: (2024)
por: Chen, Zhaorun, et al.
Publicado: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
por: Luo, Jianjie, et al.
Publicado: (2024)
por: Luo, Jianjie, et al.
Publicado: (2024)
AiGen-FoodReview: A Multimodal Dataset of Machine-Generated Restaurant Reviews and Images on Social Media
por: Gambetti, Alessandro, et al.
Publicado: (2024)
por: Gambetti, Alessandro, et al.
Publicado: (2024)
Visual Lexicon: Rich Image Features in Language Space
por: Wang, XuDong, et al.
Publicado: (2024)
por: Wang, XuDong, et al.
Publicado: (2024)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
por: Gafni, Tomer, et al.
Publicado: (2025)
por: Gafni, Tomer, et al.
Publicado: (2025)
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference
por: Miranda, Imanol, et al.
Publicado: (2026)
por: Miranda, Imanol, et al.
Publicado: (2026)
Shape-Guided Diffusion with Inside-Outside Attention
por: Park, Dong Huk, et al.
Publicado: (2022)
por: Park, Dong Huk, et al.
Publicado: (2022)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
por: Yu, Junwei, et al.
Publicado: (2025)
por: Yu, Junwei, et al.
Publicado: (2025)
Ejemplares similares
-
Vision-Language Models Create Cross-Modal Task Representations
por: Luo, Grace, et al.
Publicado: (2024) -
Readout Guidance: Learning Control from Diffusion Features
por: Luo, Grace, et al.
Publicado: (2023) -
Diffusion Hyperfeatures: Searching Through Time and Space for Semantic Correspondence
por: Luo, Grace, et al.
Publicado: (2023) -
Disentangled 3D Scene Generation with Layout Learning
por: Epstein, Dave, et al.
Publicado: (2024) -
Compositional Chain-of-Thought Prompting for Large Multimodal Models
por: Mitra, Chancharik, et al.
Publicado: (2023)