Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dell'Erba, Samuele, Bagdanov, Andrew D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
von: Feng, Yichen, et al.
Veröffentlicht: (2026)
von: Feng, Yichen, et al.
Veröffentlicht: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
von: Skripkin, Matvey, et al.
Veröffentlicht: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
von: Liu, Shanyuan, et al.
Veröffentlicht: (2023)
von: Liu, Shanyuan, et al.
Veröffentlicht: (2023)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
Chat-Driven Text Generation and Interaction for Person Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2025)
von: Xie, Zequn, et al.
Veröffentlicht: (2025)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
von: Shahin, Nada, et al.
Veröffentlicht: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
von: Xu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2025)
RoboScript: Code Generation for Free-Form Manipulation Tasks across Real and Simulation
von: Chen, Junting, et al.
Veröffentlicht: (2024)
von: Chen, Junting, et al.
Veröffentlicht: (2024)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
von: Le, Van-Truong
Veröffentlicht: (2026)
von: Le, Van-Truong
Veröffentlicht: (2026)
Analyzing Quality, Bias, and Performance in Text-to-Image Generative Models
von: Masrourisaadat, Nila, et al.
Veröffentlicht: (2024)
von: Masrourisaadat, Nila, et al.
Veröffentlicht: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2025)
A Surveillance Based Interactive Robot
von: Kavimandan, Kshitij, et al.
Veröffentlicht: (2025)
von: Kavimandan, Kshitij, et al.
Veröffentlicht: (2025)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
von: Shahin, Nada, et al.
Veröffentlicht: (2026)
Privacy-Preserving Structureless Visual Localization via Image Obfuscation
von: Panek, Vojtech, et al.
Veröffentlicht: (2026)
von: Panek, Vojtech, et al.
Veröffentlicht: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
Combining Absolute and Semi-Generalized Relative Poses for Visual Localization
von: Panek, Vojtech, et al.
Veröffentlicht: (2024)
von: Panek, Vojtech, et al.
Veröffentlicht: (2024)
Memory-Efficient Differentially Private Training with Gradient Random Projection
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
von: Mulrooney, Alex, et al.
Veröffentlicht: (2025)
OmniFusion Technical Report
von: Goncharova, Elizaveta, et al.
Veröffentlicht: (2024)
von: Goncharova, Elizaveta, et al.
Veröffentlicht: (2024)
OkanNet: A Lightweight Deep Learning Architecture for Classification of Brain Tumor from MRI Images
von: Uçar, Okan, et al.
Veröffentlicht: (2026)
von: Uçar, Okan, et al.
Veröffentlicht: (2026)
IntrinsiX: High-Quality PBR Generation using Image Priors
von: Kocsis, Peter, et al.
Veröffentlicht: (2025)
von: Kocsis, Peter, et al.
Veröffentlicht: (2025)
Towards Localizing Structural Elements: Merging Geometrical Detection with Semantic Verification in RGB-D Data
von: Tourani, Ali, et al.
Veröffentlicht: (2024)
von: Tourani, Ali, et al.
Veröffentlicht: (2024)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
von: Bonial, Claire, et al.
Veröffentlicht: (2024)
von: Bonial, Claire, et al.
Veröffentlicht: (2024)
SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
von: Lukin, Stephanie M., et al.
Veröffentlicht: (2024)
von: Lukin, Stephanie M., et al.
Veröffentlicht: (2024)
AVControl: Efficient Framework for Training Audio-Visual Controls
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
von: Niu, Yuwei, et al.
Veröffentlicht: (2025) -
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024) -
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025) -
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
von: Feng, Yichen, et al.
Veröffentlicht: (2026) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)