LVLM-Composer's Explicit Planning for Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ramsey, Spencer, Lee, Jeffrey, Grant, Amina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
von: Garcia, Fernando Gabriela, et al.
Veröffentlicht: (2025)
von: Garcia, Fernando Gabriela, et al.
Veröffentlicht: (2025)
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
von: Hwang, Yerin, et al.
Veröffentlicht: (2025)
PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
von: Mahadev, Rohan, et al.
Veröffentlicht: (2026)
von: Mahadev, Rohan, et al.
Veröffentlicht: (2026)
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
von: Lim, Jiyoung, et al.
Veröffentlicht: (2026)
von: Lim, Jiyoung, et al.
Veröffentlicht: (2026)
FashionComposer: Compositional Fashion Image Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
von: Ji, Sihui, et al.
Veröffentlicht: (2024)
Compose and Conquer: Diffusion-Based 3D Depth Aware Composable Image Synthesis
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
LAVID: An Agentic LVLM Framework for Diffusion-Generated Video Detection
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2025)
An LLM-LVLM Driven Agent for Iterative and Fine-Grained Image Editing
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2025)
von: Li, Zixu, et al.
Veröffentlicht: (2025)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2025)
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2025)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
von: Chen, Zining, et al.
Veröffentlicht: (2025)
von: Chen, Zining, et al.
Veröffentlicht: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
von: Khan, Zeeshan, et al.
Veröffentlicht: (2025)
AdaIAT: Adaptively Increasing Attention to Generated Text to Alleviate Hallucinations in LVLM
von: Zhong, Li'an, et al.
Veröffentlicht: (2026)
von: Zhong, Li'an, et al.
Veröffentlicht: (2026)
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
von: Zou, Quanchen, et al.
Veröffentlicht: (2025)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
von: Chen, Zhekai, et al.
Veröffentlicht: (2024)
LVLM-Aware Multimodal Retrieval for RAG-Based Medical Diagnosis with General-Purpose Models
von: Mazor, Nir, et al.
Veröffentlicht: (2025)
von: Mazor, Nir, et al.
Veröffentlicht: (2025)
Instance-Level Composed Image Retrieval
von: Psomas, Bill, et al.
Veröffentlicht: (2025)
von: Psomas, Bill, et al.
Veröffentlicht: (2025)
Zero Shot Composed Image Retrieval
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
von: Kakarla, Santhosh, et al.
Veröffentlicht: (2025)
Composed Image Retrieval for Remote Sensing
von: Psomas, Bill, et al.
Veröffentlicht: (2024)
von: Psomas, Bill, et al.
Veröffentlicht: (2024)
NCL-CIR: Noise-aware Contrastive Learning for Composed Image Retrieval
von: Gao, Peng, et al.
Veröffentlicht: (2025)
von: Gao, Peng, et al.
Veröffentlicht: (2025)
Generating a Paracosm for Training-Free Zero-Shot Composed Image Retrieval
von: Wang, Tong, et al.
Veröffentlicht: (2026)
von: Wang, Tong, et al.
Veröffentlicht: (2026)
From Mapping to Composing: A Two-Stage Framework for Zero-shot Composed Image Retrieval
von: Wang, Yabing, et al.
Veröffentlicht: (2025)
von: Wang, Yabing, et al.
Veröffentlicht: (2025)
Multi-Level LVLM Guidance for Untrimmed Video Action Recognition
von: Peng, Liyang, et al.
Veröffentlicht: (2025)
von: Peng, Liyang, et al.
Veröffentlicht: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
von: Stan, Gabriela Ben Melech, et al.
Veröffentlicht: (2024)
A Sanity Check on Composed Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
von: Liu, Yikun, et al.
Veröffentlicht: (2026)
Zero-shot Composed Text-Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
Controllable Image Generation with Composed Parallel Token Prediction
von: Stirling, Jamie, et al.
Veröffentlicht: (2024)
von: Stirling, Jamie, et al.
Veröffentlicht: (2024)
EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback
von: Jia, Jingyang, et al.
Veröffentlicht: (2025)
von: Jia, Jingyang, et al.
Veröffentlicht: (2025)
MedM-VL: What Makes a Good Medical LVLM?
von: Shi, Yiming, et al.
Veröffentlicht: (2025)
von: Shi, Yiming, et al.
Veröffentlicht: (2025)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
Fighting Hallucinations with Counterfactuals: Diffusion-Guided Perturbations for LVLM Hallucination Suppression
von: Dastmalchi, Hamidreza, et al.
Veröffentlicht: (2026)
von: Dastmalchi, Hamidreza, et al.
Veröffentlicht: (2026)
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
von: Yan, Bei, et al.
Veröffentlicht: (2026)
von: Yan, Bei, et al.
Veröffentlicht: (2026)
LVLM-empowered Multi-modal Representation Learning for Visual Place Recognition
von: Wang, Teng, et al.
Veröffentlicht: (2024)
von: Wang, Teng, et al.
Veröffentlicht: (2024)
VALD: Multi-Stage Vision Attack Detection for Efficient LVLM Defense
von: Kadvil, Nadav, et al.
Veröffentlicht: (2026)
von: Kadvil, Nadav, et al.
Veröffentlicht: (2026)
Composing Object Relations and Attributes for Image-Text Matching
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
von: Pham, Khoi, et al.
Veröffentlicht: (2024)
Benchmarking Composed Image Retrieval for Applied Earth Observation
von: Psomas, Bill, et al.
Veröffentlicht: (2026)
von: Psomas, Bill, et al.
Veröffentlicht: (2026)
Composed Image Retrieval for Training-Free Domain Conversion
von: Efthymiadis, Nikos, et al.
Veröffentlicht: (2024)
von: Efthymiadis, Nikos, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unlocking Compositional Control: Self-Supervision for LVLM-Based Image Generation
von: Garcia, Fernando Gabriela, et al.
Veröffentlicht: (2025) -
Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation
von: Hwang, Yerin, et al.
Veröffentlicht: (2025) -
PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
von: Mahadev, Rohan, et al.
Veröffentlicht: (2026) -
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
von: Lim, Jiyoung, et al.
Veröffentlicht: (2026) -
FashionComposer: Compositional Fashion Image Generation
von: Ji, Sihui, et al.
Veröffentlicht: (2024)