GLoD: Composing Global Contexts and Local Details in Image Generation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Yamada, Moyuru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
von: Takemoto, Kentaro, et al.
Veröffentlicht: (2023)
von: Takemoto, Kentaro, et al.
Veröffentlicht: (2023)
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
von: Kolouju, Pranavi, et al.
Veröffentlicht: (2025)
von: Kolouju, Pranavi, et al.
Veröffentlicht: (2025)
D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
von: Rahimi, Amir, et al.
Veröffentlicht: (2023)
von: Rahimi, Amir, et al.
Veröffentlicht: (2023)
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
von: Malik, Sameer, et al.
Veröffentlicht: (2025)
von: Malik, Sameer, et al.
Veröffentlicht: (2025)
Global Context or Local Detail? Adaptive Visual Grounding for Hallucination Mitigation
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
von: Jiang, Yubo, et al.
Veröffentlicht: (2026)
DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval
von: Yang, Yuxin, et al.
Veröffentlicht: (2025)
von: Yang, Yuxin, et al.
Veröffentlicht: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Generative Dataset Distillation: Balancing Global Structure and Local Details
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
von: Li, Longzhen, et al.
Veröffentlicht: (2024)
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
von: Eghbali, Niloufar, et al.
Veröffentlicht: (2025)
von: Eghbali, Niloufar, et al.
Veröffentlicht: (2025)
Generating Accurate and Detailed Captions for High-Resolution Images
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
von: Lee, Hankyeol, et al.
Veröffentlicht: (2025)
MambaBack: Bridging Local Features and Global Contexts in Whole Slide Image Analysis
von: Chen, Sicheng, et al.
Veröffentlicht: (2026)
von: Chen, Sicheng, et al.
Veröffentlicht: (2026)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
von: Chen, Lifeng, et al.
Veröffentlicht: (2025)
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
von: Saichandran, Ketan Suhaas, et al.
Veröffentlicht: (2025)
von: Saichandran, Ketan Suhaas, et al.
Veröffentlicht: (2025)
Dual Relation Alignment for Composed Image Retrieval
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
von: Jiang, Xintong, et al.
Veröffentlicht: (2023)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
von: Sun, Guolei, et al.
Veröffentlicht: (2022)
von: Sun, Guolei, et al.
Veröffentlicht: (2022)
TMCIR: Token Merge Benefits Composed Image Retrieval
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
Scale-Invariant Object Detection by Adaptive Convolution with Unified Global-Local Context
von: Singh, Amrita, et al.
Veröffentlicht: (2024)
von: Singh, Amrita, et al.
Veröffentlicht: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
von: Min, Kyungmin, et al.
Veröffentlicht: (2026)
Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
von: Lai, Zeqiang, et al.
Veröffentlicht: (2025)
von: Lai, Zeqiang, et al.
Veröffentlicht: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
von: Li, Lingen, et al.
Veröffentlicht: (2025)
von: Li, Lingen, et al.
Veröffentlicht: (2025)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
PrimeComposer: Faster Progressively Combined Diffusion for Image Composition with Attention Steering
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives
von: Feng, Zhangchi, et al.
Veröffentlicht: (2024)
von: Feng, Zhangchi, et al.
Veröffentlicht: (2024)
Spherical Linear Interpolation and Text-Anchoring for Zero-shot Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
von: Liu, Mingyu, et al.
Veröffentlicht: (2026)
von: Liu, Mingyu, et al.
Veröffentlicht: (2026)
CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
Enhancing Supervised Composed Image Retrieval via Reasoning-Augmented Representation Engineering
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
DPDEdit: Detail-Preserved Diffusion Models for Multimodal Fashion Image Editing
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolong, et al.
Veröffentlicht: (2024)
Restoring Real-World Images with an Internal Detail Enhancement Diffusion Model
von: Xiao, Peng, et al.
Veröffentlicht: (2025)
von: Xiao, Peng, et al.
Veröffentlicht: (2025)
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
von: Wei, Haoran, et al.
Veröffentlicht: (2024)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2025)
von: Li, Zixu, et al.
Veröffentlicht: (2025)
MELT: Improve Composed Image Retrieval via the Modification Frequentation-Rarity Balance Network
von: Qiu, Guozhi, et al.
Veröffentlicht: (2026)
von: Qiu, Guozhi, et al.
Veröffentlicht: (2026)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
von: Yuan, Peng, et al.
Veröffentlicht: (2026)
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
von: Huang, Wen, et al.
Veröffentlicht: (2025)
von: Huang, Wen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
von: Takemoto, Kentaro, et al.
Veröffentlicht: (2023) -
Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
von: de Margerie, Anatole Jacquin, et al.
Veröffentlicht: (2025) -
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
von: Kolouju, Pranavi, et al.
Veröffentlicht: (2025) -
D3: Data Diversity Design for Systematic Generalization in Visual Question Answering
von: Rahimi, Amir, et al.
Veröffentlicht: (2023) -
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
von: Malik, Sameer, et al.
Veröffentlicht: (2025)