Adaptively Clustering Neighbor Elements for Image-Text Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zihua, Yang, Xu, Zhang, Hanwang, Xu, Haiyang, Yan, Ming, Huang, Fei, Zhang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient and Effective In-context Demonstration Selection with Coreset
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
von: Qian, Qi, et al.
Veröffentlicht: (2024)
von: Qian, Qi, et al.
Veröffentlicht: (2024)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
von: Yu, Lu, et al.
Veröffentlicht: (2026)
von: Yu, Lu, et al.
Veröffentlicht: (2026)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)
Focus on Neighbors and Know the Whole: Towards Consistent Dense Multiview Text-to-Image Generator for 3D Creation
von: Li, Bonan, et al.
Veröffentlicht: (2024)
von: Li, Bonan, et al.
Veröffentlicht: (2024)
Enhancing Adaptive Deep Networks for Image Classification via Uncertainty-aware Decision Fusion
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
Unsupervised Image Classification with Adaptive Nearest Neighbor Selection and Cluster Ensembles
von: Baydar, Melih, et al.
Veröffentlicht: (2025)
von: Baydar, Melih, et al.
Veröffentlicht: (2025)
Deep Image Clustering Based on Curriculum Learning and Density Information
von: Zheng, Haiyang, et al.
Veröffentlicht: (2026)
von: Zheng, Haiyang, et al.
Veröffentlicht: (2026)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
von: Zhang, Liang, et al.
Veröffentlicht: (2024)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
von: Wang, Junyang, et al.
Veröffentlicht: (2025)
von: Wang, Junyang, et al.
Veröffentlicht: (2025)
L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
von: Li, Li, et al.
Veröffentlicht: (2025)
von: Li, Li, et al.
Veröffentlicht: (2025)
Doubly Abductive Counterfactual Inference for Text-based Image Editing
von: Song, Xue, et al.
Veröffentlicht: (2024)
von: Song, Xue, et al.
Veröffentlicht: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Few-shot NeRF by Adaptive Rendering Loss Regularization
von: Xu, Qingshan, et al.
Veröffentlicht: (2024)
von: Xu, Qingshan, et al.
Veröffentlicht: (2024)
Diffusion Time-step Curriculum for One Image to 3D Generation
von: Yi, Xuanyu, et al.
Veröffentlicht: (2024)
von: Yi, Xuanyu, et al.
Veröffentlicht: (2024)
Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Mobile-Agent-E: Self-Evolving Mobile Assistant for Complex Tasks
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
DragNeXt: Rethinking Drag-Based Image Editing
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
EMMA: Your Text-to-Image Diffusion Model Can Secretly Accept Multi-Modal Prompts
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
von: Han, Yucheng, et al.
Veröffentlicht: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation
von: He, Qingdong, et al.
Veröffentlicht: (2024)
von: He, Qingdong, et al.
Veröffentlicht: (2024)
ConText: Driving In-context Learning for Text Removal and Segmentation
von: Zhang, Fei, et al.
Veröffentlicht: (2025)
von: Zhang, Fei, et al.
Veröffentlicht: (2025)
MIBench: Evaluating Multimodal Large Language Models over Multiple Images
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
Hierarchical Semantic Alignment for Image Clustering
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
Personalize Your Gaussian: Consistent 3D Scene Personalization from a Single Image
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
von: Wang, Zihua, et al.
Veröffentlicht: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
TDEC: Deep Embedded Image Clustering with Transformer and Distribution Information
von: Zhang, Ruilin, et al.
Veröffentlicht: (2026)
von: Zhang, Ruilin, et al.
Veröffentlicht: (2026)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Component Adaptive Clustering for Generalized Category Discovery
von: Yan, Mingfu, et al.
Veröffentlicht: (2025)
von: Yan, Mingfu, et al.
Veröffentlicht: (2025)
SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
von: Jia, Hongrui, et al.
Veröffentlicht: (2024)
mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
von: Hu, Anwen, et al.
Veröffentlicht: (2024)
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
von: Liu, Haowei, et al.
Veröffentlicht: (2025)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
von: Liu, Haowei, et al.
Veröffentlicht: (2024)
Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
von: Zhang, Xu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient and Effective In-context Demonstration Selection with Coreset
von: Wang, Zihua, et al.
Veröffentlicht: (2025) -
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
von: Qian, Qi, et al.
Veröffentlicht: (2024) -
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024) -
Complementary Text-Guided Attention for Zero-Shot Adversarial Robustness
von: Yu, Lu, et al.
Veröffentlicht: (2026) -
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
von: Zeng, Guanning, et al.
Veröffentlicht: (2025)