Gespeichert in:
| Hauptverfasser: | Chen, Muxi, Liu, Yi, Yi, Jian, Xu, Changran, Lai, Qiuxia, Wang, Hongliang, Ho, Tsung-Yi, Xu, Qiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.05125 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RIGID: A Training-free and Model-Agnostic Framework for Robust AI-Generated Image Detection
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
von: He, Zhiyuan, et al.
Veröffentlicht: (2024)
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging
von: Chen, Muxi, et al.
Veröffentlicht: (2025)
von: Chen, Muxi, et al.
Veröffentlicht: (2025)
Unsupervised Out-of-Distribution Detection in Medical Imaging Using Multi-Exit Class Activation Maps and Feature Masking
von: Chen, Yu-Jen, et al.
Veröffentlicht: (2025)
von: Chen, Yu-Jen, et al.
Veröffentlicht: (2025)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
MMA-Diffusion: MultiModal Attack on Diffusion Models
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
von: Yang, Yijun, et al.
Veröffentlicht: (2023)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
von: Wang, Xinran, et al.
Veröffentlicht: (2025)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
Temporal Consistency-Aware Text-to-Motion Generation
von: Wang, Hongsong, et al.
Veröffentlicht: (2026)
von: Wang, Hongsong, et al.
Veröffentlicht: (2026)
3DIS: Depth-Driven Decoupled Instance Synthesis for Text-to-Image Generation
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
von: Lin, Tiancheng, et al.
Veröffentlicht: (2024)
Vector Quantization Prompting for Continual Learning
von: Jiao, Li, et al.
Veröffentlicht: (2024)
von: Jiao, Li, et al.
Veröffentlicht: (2024)
PhyCustom: Towards Realistic Physical Customization in Text-to-Image Generation
von: Wu, Fan, et al.
Veröffentlicht: (2025)
von: Wu, Fan, et al.
Veröffentlicht: (2025)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
An Empirical Study and Analysis of Text-to-Image Generation Using Large Language Model-Powered Textual Representation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2024)
Text-IF: Leveraging Semantic Text Guidance for Degradation-Aware and Interactive Image Fusion
von: Yi, Xunpeng, et al.
Veröffentlicht: (2024)
von: Yi, Xunpeng, et al.
Veröffentlicht: (2024)
Qwen-Image-Bench: From Generation to Creation in Text-to-Image Evaluation
von: Li, Niantong, et al.
Veröffentlicht: (2026)
von: Li, Niantong, et al.
Veröffentlicht: (2026)
Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model
von: Yi, Mingyang, et al.
Veröffentlicht: (2024)
von: Yi, Mingyang, et al.
Veröffentlicht: (2024)
Fine-grained Text to Image Synthesis
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
von: Ouyang, Xu, et al.
Veröffentlicht: (2024)
Information Bottleneck Approach to Spatial Attention Learning
von: Lai, Qiuxia, et al.
Veröffentlicht: (2021)
von: Lai, Qiuxia, et al.
Veröffentlicht: (2021)
HyperVQ: Enabling Hyperprior Entropy Modeling for VQ-Based Generative Image Compression
von: Yi, Niu, et al.
Veröffentlicht: (2025)
von: Yi, Niu, et al.
Veröffentlicht: (2025)
MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
von: Xu, Meng, et al.
Veröffentlicht: (2024)
von: Xu, Meng, et al.
Veröffentlicht: (2024)
HumanCoser: Layered 3D Human Generation via Semantic-Aware Diffusion Model
von: Wang, Yi, et al.
Veröffentlicht: (2024)
von: Wang, Yi, et al.
Veröffentlicht: (2024)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
Multi-rater Prompting for Ambiguous Medical Image Segmentation
von: Wang, Jinhong, et al.
Veröffentlicht: (2024)
von: Wang, Jinhong, et al.
Veröffentlicht: (2024)
Origin Identification for Text-Guided Image-to-Image Diffusion Models
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
von: Wang, Wenhao, et al.
Veröffentlicht: (2025)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
von: Cui, Tianyu, et al.
Veröffentlicht: (2025)
TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2025)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2025)
TIPO: Text to Image with Text Presampling for Prompt Optimization
von: Yeh, Shih-Ying, et al.
Veröffentlicht: (2024)
von: Yeh, Shih-Ying, et al.
Veröffentlicht: (2024)
Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2024)
von: Han, Jian, et al.
Veröffentlicht: (2024)
AnomalyDiffusion: Few-Shot Anomaly Image Generation with Diffusion Model
von: Hu, Teng, et al.
Veröffentlicht: (2023)
von: Hu, Teng, et al.
Veröffentlicht: (2023)
Layered 3D Human Generation via Semantic-Aware Diffusion Model
von: Wang, Yi, et al.
Veröffentlicht: (2023)
von: Wang, Yi, et al.
Veröffentlicht: (2023)
Iterative Online Image Synthesis via Diffusion Model for Imbalanced Classification
von: Li, Shuhan, et al.
Veröffentlicht: (2024)
von: Li, Shuhan, et al.
Veröffentlicht: (2024)
EmotiCrafter: Text-to-Emotional-Image Generation based on Valence-Arousal Model
von: Dang, Shengqi, et al.
Veröffentlicht: (2025)
von: Dang, Shengqi, et al.
Veröffentlicht: (2025)
Sparsity Meets Similarity: Leveraging Long-Tail Distribution for Dynamic Optimized Token Representation in Multimodal Large Language Models
von: Yu, Gaotong, et al.
Veröffentlicht: (2024)
von: Yu, Gaotong, et al.
Veröffentlicht: (2024)
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
Intriguing Properties of Diffusion Models: An Empirical Study of the Natural Attack Capability in Text-to-Image Generative Models
von: Sato, Takami, et al.
Veröffentlicht: (2023)
von: Sato, Takami, et al.
Veröffentlicht: (2023)
E-VRAG: Enhancing Long Video Understanding with Resource-Efficient Retrieval Augmented Generation
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
von: Xu, Zeyu, et al.
Veröffentlicht: (2025)
FBCIR: Balancing Cross-Modal Focuses in Composed Image Retrieval
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
von: Yu, Haoran, et al.
Veröffentlicht: (2025)
von: Yu, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RIGID: A Training-free and Model-Agnostic Framework for Robust AI-Generated Image Detection
von: He, Zhiyuan, et al.
Veröffentlicht: (2024) -
\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions
von: Zhao, Chenchen, et al.
Veröffentlicht: (2026) -
HiBug2: Efficient and Interpretable Error Slice Discovery for Comprehensive Model Debugging
von: Chen, Muxi, et al.
Veröffentlicht: (2025) -
Unsupervised Out-of-Distribution Detection in Medical Imaging Using Multi-Exit Class Activation Maps and Feature Masking
von: Chen, Yu-Jen, et al.
Veröffentlicht: (2025) -
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
von: Meng, Chutian, et al.
Veröffentlicht: (2024)