Science-T2I: Addressing Scientific Illusions in Image Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Jialuo, Chai, Wenhao, Fu, Xingyu, Xu, Haiyang, Xie, Saining |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VideoNSA: Native Sparse Attention Scales Video Understanding
por: Song, Enxin, et al.
Publicado: (2025)
por: Song, Enxin, et al.
Publicado: (2025)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
por: Li, Jialuo, et al.
Publicado: (2025)
por: Li, Jialuo, et al.
Publicado: (2025)
Transition Matching Distillation for Fast Video Generation
por: Nie, Weili, et al.
Publicado: (2026)
por: Nie, Weili, et al.
Publicado: (2026)
Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI
por: Newen, Carina, et al.
Publicado: (2025)
por: Newen, Carina, et al.
Publicado: (2025)
Improved Baselines with Representation Autoencoders
por: Singh, Jaskirat, et al.
Publicado: (2026)
por: Singh, Jaskirat, et al.
Publicado: (2026)
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
por: Wang, Hao, et al.
Publicado: (2025)
por: Wang, Hao, et al.
Publicado: (2025)
Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
por: Zhang, Haichao, et al.
Publicado: (2023)
por: Zhang, Haichao, et al.
Publicado: (2023)
Leveraging Geometric Visual Illusions as Perceptual Inductive Biases for Vision Models
por: Yang, Haobo, et al.
Publicado: (2025)
por: Yang, Haobo, et al.
Publicado: (2025)
CycleNet: Rethinking Cycle Consistency in Text-Guided Diffusion for Image Manipulation
por: Xu, Sihan, et al.
Publicado: (2023)
por: Xu, Sihan, et al.
Publicado: (2023)
MoDE: CLIP Data Experts via Clustering
por: Ma, Jiawei, et al.
Publicado: (2024)
por: Ma, Jiawei, et al.
Publicado: (2024)
What matters for Representation Alignment: Global Information or Spatial Structure?
por: Singh, Jaskirat, et al.
Publicado: (2025)
por: Singh, Jaskirat, et al.
Publicado: (2025)
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
por: Chu, Tianzhe, et al.
Publicado: (2025)
por: Chu, Tianzhe, et al.
Publicado: (2025)
T2I-ConBench: Text-to-Image Benchmark for Continual Post-training
por: Huang, Zhehao, et al.
Publicado: (2025)
por: Huang, Zhehao, et al.
Publicado: (2025)
Label-free Neural Semantic Image Synthesis
por: Wang, Jiayi, et al.
Publicado: (2024)
por: Wang, Jiayi, et al.
Publicado: (2024)
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
por: Jiang, Zhongyu, et al.
Publicado: (2025)
por: Jiang, Zhongyu, et al.
Publicado: (2025)
The Illusion of Forgetting: Attack Unlearned Diffusion via Initial Latent Variable Optimization
por: Li, Manyi, et al.
Publicado: (2026)
por: Li, Manyi, et al.
Publicado: (2026)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
por: Berrada, Tariq, et al.
Publicado: (2023)
por: Berrada, Tariq, et al.
Publicado: (2023)
Addressing Negative Transfer in Diffusion Models
por: Go, Hyojun, et al.
Publicado: (2023)
por: Go, Hyojun, et al.
Publicado: (2023)
Generalizable Geometric Image Caption Synthesis
por: Xin, Yue, et al.
Publicado: (2025)
por: Xin, Yue, et al.
Publicado: (2025)
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
por: Tang, Lv, et al.
Publicado: (2026)
por: Tang, Lv, et al.
Publicado: (2026)
ATLAS: Adapter-Based Multi-Modal Continual Learning with a Two-Stage Learning Strategy
por: Li, Hong, et al.
Publicado: (2024)
por: Li, Hong, et al.
Publicado: (2024)
Accessing Vision Foundation Models via ImageNet-1K
por: Zhang, Yitian, et al.
Publicado: (2024)
por: Zhang, Yitian, et al.
Publicado: (2024)
Adaptive Training Meets Progressive Scaling: Elevating Efficiency in Diffusion Models
por: Li, Wenhao, et al.
Publicado: (2023)
por: Li, Wenhao, et al.
Publicado: (2023)
H$_{2}$OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
por: Li, Wenhao, et al.
Publicado: (2025)
por: Li, Wenhao, et al.
Publicado: (2025)
Improving Diffusion-Based Image Synthesis with Context Prediction
por: Yang, Ling, et al.
Publicado: (2024)
por: Yang, Ling, et al.
Publicado: (2024)
ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation
por: Zhang, Yi, et al.
Publicado: (2024)
por: Zhang, Yi, et al.
Publicado: (2024)
ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion
por: Fang, Zhenghan, et al.
Publicado: (2025)
por: Fang, Zhenghan, et al.
Publicado: (2025)
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
por: Kao, Kuei-Chun, et al.
Publicado: (2026)
por: Kao, Kuei-Chun, et al.
Publicado: (2026)
High-Resolution Image Synthesis via Next-Token Prediction
por: Chen, Dengsheng, et al.
Publicado: (2024)
por: Chen, Dengsheng, et al.
Publicado: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
por: Zhou, Guanyu, et al.
Publicado: (2026)
por: Zhou, Guanyu, et al.
Publicado: (2026)
Beyond Visual Fidelity: Benchmarking Super-Resolution Models for Large-Scale Remote Sensing Imagery via Downstream Task Integration
por: Li, Zhili, et al.
Publicado: (2026)
por: Li, Zhili, et al.
Publicado: (2026)
SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
por: Ye, Hanrong, et al.
Publicado: (2023)
por: Ye, Hanrong, et al.
Publicado: (2023)
When are Foundation Models Effective? Understanding the Suitability for Pixel-Level Classification Using Multispectral Imagery
por: Xie, Yiqun, et al.
Publicado: (2024)
por: Xie, Yiqun, et al.
Publicado: (2024)
BIP3D: Bridging 2D Images and 3D Perception for Embodied Intelligence
por: Lin, Xuewu, et al.
Publicado: (2024)
por: Lin, Xuewu, et al.
Publicado: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
por: Xiong, Tianwei, et al.
Publicado: (2024)
por: Xiong, Tianwei, et al.
Publicado: (2024)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
por: Izadi, Amirmohammad, et al.
Publicado: (2025)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
por: Gu, Jiatao, et al.
Publicado: (2025)
por: Gu, Jiatao, et al.
Publicado: (2025)
Improved Sub-Visible Particle Classification in Flow Imaging Microscopy via Generative AI-Based Image Synthesis
por: Ozbulak, Utku, et al.
Publicado: (2025)
por: Ozbulak, Utku, et al.
Publicado: (2025)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
por: Zhou, Wenhao, et al.
Publicado: (2025)
por: Zhou, Wenhao, et al.
Publicado: (2025)
Leveraging Image Generators to Address Training Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
por: Jeanson, Gabriel, et al.
Publicado: (2026)
por: Jeanson, Gabriel, et al.
Publicado: (2026)
Ejemplares similares
-
VideoNSA: Native Sparse Attention Scales Video Understanding
por: Song, Enxin, et al.
Publicado: (2025) -
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
por: Li, Jialuo, et al.
Publicado: (2025) -
Transition Matching Distillation for Fast Video Generation
por: Nie, Weili, et al.
Publicado: (2026) -
Do you see what I see? An Ambiguous Optical Illusion Dataset exposing limitations of Explainable AI
por: Newen, Carina, et al.
Publicado: (2025) -
Improved Baselines with Representation Autoencoders
por: Singh, Jaskirat, et al.
Publicado: (2026)