MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oshima, Yuta, Miyake, Daiki, Matsutani, Kohsei, Iwasawa, Yusuke, Suzuki, Masahiro, Matsuo, Yutaka, Furuta, Hiroki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
von: Oshima, Yuta, et al.
Veröffentlicht: (2025)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026)
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2025)
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2025)
ColorizeDiffusion v2: Enhancing Reference-based Sketch Colorization Through Separating Utilities
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
CLIP-like Model as a Foundational Density Ratio Estimator
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2025)
von: Uchiyama, Fumiya, et al.
Veröffentlicht: (2025)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
von: Oshima, Yuta, et al.
Veröffentlicht: (2024)
Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
von: Kim, Bum Jun, et al.
Veröffentlicht: (2025)
von: Kim, Bum Jun, et al.
Veröffentlicht: (2025)
Realtime Data-Efficient Portrait Stylization Based On Geometric Alignment
von: Wang, Xinrui, et al.
Veröffentlicht: (2022)
von: Wang, Xinrui, et al.
Veröffentlicht: (2022)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
von: Tsuchiya, Fumihiko, et al.
Veröffentlicht: (2026)
von: Tsuchiya, Fumihiko, et al.
Veröffentlicht: (2026)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
von: Furuta, Hiroki, et al.
Veröffentlicht: (2024)
Towards High-resolution and Disentangled Reference-based Sketch Colorization
von: Yan, Dingkun, et al.
Veröffentlicht: (2026)
von: Yan, Dingkun, et al.
Veröffentlicht: (2026)
Image Referenced Sketch Colorization Based on Animation Creation Workflow
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
von: Yan, Dingkun, et al.
Veröffentlicht: (2025)
Leave No Observation Behind: Real-time Correction for VLA Action Chunks
von: Sendai, Kohei, et al.
Veröffentlicht: (2025)
von: Sendai, Kohei, et al.
Veröffentlicht: (2025)
E3VS-Bench: A Benchmark for Viewpoint-Dependent Active Perception in 3D Gaussian Splatting Scenes
von: Sakamoto, Koya, et al.
Veröffentlicht: (2026)
von: Sakamoto, Koya, et al.
Veröffentlicht: (2026)
Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion Models
von: Miyake, Daiki, et al.
Veröffentlicht: (2023)
von: Miyake, Daiki, et al.
Veröffentlicht: (2023)
Topology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph Properties
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
von: Minegishi, Gouki, et al.
Veröffentlicht: (2026)
JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation
von: Xun, Yue, et al.
Veröffentlicht: (2026)
von: Xun, Yue, et al.
Veröffentlicht: (2026)
Learning from Majority Label: A Novel Problem in Multi-class Multiple-Instance Learning
von: Kaito, Shiku, et al.
Veröffentlicht: (2025)
von: Kaito, Shiku, et al.
Veröffentlicht: (2025)
Frequency-Calibrated Membership Inference Attacks on Medical Image Diffusion Models
von: Zhao, Xinkai, et al.
Veröffentlicht: (2025)
von: Zhao, Xinkai, et al.
Veröffentlicht: (2025)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
Inference-time Trajectory Optimization for Manga Image Editing
von: Furuta, Ryosuke
Veröffentlicht: (2026)
von: Furuta, Ryosuke
Veröffentlicht: (2026)
Conditional Text-to-Image Generation with Reference Guidance
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
von: Kim, Taewook, et al.
Veröffentlicht: (2024)
MultiRef: Controllable Image Generation with Multiple Visual References
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
von: Chen, Ruoxi, et al.
Veröffentlicht: (2025)
MMMG: A Massive, Multidisciplinary, Multi-Tier Generation Benchmark for Text-to-Image Reasoning
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
von: Luo, Yuxuan, et al.
Veröffentlicht: (2025)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
von: Xu, Yiyan, et al.
Veröffentlicht: (2026)
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
von: Liang, Xiaoye, et al.
Veröffentlicht: (2026)
von: Liang, Xiaoye, et al.
Veröffentlicht: (2026)
Learning Multi-dimensional Human Preference for Text-to-Image Generation
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
von: Zhang, Sixian, et al.
Veröffentlicht: (2024)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
Instilling Multi-round Thinking to Text-guided Image Generation
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
von: Zeng, Lidong, et al.
Veröffentlicht: (2024)
Benchmarking and Learning Multi-Dimensional Quality Evaluator for Text-to-3D Generation
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
von: Zhang, Yujie, et al.
Veröffentlicht: (2024)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
von: Chen, Zhekai, et al.
Veröffentlicht: (2026)
von: Chen, Zhekai, et al.
Veröffentlicht: (2026)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
von: Oshima, Yuta, et al.
Veröffentlicht: (2025) -
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
von: Oshima, Yuta, et al.
Veröffentlicht: (2025) -
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
von: Oshima, Yuta, et al.
Veröffentlicht: (2024) -
Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2026) -
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
von: Matsutani, Kohsei, et al.
Veröffentlicht: (2025)