CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Tong, Chengzhuo, Chang, Mingkun, Zhang, Shenglong, Wang, Yuran, Liang, Cheng, Zhao, Zhizheng, An, Ruichuan, Zeng, Bohan, Shi, Yang, Dai, Yifan, Zhao, Ziming, Li, Guanbin, Wan, Pengfei, Zhang, Yuanxing, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
by: Qi, Yu, et al.
Published: (2026)
by: Qi, Yu, et al.
Published: (2026)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
yuzhimanhua/CoF: v1.0
by: Yu Zhang
Published: (2025)
by: Yu Zhang
Published: (2025)
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks
by: Zeng, Bohan, et al.
Published: (2026)
by: Zeng, Bohan, et al.
Published: (2026)
CoF3: a g-wave Altermagnet
by: Tagani, Meysam Bagheri
Published: (2024)
by: Tagani, Meysam Bagheri
Published: (2024)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
Polar phonons and magnetic excitations in the antiferromagnet CoF$_2$
by: Dubrovin, R. M., et al.
Published: (2024)
by: Dubrovin, R. M., et al.
Published: (2024)
Impulsive Fermi magnon-phonon resonance in antiferromagnetic $CoF_{2}$
by: Metzger, Thomas W. J., et al.
Published: (2023)
by: Metzger, Thomas W. J., et al.
Published: (2023)
Exposing Image Classifier Shortcuts with Counterfactual Frequency (CoF) Tables
by: Hinns, James, et al.
Published: (2024)
by: Hinns, James, et al.
Published: (2024)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
by: Shen, Zijun, et al.
Published: (2026)
by: Shen, Zijun, et al.
Published: (2026)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
by: Wang, Yeyuan, et al.
Published: (2024)
by: Wang, Yeyuan, et al.
Published: (2024)
Synthesis and Performance of Multifunctional Cobalt‐Doped Polydopamine‐Derived Carbon‐Based Electrocatalysts
by: Chengzhuo Xiao, et al.
Published: (2026)
by: Chengzhuo Xiao, et al.
Published: (2026)
SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
PI-CoF: A Bilevel Optimization Framework for Solving Active Learning Problems using Physics-Information
by: Dong, Liqiu, et al.
Published: (2024)
by: Dong, Liqiu, et al.
Published: (2024)
Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
by: Li, Bozhou, et al.
Published: (2026)
by: Li, Bozhou, et al.
Published: (2026)
Giant intrinsic nonlinear phonon-magnon coupling in the antiferromagnet CoF$_2$
by: Prosnikov, M. A., et al.
Published: (2024)
by: Prosnikov, M. A., et al.
Published: (2024)
Derived logarithmic deformation theory and moduli stacks of derived logarithmic structures
by: Zhang, Ruichuan
Published: (2026)
by: Zhang, Ruichuan
Published: (2026)
Earthworm‐Inspired Co/Co3O4/CoF2@NSC Nanofibrous Electrocatalyst with Confined Channels for Enhanced ORR/OER Performance
by: Han Li, et al.
Published: (2024)
by: Han Li, et al.
Published: (2024)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models
by: DataFlow Team, et al.
Published: (2026)
by: DataFlow Team, et al.
Published: (2026)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
by: Wang, Qunzhong, et al.
Published: (2025)
by: Wang, Qunzhong, et al.
Published: (2025)
PEARL: Personalized Streaming Video Understanding Model
by: Zheng, Yuanhong, et al.
Published: (2026)
by: Zheng, Yuanhong, et al.
Published: (2026)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
by: Li, Yuying, et al.
Published: (2025)
by: Li, Yuying, et al.
Published: (2025)
Representational power of selected neural network quantum states in second quantization
by: Li, Zhendong, et al.
Published: (2025)
by: Li, Zhendong, et al.
Published: (2025)
Ligand field and interference effects in L-edge x-ray raman scattering of MnF2 and CoF2
by: J. Jiménez-Mier
Published: (2008)
by: J. Jiménez-Mier
Published: (2008)
Ligand field and interference effects in L-edge x-ray Raman scattering of MnF2 and CoF2
by: J. Jiménez-Mier
Published: (2008)
by: J. Jiménez-Mier
Published: (2008)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos
by: Feng, Hengyi, et al.
Published: (2026)
by: Feng, Hengyi, et al.
Published: (2026)
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
ExCoT: Optimizing Reasoning for Text-to-SQL with Execution Feedback
by: Zhai, Bohan, et al.
Published: (2025)
by: Zhai, Bohan, et al.
Published: (2025)
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
by: Liang, Zhijia, et al.
Published: (2026)
by: Liang, Zhijia, et al.
Published: (2026)
VGGT-Edit: Feed-forward Native 3D Scene Editing with Residual Field Prediction
by: Zhu, Kaixin, et al.
Published: (2026)
by: Zhu, Kaixin, et al.
Published: (2026)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
by: Lin, Weifeng, et al.
Published: (2025)
by: Lin, Weifeng, et al.
Published: (2025)
A Deep Learning Framework with Geographic Information Adaptive Loss for Remote Sensing Images based UAV Self-Positioning
by: Li, Mingkun, et al.
Published: (2025)
by: Li, Mingkun, et al.
Published: (2025)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Similar Items
-
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
by: Qi, Yu, et al.
Published: (2026) -
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
by: Guo, Ziyu, et al.
Published: (2025) -
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025) -
yuzhimanhua/CoF: v1.0
by: Yu Zhang
Published: (2025) -
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)