FR-TTS: Test-Time Scaling for NTP-based Image Generation with Effective Filling-based Reward Signal
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Hang, Huang, Linjiang, Zhao, Feng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Fill in the ____ (a Diffusion-based Image Inpainting Pipeline)
by: Gebre, Eyoel, et al.
Published: (2024)
by: Gebre, Eyoel, et al.
Published: (2024)
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
by: He, Runze, et al.
Published: (2024)
by: He, Runze, et al.
Published: (2024)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Video-T1: Test-Time Scaling for Video Generation
by: Liu, Fangfu, et al.
Published: (2025)
by: Liu, Fangfu, et al.
Published: (2025)
RewardFlow: Generate Images by Optimizing What You Reward
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
InfoScale: Unleashing Training-free Variable-scaled Image Generation via Effective Utilization of Information
by: Zhang, Guohui, et al.
Published: (2025)
by: Zhang, Guohui, et al.
Published: (2025)
Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned
by: Ong, Brandon, et al.
Published: (2025)
by: Ong, Brandon, et al.
Published: (2025)
EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback
by: Jia, Jingyang, et al.
Published: (2025)
by: Jia, Jingyang, et al.
Published: (2025)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
by: Wang, Kaishen, et al.
Published: (2025)
by: Wang, Kaishen, et al.
Published: (2025)
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
by: Ren, Tianfei, et al.
Published: (2026)
by: Ren, Tianfei, et al.
Published: (2026)
Scaling Image and Video Generation via Test-Time Evolutionary Search
by: He, Haoran, et al.
Published: (2025)
by: He, Haoran, et al.
Published: (2025)
SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Personalized Reward Modeling for Text-to-Image Generation
by: Lee, Jeongeun, et al.
Published: (2025)
by: Lee, Jeongeun, et al.
Published: (2025)
Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolution
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
by: Ye, Wen, et al.
Published: (2025)
by: Ye, Wen, et al.
Published: (2025)
Self-Corrected Image Generation with Explainable Latent Rewards
by: Luo, Yinyi, et al.
Published: (2026)
by: Luo, Yinyi, et al.
Published: (2026)
Stroke-based Cyclic Amplifier: Image Super-Resolution at Arbitrary Ultra-Large Scales
by: Guo, Wenhao, et al.
Published: (2025)
by: Guo, Wenhao, et al.
Published: (2025)
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
by: Ye, Junyan, et al.
Published: (2025)
by: Ye, Junyan, et al.
Published: (2025)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
by: Lu, Guansong, et al.
Published: (2023)
by: Lu, Guansong, et al.
Published: (2023)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
Retrieval, Refinement, and Ranking for Text-to-Video Generation via Prompt Optimization and Test-Time Scaling
by: Rahman, Zillur, et al.
Published: (2026)
by: Rahman, Zillur, et al.
Published: (2026)
Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale
by: Li, Zhengcen, et al.
Published: (2026)
by: Li, Zhengcen, et al.
Published: (2026)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
NI-Tex: Non-isometric Image-based Garment Texture Generation
by: Shan, Hui, et al.
Published: (2025)
by: Shan, Hui, et al.
Published: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Test-Time Domain Generalization via Universe Learning: A Multi-Graph Matching Approach for Medical Image Segmentation
by: Lv, Xingguo, et al.
Published: (2025)
by: Lv, Xingguo, et al.
Published: (2025)
Force Prompting: Video Generation Models Can Learn and Generalize Physics-based Control Signals
by: Gillman, Nate, et al.
Published: (2025)
by: Gillman, Nate, et al.
Published: (2025)
Surfel-based 3D Registration with Equivariant SE(3) Features
by: Kang, Xueyang, et al.
Published: (2025)
by: Kang, Xueyang, et al.
Published: (2025)
Guiding Diffusion-based Reconstruction with Contrastive Signals for Balanced Visual Representation
by: Han, Boyu, et al.
Published: (2026)
by: Han, Boyu, et al.
Published: (2026)
Sample-Aware Test-Time Adaptation for Medical Image-to-Image Translation
by: Iele, Irene, et al.
Published: (2025)
by: Iele, Irene, et al.
Published: (2025)
Structured Click Control in Transformer-based Interactive Segmentation
by: Xu, Long, et al.
Published: (2024)
by: Xu, Long, et al.
Published: (2024)
METAL: A Multi-Agent Framework for Chart Generation with Test-Time Scaling
by: Li, Bingxuan, et al.
Published: (2025)
by: Li, Bingxuan, et al.
Published: (2025)
Camyla: Scaling Autonomous Research in Medical Image Segmentation
by: Gao, Yifan, et al.
Published: (2026)
by: Gao, Yifan, et al.
Published: (2026)
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
by: Xu, Xinli, et al.
Published: (2024)
by: Xu, Xinli, et al.
Published: (2024)
Compositional Image Synthesis with Inference-Time Scaling
by: Ji, Minsuk, et al.
Published: (2025)
by: Ji, Minsuk, et al.
Published: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
by: Zhang, Zhihong, et al.
Published: (2025)
by: Zhang, Zhihong, et al.
Published: (2025)
Similar Items
-
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
by: Xu, Hang, et al.
Published: (2025) -
FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment
by: Xu, Hang, et al.
Published: (2025) -
Fill in the ____ (a Diffusion-based Image Inpainting Pipeline)
by: Gebre, Eyoel, et al.
Published: (2024) -
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction
by: He, Runze, et al.
Published: (2024) -
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
by: Park, Joonhyung, et al.
Published: (2025)