Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Jinyan, Duan, Zhongjie, Li, Zhiwen, Chen, Cen, Chen, Daoyuan, Li, Yaliang, Chen, Yingda |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
by: Li, Zhiwen, et al.
Published: (2026)
by: Li, Zhiwen, et al.
Published: (2026)
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
by: Li, Zhiwen, et al.
Published: (2025)
by: Li, Zhiwen, et al.
Published: (2025)
AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
Comprehensive Assessment and Analysis for NSFW Content Erasure in Text-to-Image Diffusion Models
by: Chen, Die, et al.
Published: (2025)
by: Chen, Die, et al.
Published: (2025)
EliGen: Entity-Level Controlled Image Generation with Regional Attention
by: Zhang, Hong, et al.
Published: (2025)
by: Zhang, Hong, et al.
Published: (2025)
Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion Models
by: Chen, Die, et al.
Published: (2024)
by: Chen, Die, et al.
Published: (2024)
Diffusion Templates: A Unified Plugin Framework for Controllable Diffusion
by: Duan, Zhongjie, et al.
Published: (2026)
by: Duan, Zhongjie, et al.
Published: (2026)
VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
by: Li, Yuyi, et al.
Published: (2025)
by: Li, Yuyi, et al.
Published: (2025)
DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?
by: Jiao, Qirui, et al.
Published: (2025)
by: Jiao, Qirui, et al.
Published: (2025)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
Diffutoon: High-Resolution Editable Toon Shading via Diffusion Models
by: Duan, Zhongjie, et al.
Published: (2024)
by: Duan, Zhongjie, et al.
Published: (2024)
MindGYM: What Matters in Question Synthesis for Thinking-Centric Fine-Tuning?
by: Xu, Zhe, et al.
Published: (2025)
by: Xu, Zhe, et al.
Published: (2025)
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
by: Zhang, Hong, et al.
Published: (2025)
by: Zhang, Hong, et al.
Published: (2025)
Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
by: Jiao, Qirui, et al.
Published: (2024)
by: Jiao, Qirui, et al.
Published: (2024)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
Transferability Bound Theory: Exploring Relationship between Adversarial Transferability and Flatness
by: Fan, Mingyuan, et al.
Published: (2023)
by: Fan, Mingyuan, et al.
Published: (2023)
A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
by: Yang, Shentao, et al.
Published: (2024)
by: Yang, Shentao, et al.
Published: (2024)
RewardDance: Reward Scaling in Visual Generation
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025)
by: Xie, Enze, et al.
Published: (2025)
Diffusion-Classifier Synergy: Reward-Aligned Learning via Mutual Boosting Loop for FSCIL
by: Wu, Ruitao, et al.
Published: (2025)
by: Wu, Ruitao, et al.
Published: (2025)
Data-Juicer Sandbox: A Feedback-Driven Suite for Multimodal Data-Model Co-development
by: Chen, Daoyuan, et al.
Published: (2024)
by: Chen, Daoyuan, et al.
Published: (2024)
Multi-dimensional Visual Prompt Enhanced Image Restoration via Mamba-Transformer Aggregation
by: Jiang, Aiwen, et al.
Published: (2024)
by: Jiang, Aiwen, et al.
Published: (2024)
Bidirectional Prototype-Reward co-Evolution for Test-Time Adaptation of Vision-Language Models
by: Qiao, Xiaozhen, et al.
Published: (2025)
by: Qiao, Xiaozhen, et al.
Published: (2025)
SVGThinker: Instruction-Aligned and Reasoning-Driven Text-to-SVG Generation
by: Chen, Hanqi, et al.
Published: (2025)
by: Chen, Hanqi, et al.
Published: (2025)
RefReward-SR: LR-Conditioned Reward Modeling for Preference-Aligned Super-Resolution
by: Song, Yushuai, et al.
Published: (2026)
by: Song, Yushuai, et al.
Published: (2026)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
by: Li, Zizun, et al.
Published: (2026)
by: Li, Zizun, et al.
Published: (2026)
Physics-Aligned Spectral Mamba: Decoupling Semantics and Dynamics for Few-Shot Hyperspectral Target Detection
by: Gong, Luqi, et al.
Published: (2026)
by: Gong, Luqi, et al.
Published: (2026)
Hyperspectral Image Classification via Efficient Global Spectral Supertoken Clustering
by: Liu, Peifu, et al.
Published: (2026)
by: Liu, Peifu, et al.
Published: (2026)
Enhancing Spatial Understanding in Image Generation via Reward Modeling
by: Tang, Zhenyu, et al.
Published: (2026)
by: Tang, Zhenyu, et al.
Published: (2026)
PoreTrack3D: A Benchmark for Dynamic 3D Gaussian Splatting in Pore-Scale Facial Trajectory Tracking
by: Li, Dong, et al.
Published: (2025)
by: Li, Dong, et al.
Published: (2025)
Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM
by: Li, Shen, et al.
Published: (2025)
by: Li, Shen, et al.
Published: (2025)
The Synergy between Data and Multi-Modal Large Language Models: A Survey from Co-Development Perspective
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation
by: Fu, Minghao, et al.
Published: (2025)
by: Fu, Minghao, et al.
Published: (2025)
AlignedGen: Aligning Style Across Generated Images
by: Zhang, Jiexuan, et al.
Published: (2025)
by: Zhang, Jiexuan, et al.
Published: (2025)
Similar Items
-
VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
by: Li, Zhiwen, et al.
Published: (2026) -
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
by: Li, Zhiwen, et al.
Published: (2025) -
AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
by: Chen, Die, et al.
Published: (2025) -
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
by: Duan, Zhongjie, et al.
Published: (2024) -
Comprehensive Evaluation and Analysis for NSFW Concept Erasure in Text-to-Image Diffusion Models
by: Chen, Die, et al.
Published: (2025)