Salvato in:
| Autori principali: | Rahman, Zillur, Sheng, Alex, Meo, Cristian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.01509 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
di: Wu, Shang, et al.
Pubblicazione: (2026)
di: Wu, Shang, et al.
Pubblicazione: (2026)
Video-T1: Test-Time Scaling for Video Generation
di: Liu, Fangfu, et al.
Pubblicazione: (2025)
di: Liu, Fangfu, et al.
Pubblicazione: (2025)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
di: Paul, Dhiman, et al.
Pubblicazione: (2024)
di: Paul, Dhiman, et al.
Pubblicazione: (2024)
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
di: Jeong, Suchae, et al.
Pubblicazione: (2025)
di: Jeong, Suchae, et al.
Pubblicazione: (2025)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
di: Wang, Ni, et al.
Pubblicazione: (2024)
di: Wang, Ni, et al.
Pubblicazione: (2024)
Dynamic Prompt Optimizing for Text-to-Image Generation
di: Mo, Wenyi, et al.
Pubblicazione: (2024)
di: Mo, Wenyi, et al.
Pubblicazione: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
di: He, Haoran, et al.
Pubblicazione: (2025)
di: He, Haoran, et al.
Pubblicazione: (2025)
MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
di: Han, Donghoon, et al.
Pubblicazione: (2024)
di: Han, Donghoon, et al.
Pubblicazione: (2024)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
di: Rahman, Kazi Mahathir, et al.
Pubblicazione: (2025)
di: Rahman, Kazi Mahathir, et al.
Pubblicazione: (2025)
GenPilot: A Multi-Agent System for Test-Time Prompt Optimization in Image Generation
di: Ye, Wen, et al.
Pubblicazione: (2025)
di: Ye, Wen, et al.
Pubblicazione: (2025)
TimeRefine: Temporal Grounding with Time Refining Video LLM
di: Wang, Xizi, et al.
Pubblicazione: (2024)
di: Wang, Xizi, et al.
Pubblicazione: (2024)
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
di: Kim, Subin, et al.
Pubblicazione: (2025)
di: Kim, Subin, et al.
Pubblicazione: (2025)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
di: Lee, Daeun, et al.
Pubblicazione: (2024)
di: Lee, Daeun, et al.
Pubblicazione: (2024)
Minority-Focused Text-to-Image Generation via Prompt Optimization
di: Um, Soobin, et al.
Pubblicazione: (2024)
di: Um, Soobin, et al.
Pubblicazione: (2024)
Long-Text-to-Image Generation via Compositional Prompt Decomposition
di: Huang, Jen-Yuan, et al.
Pubblicazione: (2026)
di: Huang, Jen-Yuan, et al.
Pubblicazione: (2026)
PrismVAU: Prompt-Refined Inference System for Multimodal Video Anomaly Understanding
di: Erregue, Iñaki, et al.
Pubblicazione: (2026)
di: Erregue, Iñaki, et al.
Pubblicazione: (2026)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
di: Xue, Qiyao, et al.
Pubblicazione: (2024)
di: Xue, Qiyao, et al.
Pubblicazione: (2024)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
di: Zhou, Yinan, et al.
Pubblicazione: (2025)
di: Zhou, Yinan, et al.
Pubblicazione: (2025)
Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval
di: Yang, Junkai, et al.
Pubblicazione: (2026)
di: Yang, Junkai, et al.
Pubblicazione: (2026)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
di: Tzachor, Issar, et al.
Pubblicazione: (2026)
di: Tzachor, Issar, et al.
Pubblicazione: (2026)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
di: Park, Joonhyung, et al.
Pubblicazione: (2025)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
di: Yang, Bowen, et al.
Pubblicazione: (2025)
di: Yang, Bowen, et al.
Pubblicazione: (2025)
VC4VG: Optimizing Video Captions for Text-to-Video Generation
di: Du, Yang, et al.
Pubblicazione: (2025)
di: Du, Yang, et al.
Pubblicazione: (2025)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
di: Cho, CH, et al.
Pubblicazione: (2025)
di: Cho, CH, et al.
Pubblicazione: (2025)
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
di: Xu, Hang, et al.
Pubblicazione: (2025)
di: Xu, Hang, et al.
Pubblicazione: (2025)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
di: Yang, Xiaomeng, et al.
Pubblicazione: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
di: Menapace, Willi, et al.
Pubblicazione: (2024)
di: Menapace, Willi, et al.
Pubblicazione: (2024)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
di: Zeng, Gangyan, et al.
Pubblicazione: (2024)
Video-As-Prompt: Unified Semantic Control for Video Generation
di: Bian, Yuxuan, et al.
Pubblicazione: (2025)
di: Bian, Yuxuan, et al.
Pubblicazione: (2025)
Singular Value Scaling: Efficient Generative Model Compression via Pruned Weights Refinement
di: Kim, Hyeonjin, et al.
Pubblicazione: (2024)
di: Kim, Hyeonjin, et al.
Pubblicazione: (2024)
Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
di: Kang, Peng, et al.
Pubblicazione: (2025)
di: Kang, Peng, et al.
Pubblicazione: (2025)
VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
di: Yesiltepe, Hidir, et al.
Pubblicazione: (2026)
di: Yesiltepe, Hidir, et al.
Pubblicazione: (2026)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
di: Ji, Longbin, et al.
Pubblicazione: (2026)
di: Ji, Longbin, et al.
Pubblicazione: (2026)
DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation
di: Nithisopa, Naphat, et al.
Pubblicazione: (2025)
di: Nithisopa, Naphat, et al.
Pubblicazione: (2025)
Batch-Instructed Gradient for Prompt Evolution:Systematic Prompt Optimization for Enhanced Text-to-Image Synthesis
di: Yang, Xinrui, et al.
Pubblicazione: (2024)
di: Yang, Xinrui, et al.
Pubblicazione: (2024)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
di: Wu, Mingrui, et al.
Pubblicazione: (2025)
di: Wu, Mingrui, et al.
Pubblicazione: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
di: Gao, Bingjie, et al.
Pubblicazione: (2025)
Text-to-Image Diffusion Models Cannot Count, and Prompt Refinement Cannot Help
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
di: Guo, Xuyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
di: Wu, Shang, et al.
Pubblicazione: (2026) -
Video-T1: Test-Time Scaling for Video Generation
di: Liu, Fangfu, et al.
Pubblicazione: (2025) -
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
di: Paul, Dhiman, et al.
Pubblicazione: (2024) -
Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement
di: Jeong, Suchae, et al.
Pubblicazione: (2025) -
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
di: Wang, Ni, et al.
Pubblicazione: (2024)