Saved in:
| Main Authors: | Ma, Shijie, Xu, Huayi, Li, Mengjian, Geng, Weidong, Wang, Yaxiong, Wang, Meng |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.00949 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Geometric-Photometric Joint Alignment for Facial Mesh Registration
by: Wang, Xizhi, et al.
Published: (2024)
by: Wang, Xizhi, et al.
Published: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
by: Jiang, Xintong, et al.
Published: (2024)
by: Jiang, Xintong, et al.
Published: (2024)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
by: Huang, Ziqi, et al.
Published: (2024)
by: Huang, Ziqi, et al.
Published: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
TDEdit: A Unified Diffusion Framework for Text-Drag Guided Image Manipulation
by: Wang, Qihang, et al.
Published: (2025)
by: Wang, Qihang, et al.
Published: (2025)
RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
by: Wang, Yaxiong, et al.
Published: (2024)
by: Wang, Yaxiong, et al.
Published: (2024)
Consistency-aware Fake Videos Detection on Short Video Platforms
by: Wang, Junxi, et al.
Published: (2025)
by: Wang, Junxi, et al.
Published: (2025)
Universal Prompt Optimizer for Safe Text-to-Image Generation
by: Wu, Zongyu, et al.
Published: (2024)
by: Wu, Zongyu, et al.
Published: (2024)
Text-Driven Diffusion Model for Sign Language Production
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
by: Chen, Guo, et al.
Published: (2024)
by: Chen, Guo, et al.
Published: (2024)
Optimizing Prompts for Text-to-Image Generation
by: Hao, Yaru, et al.
Published: (2022)
by: Hao, Yaru, et al.
Published: (2022)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Minimizing the Pretraining Gap: Domain-aligned Text-Based Person Retrieval
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
by: Huang, Jinghao, et al.
Published: (2025)
by: Huang, Jinghao, et al.
Published: (2025)
Scale Up Composed Image Retrieval Learning via Modification Text Generation
by: Zhou, Yinan, et al.
Published: (2025)
by: Zhou, Yinan, et al.
Published: (2025)
TIPO: Text to Image with Text Presampling for Prompt Optimization
by: Yeh, Shih-Ying, et al.
Published: (2024)
by: Yeh, Shih-Ying, et al.
Published: (2024)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
by: Ma, Wufei, et al.
Published: (2024)
by: Ma, Wufei, et al.
Published: (2024)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
by: Liu, Lingyu, et al.
Published: (2026)
by: Liu, Lingyu, et al.
Published: (2026)
SoPo: Text-to-Motion Generation Using Semi-Online Preference Optimization
by: Tan, Xiaofeng, et al.
Published: (2024)
by: Tan, Xiaofeng, et al.
Published: (2024)
VTD-CLIP: Video-to-Text Discretization via Prompting CLIP
by: Zhu, Wencheng, et al.
Published: (2025)
by: Zhu, Wencheng, et al.
Published: (2025)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
Every Painting Awakened: A Training-free Framework for Painting-to-Animation Generation
by: Liu, Lingyu, et al.
Published: (2025)
by: Liu, Lingyu, et al.
Published: (2025)
VideoDirector: Precise Video Editing via Text-to-Video Models
by: Wang, Yukun, et al.
Published: (2024)
by: Wang, Yukun, et al.
Published: (2024)
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
by: Yang, Jiahui, et al.
Published: (2024)
by: Yang, Jiahui, et al.
Published: (2024)
TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction
by: Huang, Yiyao, et al.
Published: (2025)
by: Huang, Yiyao, et al.
Published: (2025)
Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
Text Prompting for Multi-Concept Video Customization by Autoregressive Generation
by: Kothandaraman, Divya, et al.
Published: (2024)
by: Kothandaraman, Divya, et al.
Published: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Motion is the Choreographer: Learning Latent Pose Dynamics for Seamless Sign Language Generation
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories
by: Geng, Daniel, et al.
Published: (2024)
by: Geng, Daniel, et al.
Published: (2024)
V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
by: Luo, Yang, et al.
Published: (2025)
by: Luo, Yang, et al.
Published: (2025)
Pretrain-then-Adapt: Uncertainty-Aware Test-Time Adaptation for Text-based Person Search
by: Zhang, Jiahao, et al.
Published: (2026)
by: Zhang, Jiahao, et al.
Published: (2026)
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
by: Wang, Zhao, et al.
Published: (2024)
by: Wang, Zhao, et al.
Published: (2024)
VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
by: Sun, Shangkun, et al.
Published: (2024)
by: Sun, Shangkun, et al.
Published: (2024)
Similar Items
-
Towards Geometric-Photometric Joint Alignment for Facial Mesh Registration
by: Wang, Xizhi, et al.
Published: (2024) -
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025) -
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval
by: Jiang, Xintong, et al.
Published: (2024) -
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025) -
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)