Learning Visual Generative Priors without Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Shuailei, Zheng, Kecheng, Wei, Ying, Wu, Wei, Lu, Fan, Zhang, Yifei, Xie, Chen-Wei, Gong, Biao, Zhu, Jiapeng, Shen, Yujun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DreamLIP: Language-Image Pre-training with Long Captions
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024)
von: Lu, Fan, et al.
Veröffentlicht: (2024)
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
von: Tan, Shuai, et al.
Veröffentlicht: (2025)
Aligned Better, Listen Better for Audio-Visual Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
von: Gong, Biao, et al.
Veröffentlicht: (2023)
von: Gong, Biao, et al.
Veröffentlicht: (2023)
Framer: Interactive Frame Interpolation
von: Wang, Wen, et al.
Veröffentlicht: (2024)
von: Wang, Wen, et al.
Veröffentlicht: (2024)
UCD: Unconditional Discriminator Promotes Nash Equilibrium in GANs
von: Xia, Mengfei, et al.
Veröffentlicht: (2025)
von: Xia, Mengfei, et al.
Veröffentlicht: (2025)
Ranni: Taming Text-to-Image Diffusion for Accurate Instruction Following
von: Feng, Yutong, et al.
Veröffentlicht: (2023)
von: Feng, Yutong, et al.
Veröffentlicht: (2023)
SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-world Object Detector
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
von: Ma, Shuailei, et al.
Veröffentlicht: (2023)
TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
von: Liu, Qinying, et al.
Veröffentlicht: (2023)
von: Liu, Qinying, et al.
Veröffentlicht: (2023)
PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
von: Liu, Shi, et al.
Veröffentlicht: (2024)
von: Liu, Shi, et al.
Veröffentlicht: (2024)
CoReS: Orchestrating the Dance of Reasoning and Segmentation
von: Bao, Xiaoyi, et al.
Veröffentlicht: (2024)
von: Bao, Xiaoyi, et al.
Veröffentlicht: (2024)
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
von: Zhu, Lei, et al.
Veröffentlicht: (2024)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
von: Yan, Junkai, et al.
Veröffentlicht: (2024)
von: Yan, Junkai, et al.
Veröffentlicht: (2024)
DreamDissector: Learning Disentangled Text-to-3D Generation from 2D Diffusion Priors
von: Yan, Zizheng, et al.
Veröffentlicht: (2024)
von: Yan, Zizheng, et al.
Veröffentlicht: (2024)
Diffusion Models Need Visual Priors for Image Generation
von: Yue, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yue, Xiaoyu, et al.
Veröffentlicht: (2024)
Supervised Learning without Backpropagation using Spike-Timing-Dependent Plasticity for Image Recognition
von: Xie, Wei
Veröffentlicht: (2024)
von: Xie, Wei
Veröffentlicht: (2024)
Harmonizing Visual Text Comprehension and Generation
von: Zhao, Zhen, et al.
Veröffentlicht: (2024)
von: Zhao, Zhen, et al.
Veröffentlicht: (2024)
Motion2VecSets: 4D Latent Vector Set Diffusion for Non-rigid Shape Reconstruction and Tracking
von: Cao, Wei, et al.
Veröffentlicht: (2024)
von: Cao, Wei, et al.
Veröffentlicht: (2024)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
Knowledge Visualization: A Benchmark and Method for Knowledge-Intensive Text-to-Image Generation
von: Zhao, Ran, et al.
Veröffentlicht: (2026)
von: Zhao, Ran, et al.
Veröffentlicht: (2026)
Learning Visual Affordance from Audio
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
von: Lu, Lidong, et al.
Veröffentlicht: (2025)
GenPC: Zero-shot Point Cloud Completion via 3D Generative Priors
von: Li, An, et al.
Veröffentlicht: (2025)
von: Li, An, et al.
Veröffentlicht: (2025)
Advancing Open-source World Models
von: Robbyant Team, et al.
Veröffentlicht: (2026)
von: Robbyant Team, et al.
Veröffentlicht: (2026)
Learning Naturally Aggregated Appearance for Efficient 3D Editing
von: Cheng, Ka Leong, et al.
Veröffentlicht: (2023)
von: Cheng, Ka Leong, et al.
Veröffentlicht: (2023)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
A Pragmatic VLA Foundation Model
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
MePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model
von: Wang, Xinyang, et al.
Veröffentlicht: (2024)
von: Wang, Xinyang, et al.
Veröffentlicht: (2024)
DP-TTA: Test-time Adaptation for Transient Electromagnetic Signal Denoising via Dictionary-driven Prior Regularization
von: Yang, Meng, et al.
Veröffentlicht: (2025)
von: Yang, Meng, et al.
Veröffentlicht: (2025)
TOOLCAD: Exploring Tool-Using Large Language Models in Text-to-CAD Generation with Reinforcement Learning
von: Gong, Yifei, et al.
Veröffentlicht: (2026)
von: Gong, Yifei, et al.
Veröffentlicht: (2026)
MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Adaptive Fused Prior Transfer for Controllable Generative Image Compression
von: Pei, Yifei, et al.
Veröffentlicht: (2026)
von: Pei, Yifei, et al.
Veröffentlicht: (2026)
The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
von: Wang, Hanlin, et al.
Veröffentlicht: (2025)
Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
von: Huang, Siteng, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DreamLIP: Language-Image Pre-training with Long Captions
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024) -
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
von: Wu, Wei, et al.
Veröffentlicht: (2024) -
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024) -
SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
von: Tan, Shuai, et al.
Veröffentlicht: (2025) -
Aligned Better, Listen Better for Audio-Visual Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)