Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Haozhuo, Zhu, Bin, Cao, Yu, Hao, Yanbin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffBrush:Just Painting the Art by Your Hands
di: Chu, Jiaming, et al.
Pubblicazione: (2025)
di: Chu, Jiaming, et al.
Pubblicazione: (2025)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
di: Wang, Zhihua, et al.
Pubblicazione: (2025)
di: Wang, Zhihua, et al.
Pubblicazione: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
di: Li, Chunyi, et al.
Pubblicazione: (2024)
di: Li, Chunyi, et al.
Pubblicazione: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
di: Zhang, Xinjie, et al.
Pubblicazione: (2024)
di: Zhang, Xinjie, et al.
Pubblicazione: (2024)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
di: Wang, Xidong, et al.
Pubblicazione: (2024)
di: Wang, Xidong, et al.
Pubblicazione: (2024)
Text-Only Data Synthesis for Vision Language Model Training
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
di: Yu, Xiaomin, et al.
Pubblicazione: (2025)
Accelerating Controllable Generation via Hybrid-grained Cache
di: Liu, Lin, et al.
Pubblicazione: (2025)
di: Liu, Lin, et al.
Pubblicazione: (2025)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
di: Duan, Yiqun, et al.
Pubblicazione: (2025)
di: Duan, Yiqun, et al.
Pubblicazione: (2025)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
di: Wu, Yi, et al.
Pubblicazione: (2025)
di: Wu, Yi, et al.
Pubblicazione: (2025)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
di: Ge, Shuheng, et al.
Pubblicazione: (2024)
di: Ge, Shuheng, et al.
Pubblicazione: (2024)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
di: Yang, Shuyu, et al.
Pubblicazione: (2024)
VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
di: Cha, SeungJu, et al.
Pubblicazione: (2025)
di: Cha, SeungJu, et al.
Pubblicazione: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
di: Chen, Junyu, et al.
Pubblicazione: (2025)
di: Chen, Junyu, et al.
Pubblicazione: (2025)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
di: Huang, Yuesheng, et al.
Pubblicazione: (2025)
di: Huang, Yuesheng, et al.
Pubblicazione: (2025)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
di: Gao, Jiayi, et al.
Pubblicazione: (2025)
di: Gao, Jiayi, et al.
Pubblicazione: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
di: Mao, Yuxin, et al.
Pubblicazione: (2024)
di: Mao, Yuxin, et al.
Pubblicazione: (2024)
Fine-grained Image Retrieval via Dual-Vision Adaptation
di: Jiang, Xin, et al.
Pubblicazione: (2025)
di: Jiang, Xin, et al.
Pubblicazione: (2025)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
di: Cai, Qi, et al.
Pubblicazione: (2026)
di: Cai, Qi, et al.
Pubblicazione: (2026)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
di: Qiao, Xiangshuo, et al.
Pubblicazione: (2024)
di: Qiao, Xiangshuo, et al.
Pubblicazione: (2024)
Seeing Text in the Dark: Algorithm and Benchmark
di: Xu, Chengpei, et al.
Pubblicazione: (2024)
di: Xu, Chengpei, et al.
Pubblicazione: (2024)
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
di: Gutflaish, Eyal, et al.
Pubblicazione: (2025)
di: Gutflaish, Eyal, et al.
Pubblicazione: (2025)
Deep Boosting Learning: A Brand-new Cooperative Approach for Image-Text Matching
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
di: Diao, Haiwen, et al.
Pubblicazione: (2024)
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
di: Ning, Hailong, et al.
Pubblicazione: (2025)
di: Ning, Hailong, et al.
Pubblicazione: (2025)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
di: Zhou, Pengyuan, et al.
Pubblicazione: (2024)
di: Zhou, Pengyuan, et al.
Pubblicazione: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
di: Chen, Nan, et al.
Pubblicazione: (2024)
di: Chen, Nan, et al.
Pubblicazione: (2024)
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
di: Fan, Chenyang, et al.
Pubblicazione: (2026)
di: Fan, Chenyang, et al.
Pubblicazione: (2026)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
di: Zhu, Lingsi, et al.
Pubblicazione: (2026)
di: Zhu, Lingsi, et al.
Pubblicazione: (2026)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
di: Qin, Yang, et al.
Pubblicazione: (2023)
di: Qin, Yang, et al.
Pubblicazione: (2023)
Generating Attribute-Aware Human Motions from Textual Prompt
di: Wang, Xinghan, et al.
Pubblicazione: (2025)
di: Wang, Xinghan, et al.
Pubblicazione: (2025)
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection
di: Nie, Fan, et al.
Pubblicazione: (2024)
di: Nie, Fan, et al.
Pubblicazione: (2024)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
di: Wang, Qing, et al.
Pubblicazione: (2025)
di: Wang, Qing, et al.
Pubblicazione: (2025)
PianoMotion10M: Dataset and Benchmark for Hand Motion Generation in Piano Performance
di: Gan, Qijun, et al.
Pubblicazione: (2024)
di: Gan, Qijun, et al.
Pubblicazione: (2024)
DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation
di: Yang, Haibo, et al.
Pubblicazione: (2024)
di: Yang, Haibo, et al.
Pubblicazione: (2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
di: Yuan, Shenghai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DiffBrush:Just Painting the Art by Your Hands
di: Chu, Jiaming, et al.
Pubblicazione: (2025) -
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
di: Wang, Zhihua, et al.
Pubblicazione: (2025) -
Selective Vision-Language Subspace Projection for Few-shot CLIP
di: Zhu, Xingyu, et al.
Pubblicazione: (2024) -
G-Refine: A General Quality Refiner for Text-to-Image Generation
di: Li, Chunyi, et al.
Pubblicazione: (2024) -
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
di: Zhang, Xinjie, et al.
Pubblicazione: (2024)