Rethink Sparse Signals for Pose-guided Text-to-image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Xuan, Wenjie, Zhang, Jing, Liu, Juhua, Du, Bo, Tao, Dacheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
por: Zhong, Qihuang, et al.
Publicado: (2026)
por: Zhong, Qihuang, et al.
Publicado: (2026)
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
por: He, Haibin, et al.
Publicado: (2025)
por: He, Haibin, et al.
Publicado: (2025)
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
por: He, Haibin, et al.
Publicado: (2024)
por: He, Haibin, et al.
Publicado: (2024)
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
por: Ye, Maoyuan, et al.
Publicado: (2023)
por: Ye, Maoyuan, et al.
Publicado: (2023)
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
por: Xuan, Wenjie, et al.
Publicado: (2024)
por: Xuan, Wenjie, et al.
Publicado: (2024)
Hi-SAM: Marrying Segment Anything Model for Hierarchical Text Segmentation
por: Ye, Maoyuan, et al.
Publicado: (2024)
por: Ye, Maoyuan, et al.
Publicado: (2024)
RFL-CDNet: Towards Accurate Change Detection via Richer Feature Learning
por: Gan, Yuhang, et al.
Publicado: (2024)
por: Gan, Yuhang, et al.
Publicado: (2024)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
por: He, Haibin, et al.
Publicado: (2025)
por: He, Haibin, et al.
Publicado: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
por: He, Haibin, et al.
Publicado: (2026)
por: He, Haibin, et al.
Publicado: (2026)
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
por: Zhang, Xike, et al.
Publicado: (2026)
por: Zhang, Xike, et al.
Publicado: (2026)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
por: He, Haibin, et al.
Publicado: (2025)
por: He, Haibin, et al.
Publicado: (2025)
Detect Changes like Humans: Incorporating Semantic Priors for Improved Change Detection
por: Gan, Yuhang, et al.
Publicado: (2024)
por: Gan, Yuhang, et al.
Publicado: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
por: Ye, Maoyuan, et al.
Publicado: (2025)
por: Ye, Maoyuan, et al.
Publicado: (2025)
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions
por: Ma, Sihan, et al.
Publicado: (2024)
por: Ma, Sihan, et al.
Publicado: (2024)
TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
por: Wang, Yanwen, et al.
Publicado: (2025)
por: Wang, Yanwen, et al.
Publicado: (2025)
LAB-Det: Language as a Domain-Invariant Bridge for Training-Free One-Shot Domain Generalization in Object Detection
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
PolarMAE: Efficient Fetal Ultrasound Pre-training via Semantic Screening and Polar-Guided Masking
por: Lv, Meng, et al.
Publicado: (2026)
por: Lv, Meng, et al.
Publicado: (2026)
Free-Form Motion Control: Controlling the 6D Poses of Camera and Objects in Video Generation
por: Shuai, Xincheng, et al.
Publicado: (2025)
por: Shuai, Xincheng, et al.
Publicado: (2025)
On Robust Cross-View Consistency in Self-Supervised Monocular Depth Estimation
por: Zhao, Haimei, et al.
Publicado: (2022)
por: Zhao, Haimei, et al.
Publicado: (2022)
CogMorph: Cognitive Morphing Attacks for Text-to-Image Models
por: Jing, Zonglei, et al.
Publicado: (2025)
por: Jing, Zonglei, et al.
Publicado: (2025)
Rethinking Model Efficiency: Multi-Agent Inference with Large Models
por: Dong, Sixun, et al.
Publicado: (2026)
por: Dong, Sixun, et al.
Publicado: (2026)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
por: Lu, Feichi, et al.
Publicado: (2024)
por: Lu, Feichi, et al.
Publicado: (2024)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
por: Lu, Wenquan, et al.
Publicado: (2023)
por: Lu, Wenquan, et al.
Publicado: (2023)
PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence
por: Wang, Ruiyan, et al.
Publicado: (2025)
por: Wang, Ruiyan, et al.
Publicado: (2025)
Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
por: Chen, Hang, et al.
Publicado: (2025)
por: Chen, Hang, et al.
Publicado: (2025)
Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
Rethinking Structure Preservation in Text-Guided Image Editing with Visual Autoregressive Models
por: Xia, Tao, et al.
Publicado: (2026)
por: Xia, Tao, et al.
Publicado: (2026)
Reverse Prompt: Cracking the Recipe Inside Text-to-Image Generation
por: Ren, Zhiyao, et al.
Publicado: (2025)
por: Ren, Zhiyao, et al.
Publicado: (2025)
Contact-aware Human Motion Generation from Textual Descriptions
por: Ma, Sihan, et al.
Publicado: (2024)
por: Ma, Sihan, et al.
Publicado: (2024)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
por: Wang, Jiajun, et al.
Publicado: (2024)
por: Wang, Jiajun, et al.
Publicado: (2024)
Object Gaussian for Monocular 6D Pose Estimation from Sparse Views
por: Luo, Luqing, et al.
Publicado: (2024)
por: Luo, Luqing, et al.
Publicado: (2024)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
por: Ma, Yue, et al.
Publicado: (2023)
por: Ma, Yue, et al.
Publicado: (2023)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
por: Qian, Qi, et al.
Publicado: (2024)
por: Qian, Qi, et al.
Publicado: (2024)
VSDiffusion: Taming Ill-Posed Shadow Generation via Visibility-Constrained Diffusion
por: Li, Jing, et al.
Publicado: (2026)
por: Li, Jing, et al.
Publicado: (2026)
Scriboora: Rethinking Human Pose Forecasting
por: Bermuth, Daniel, et al.
Publicado: (2025)
por: Bermuth, Daniel, et al.
Publicado: (2025)
FashionPose: Text to Pose to Relight Image Generation for Personalized Fashion Visualization
por: Shi, Chuancheng, et al.
Publicado: (2025)
por: Shi, Chuancheng, et al.
Publicado: (2025)
Text-guided Zero-Shot Object Localization
por: Wang, Jingjing, et al.
Publicado: (2024)
por: Wang, Jingjing, et al.
Publicado: (2024)
ER-Pose: Rethinking Keypoint-Driven Representation Learning for Real-Time Human Pose Estimation
por: Li, Nanjun, et al.
Publicado: (2026)
por: Li, Nanjun, et al.
Publicado: (2026)
LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis
por: Zhao, Shitian, et al.
Publicado: (2025)
por: Zhao, Shitian, et al.
Publicado: (2025)
GlyphPrinter: Region-Grouped Direct Preference Optimization for Glyph-Accurate Visual Text Rendering
por: Shuai, Xincheng, et al.
Publicado: (2026)
por: Shuai, Xincheng, et al.
Publicado: (2026)
Ejemplares similares
-
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
por: Zhong, Qihuang, et al.
Publicado: (2026) -
GoMatching++: Parameter- and Data-Efficient Arbitrary-Shaped Video Text Spotting and Benchmarking
por: He, Haibin, et al.
Publicado: (2025) -
GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching
por: He, Haibin, et al.
Publicado: (2024) -
DeepSolo++: Let Transformer Decoder with Explicit Points Solo for Multilingual Text Spotting
por: Ye, Maoyuan, et al.
Publicado: (2023) -
When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
por: Xuan, Wenjie, et al.
Publicado: (2024)