VerbDiff: Text-Only Diffusion Models with Enhanced Interaction Awareness
Fuente:
arXiv
Saved in:
| Main Authors: | Cha, SeungJu, Lee, Kwanyoung, Kim, Ye-Chan, Oh, Hyunwoo, Kim, Dong-Jin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
by: Oh, Hyunwoo, et al.
Published: (2025)
by: Oh, Hyunwoo, et al.
Published: (2025)
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
by: Hoe, Jiun Tian, et al.
Published: (2023)
by: Hoe, Jiun Tian, et al.
Published: (2023)
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
by: Kim, Ye-Chan, et al.
Published: (2025)
by: Kim, Ye-Chan, et al.
Published: (2025)
ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
by: Koh, Sungho, et al.
Published: (2025)
by: Koh, Sungho, et al.
Published: (2025)
ADAPT: Attention Driven Adaptive Prompt Scheduling and InTerpolating Orthogonal Complements for Rare Concepts Generation
by: Lee, Kwanyoung, et al.
Published: (2026)
by: Lee, Kwanyoung, et al.
Published: (2026)
Adaptive Auxiliary Prompt Blending for Target-Faithful Diffusion Generation
by: Lee, Kwanyoung, et al.
Published: (2026)
by: Lee, Kwanyoung, et al.
Published: (2026)
InteractEdit: Zero-Shot Editing of Human-Object Interactions in Images
by: Hoe, Jiun Tian, et al.
Published: (2025)
by: Hoe, Jiun Tian, et al.
Published: (2025)
KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial Animation
by: Lyu, Tianle, et al.
Published: (2025)
by: Lyu, Tianle, et al.
Published: (2025)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
by: Wu, Yuheng, et al.
Published: (2026)
by: Wu, Yuheng, et al.
Published: (2026)
On Copyright Risks of Text-to-Image Diffusion Models
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Exploring Palette based Color Guidance in Diffusion Models
by: Qiu, Qianru, et al.
Published: (2025)
by: Qiu, Qianru, et al.
Published: (2025)
STAR: Skeleton-aware Text-based 4D Avatar Generation with In-Network Motion Retargeting
by: Chai, Zenghao, et al.
Published: (2024)
by: Chai, Zenghao, et al.
Published: (2024)
Real-Time Position-Aware View Synthesis from Single-View Input
by: Gond, Manu, et al.
Published: (2024)
by: Gond, Manu, et al.
Published: (2024)
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
by: Guan, Jiazhi, et al.
Published: (2025)
by: Guan, Jiazhi, et al.
Published: (2025)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
by: Zhao, Junchuan, et al.
Published: (2026)
by: Zhao, Junchuan, et al.
Published: (2026)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
by: Lin, Jiantao, et al.
Published: (2025)
by: Lin, Jiantao, et al.
Published: (2025)
SVGS: Enhancing Gaussian Splatting Using Primitives with Spatially Varying Colors
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
Generating Digital Models Using Text-to-3D and Image-to-3D Prompts: Critical Case Study
by: Ziatdinov, Rushan, et al.
Published: (2025)
by: Ziatdinov, Rushan, et al.
Published: (2025)
DASC: Depth-of-Field Aware Scene Complexity Metric for 3D Visualization on Light Field Display
by: Akbar, Kamran, et al.
Published: (2025)
by: Akbar, Kamran, et al.
Published: (2025)
TreeMeshGPT: Artistic Mesh Generation with Autoregressive Tree Sequencing
by: Lionar, Stefan, et al.
Published: (2025)
by: Lionar, Stefan, et al.
Published: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
by: Razlighi, AmirHossein Naghi, et al.
Published: (2026)
ComVi: Context-Aware Optimized Comment Display in Video Playback
by: Kim, Minsun, et al.
Published: (2026)
by: Kim, Minsun, et al.
Published: (2026)
Closing the Loop: Unified 3D Scene Generation and Immersive Interaction via LLM-RL Coupling
by: Vo, Anh H., et al.
Published: (2026)
by: Vo, Anh H., et al.
Published: (2026)
GTLR-GS: Geometry-Texture Aware LiDAR-Regularized 3D Gaussian Splatting for Realistic Scene Reconstruction
by: Fang, Yan, et al.
Published: (2026)
by: Fang, Yan, et al.
Published: (2026)
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
by: Gupta, Prerit, et al.
Published: (2025)
by: Gupta, Prerit, et al.
Published: (2025)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
by: Dong, Wenqi, et al.
Published: (2025)
by: Dong, Wenqi, et al.
Published: (2025)
Instruction-Driven 3D Facial Expression Generation and Transition
by: Vo, Anh H., et al.
Published: (2026)
by: Vo, Anh H., et al.
Published: (2026)
Real-Time Interactive Hybrid Ocean: Spectrum-Consistent Wave Particle-FFT Coupling
by: Xue, Shengze, et al.
Published: (2025)
by: Xue, Shengze, et al.
Published: (2025)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
ToonAging: Face Re-Aging upon Artistic Portrait Style Transfer
by: Kim, Bumsoo, et al.
Published: (2024)
by: Kim, Bumsoo, et al.
Published: (2024)
FairyGen: Storied Cartoon Video from a Single Child-Drawn Character
by: Zheng, Jiayi, et al.
Published: (2025)
by: Zheng, Jiayi, et al.
Published: (2025)
Perceive-Sample-Compress: Towards Real-Time 3D Gaussian Splatting
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Splatography: Sparse multi-view dynamic Gaussian Splatting for filmmaking challenges
by: Azzarelli, Adrian, et al.
Published: (2025)
by: Azzarelli, Adrian, et al.
Published: (2025)
DreamCinema: Cinematic Transfer with Free Camera and 3D Character
by: Chen, Weiliang, et al.
Published: (2024)
by: Chen, Weiliang, et al.
Published: (2024)
MesonGS++: Post-training Compression of 3D Gaussian Splatting with Hyperparameter Searching
by: Xie, Shuzhao, et al.
Published: (2026)
by: Xie, Shuzhao, et al.
Published: (2026)
altiro3D: Scene representation from single image and novel view synthesis
by: Canessa, E., et al.
Published: (2023)
by: Canessa, E., et al.
Published: (2023)
Unveiling Deep Shadows: A Survey and Benchmark on Image and Video Shadow Detection, Removal, and Generation in the Deep Learning Era
by: Hu, Xiaowei, et al.
Published: (2024)
by: Hu, Xiaowei, et al.
Published: (2024)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
by: Xu, Zhen, et al.
Published: (2024)
by: Xu, Zhen, et al.
Published: (2024)
Neural Network-Based Tracking and 3D Reconstruction of Baseball Pitch Trajectories from Single-View 2D Video
by: Hsieh, Jhen
Published: (2024)
by: Hsieh, Jhen
Published: (2024)
Similar Items
-
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
by: Oh, Hyunwoo, et al.
Published: (2025) -
InteractDiffusion: Interaction Control in Text-to-Image Diffusion Models
by: Hoe, Jiun Tian, et al.
Published: (2023) -
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
by: Kim, Ye-Chan, et al.
Published: (2025) -
ScaleDiff: Higher-Resolution Image Synthesis via Efficient and Model-Agnostic Diffusion
by: Koh, Sungho, et al.
Published: (2025) -
ADAPT: Attention Driven Adaptive Prompt Scheduling and InTerpolating Orthogonal Complements for Rare Concepts Generation
by: Lee, Kwanyoung, et al.
Published: (2026)