DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Siniukov, Maksim, Chang, Di, Tran, Minh, Gong, Hongkun, Chaubey, Ashutosh, Soleymani, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
von: Jin, Zhangyu, et al.
Veröffentlicht: (2026)
von: Jin, Zhangyu, et al.
Veröffentlicht: (2026)
DreamMapping: High-Fidelity Text-to-3D Generation via Variational Distribution Mapping
von: Cai, Zeyu, et al.
Veröffentlicht: (2024)
von: Cai, Zeyu, et al.
Veröffentlicht: (2024)
Data Augmentation in Earth Observation: A Diffusion Model Approach
von: Sousa, Tiago, et al.
Veröffentlicht: (2024)
von: Sousa, Tiago, et al.
Veröffentlicht: (2024)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
Deepfake Detection Generalization with Diffusion Noise
von: Qi, Hongyuan, et al.
Veröffentlicht: (2026)
von: Qi, Hongyuan, et al.
Veröffentlicht: (2026)
CerberusDet: Unified Multi-Dataset Object Detection
von: Tolstykh, Irina, et al.
Veröffentlicht: (2024)
von: Tolstykh, Irina, et al.
Veröffentlicht: (2024)
Orientation-conditioned Facial Texture Mapping for Video-based Facial Remote Photoplethysmography Estimation
von: Cantrill, Sam, et al.
Veröffentlicht: (2024)
von: Cantrill, Sam, et al.
Veröffentlicht: (2024)
Dyadic Interaction Modeling for Social Behavior Generation
von: Tran, Minh, et al.
Veröffentlicht: (2024)
von: Tran, Minh, et al.
Veröffentlicht: (2024)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
NumeriKontrol: Adding Numeric Control to Diffusion Transformers for Instruction-based Image Editing
von: Xu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Xu, Zhenyu, et al.
Veröffentlicht: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
Chat-Driven Text Generation and Interaction for Person Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2025)
von: Xie, Zequn, et al.
Veröffentlicht: (2025)
Novel Concept-Oriented Synthetic Data approach for Training Generative AI-Driven Crystal Grain Analysis Using Diffusion Model
von: Saleh, Ahmed Sobhi, et al.
Veröffentlicht: (2025)
von: Saleh, Ahmed Sobhi, et al.
Veröffentlicht: (2025)
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
von: Liu, Shanyuan, et al.
Veröffentlicht: (2023)
von: Liu, Shanyuan, et al.
Veröffentlicht: (2023)
Exploring Diffusion with Test-Time Training on Efficient Image Restoration
von: Lu, Rongchang, et al.
Veröffentlicht: (2025)
von: Lu, Rongchang, et al.
Veröffentlicht: (2025)
Video-Based Human Pose Regression via Decoupled Space-Time Aggregation
von: He, Jijie, et al.
Veröffentlicht: (2024)
von: He, Jijie, et al.
Veröffentlicht: (2024)
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Ma, Zhiyuan, et al.
Veröffentlicht: (2024)
Can LLM Agents Respond to Disasters? Benchmarking Heterogeneous Geospatial Reasoning in Emergency Operations
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
WPDA: Frequency-based Backdoor Attack with Wavelet Packet Decomposition
von: Song, Zhengyao, et al.
Veröffentlicht: (2024)
von: Song, Zhengyao, et al.
Veröffentlicht: (2024)
EEG Signal Denoising Using pix2pix GAN: Enhancing Neurological Data Analysis
von: Wang, Haoyi, et al.
Veröffentlicht: (2024)
von: Wang, Haoyi, et al.
Veröffentlicht: (2024)
AVadCLIP: Audio-Visual Collaboration for Robust Video Anomaly Detection
von: Wu, Peng, et al.
Veröffentlicht: (2025)
von: Wu, Peng, et al.
Veröffentlicht: (2025)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
von: Zheng, Hantao, et al.
Veröffentlicht: (2026)
von: Zheng, Hantao, et al.
Veröffentlicht: (2026)
DIsoN: Decentralized Isolation Networks for Out-of-Distribution Detection in Medical Imaging
von: Wagner, Felix, et al.
Veröffentlicht: (2025)
von: Wagner, Felix, et al.
Veröffentlicht: (2025)
DreamPBR: Text-driven Generation of High-resolution SVBRDF with Multi-modal Guidance
von: Xin, Linxuan, et al.
Veröffentlicht: (2024)
von: Xin, Linxuan, et al.
Veröffentlicht: (2024)
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
von: Kuprashevich, Maksim, et al.
Veröffentlicht: (2024)
von: Kuprashevich, Maksim, et al.
Veröffentlicht: (2024)
Facial Spatiotemporal Graphs: Leveraging the 3D Facial Surface for Remote Physiological Measurement
von: Cantrill, Sam, et al.
Veröffentlicht: (2026)
von: Cantrill, Sam, et al.
Veröffentlicht: (2026)
Human-Centric Anomaly Detection in Surveillance Videos Using YOLO-World and Spatio-Temporal Deep Learning
von: Naeen, Mohammad Ali Etemadi, et al.
Veröffentlicht: (2025)
von: Naeen, Mohammad Ali Etemadi, et al.
Veröffentlicht: (2025)
Diffusion-Based Ukrainian Handwritten Text Generation with Cross-Domain Style Transfer
von: Ahitoliev, Andrii, et al.
Veröffentlicht: (2026)
von: Ahitoliev, Andrii, et al.
Veröffentlicht: (2026)
VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2024)
Cost Savings from Automatic Quality Assessment of Generated Images
von: Giro-i-Nieto, Xavier, et al.
Veröffentlicht: (2025)
von: Giro-i-Nieto, Xavier, et al.
Veröffentlicht: (2025)
AVControl: Efficient Framework for Training Audio-Visual Controls
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
Advanced Long-term Earth System Forecasting
von: Wu, Hao, et al.
Veröffentlicht: (2025)
von: Wu, Hao, et al.
Veröffentlicht: (2025)
IAMAP: Unlocking Deep Learning in QGIS for non-coders and limited computing resources
von: Tresson, Paul, et al.
Veröffentlicht: (2025)
von: Tresson, Paul, et al.
Veröffentlicht: (2025)
Optimizing Multi-Scale Representations to Detect Effect Heterogeneity Using Earth Observation and Computer Vision: Applications to Two Anti-Poverty RCTs
von: Zhu, Fucheng Warren, et al.
Veröffentlicht: (2024)
von: Zhu, Fucheng Warren, et al.
Veröffentlicht: (2024)
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
von: Pang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Pang, Jiacheng, et al.
Veröffentlicht: (2026)
Fingerprint Membership and Identity Inference Against Generative Adversarial Networks
von: Cavasin, Saverio, et al.
Veröffentlicht: (2024)
von: Cavasin, Saverio, et al.
Veröffentlicht: (2024)
SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation
von: Cai, Changpeng, et al.
Veröffentlicht: (2024)
von: Cai, Changpeng, et al.
Veröffentlicht: (2024)
Egocentric Video: A New Tool for Capturing Hand Use of Individuals with Spinal Cord Injury at Home
von: Likitlersuang, Jirapat, et al.
Veröffentlicht: (2018)
von: Likitlersuang, Jirapat, et al.
Veröffentlicht: (2018)
Ähnliche Einträge
-
GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
von: Jin, Zhangyu, et al.
Veröffentlicht: (2026) -
DreamMapping: High-Fidelity Text-to-3D Generation via Variational Distribution Mapping
von: Cai, Zeyu, et al.
Veröffentlicht: (2024) -
Data Augmentation in Earth Observation: A Diffusion Model Approach
von: Sousa, Tiago, et al.
Veröffentlicht: (2024) -
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024) -
Deepfake Detection Generalization with Diffusion Noise
von: Qi, Hongyuan, et al.
Veröffentlicht: (2026)