GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Zhangyu, Siniukov, Maksim, Kwon, Deuksin, Chaubey, Ashutosh, Soleymani, Mohammad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
by: Siniukov, Maksim, et al.
Published: (2025)
by: Siniukov, Maksim, et al.
Published: (2025)
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026)
by: Chaubey, Ashutosh, et al.
Published: (2026)
Discrete Facial Encoding: : A Framework for Data-driven Facial Display Discovery
by: Tran, Minh, et al.
Published: (2025)
by: Tran, Minh, et al.
Published: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Dyadic Interaction Modeling for Social Behavior Generation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Do Audio LLMs Listen or Read? Analyzing and Mitigating Paralinguistic Failures with VoxParadox
by: Pang, Jiacheng, et al.
Published: (2026)
by: Pang, Jiacheng, et al.
Published: (2026)
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026)
by: Chaubey, Ashutosh, et al.
Published: (2026)
Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
by: Chaubey, Ashutosh, et al.
Published: (2025)
by: Chaubey, Ashutosh, et al.
Published: (2025)
Towards a Generalizable Speech Marker for Parkinson's Disease Diagnosis
by: Siniukov, Maksim, et al.
Published: (2025)
by: Siniukov, Maksim, et al.
Published: (2025)
GDPO-SR: Group Direct Preference Optimization for One-Step Generative Image Super-Resolution
by: Yi, Qiaosi, et al.
Published: (2026)
by: Yi, Qiaosi, et al.
Published: (2026)
MetaGDPO: Alleviating Catastrophic Forgetting with Metacognitive Knowledge through Group Direct Preference Optimization
by: Zhang, Lanxue, et al.
Published: (2025)
by: Zhang, Lanxue, et al.
Published: (2025)
Personality Expression Across Contexts: Linguistic and Behavioral Variation in LLM Agents
by: Han, Bin, et al.
Published: (2026)
by: Han, Bin, et al.
Published: (2026)
Active Listener: Continuous Generation of Listener's Head Motion Response in Dyadic Interactions
by: Ghosh, Bishal, et al.
Published: (2024)
by: Ghosh, Bishal, et al.
Published: (2024)
GDPO: Learning to Directly Align Language Models with Diversity Using GFlowNets
by: Kwon, Oh Joon, et al.
Published: (2024)
by: Kwon, Oh Joon, et al.
Published: (2024)
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
by: Liu, Xi, et al.
Published: (2024)
by: Liu, Xi, et al.
Published: (2024)
VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
by: Li, Shiying, et al.
Published: (2025)
by: Li, Shiying, et al.
Published: (2025)
Can LLMs Generate Behaviors for Embodied Virtual Agents Based on Personality Traits?
by: Han, Bin, et al.
Published: (2025)
by: Han, Bin, et al.
Published: (2025)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
Virtual Reality in Social Media: A New Era of Immersive Social Interactions
by: Chaubey, Priyanshu
Published: (2025)
by: Chaubey, Priyanshu
Published: (2025)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
by: Tal, Or, et al.
Published: (2025)
by: Tal, Or, et al.
Published: (2025)
RateCount: Learning-Free Device Counting by Wi-Fi Probe Listening
by: He, Tianlang, et al.
Published: (2025)
by: He, Tianlang, et al.
Published: (2025)
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
by: Wang, Haotian, et al.
Published: (2024)
by: Wang, Haotian, et al.
Published: (2024)
Discrete Flow Matching Policy Optimization
by: Su, Maojiang, et al.
Published: (2026)
by: Su, Maojiang, et al.
Published: (2026)
SAT-SKYLINES: 3D Building Generation from Satellite Imagery and Coarse Geometric Priors
by: Jin, Zhangyu, et al.
Published: (2025)
by: Jin, Zhangyu, et al.
Published: (2025)
Quasiparticle solutions to the 1D nonlocal Fisher--KPP equation with a fractal time derivative in the weak diffusion approximation
by: Shapovalov, A. V., et al.
Published: (2025)
by: Shapovalov, A. V., et al.
Published: (2025)
ASTRA: A Negotiation Agent with Adaptive and Strategic Reasoning via Tool-integrated Action for Dynamic Offer Optimization
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Evaluation of Hemoglobin by Using Sahli's Method and Automated Analyser, A Comparative Study in Paediatric Age Group in A Tertiary Care Hospital of Assam
by: Dr. Jyoti Chaubey
Published: (2025)
by: Dr. Jyoti Chaubey
Published: (2025)
Transformers are Expressive, But Are They Expressive Enough for Regression?
by: Nath, Swaroop, et al.
Published: (2024)
by: Nath, Swaroop, et al.
Published: (2024)
Design and Analysis of Binaural Signal Matching with Arbitrary Microphone Arrays and Listener Head Rotations
by: Madmoni, Lior, et al.
Published: (2024)
by: Madmoni, Lior, et al.
Published: (2024)
PARCO: Parallel AutoRegressive Models for Multi-Agent Combinatorial Optimization
by: Berto, Federico, et al.
Published: (2024)
by: Berto, Federico, et al.
Published: (2024)
Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans
by: Kwon, Deuksin, et al.
Published: (2025)
by: Kwon, Deuksin, et al.
Published: (2025)
Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
by: Geng, Zichen, et al.
Published: (2025)
by: Geng, Zichen, et al.
Published: (2025)
Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
by: Doo, JaeHyeok, et al.
Published: (2026)
by: Doo, JaeHyeok, et al.
Published: (2026)
Soft Sequence Policy Optimization
by: Glazyrina, Svetlana, et al.
Published: (2026)
by: Glazyrina, Svetlana, et al.
Published: (2026)
F5R-TTS: Improving Flow-Matching based Text-to-Speech with Group Relative Policy Optimization
by: Sun, Xiaohui, et al.
Published: (2025)
by: Sun, Xiaohui, et al.
Published: (2025)
FARMER: Flow AutoRegressive Transformer over Pixels
by: Zheng, Guangting, et al.
Published: (2025)
by: Zheng, Guangting, et al.
Published: (2025)
Ansätz Expressivity and Optimization in Variational Quantum Simulations of Transverse-field Ising Model Across System Sizes
by: Tripathi, Ashutosh P., et al.
Published: (2026)
by: Tripathi, Ashutosh P., et al.
Published: (2026)
PromptGAR: Flexible Promptive Group Activity Recognition
by: Jin, Zhangyu, et al.
Published: (2025)
by: Jin, Zhangyu, et al.
Published: (2025)
On the Generalization in Topology Optimization via Sensitivity-Conditioned Bernoulli Flow Matching
by: Rashed, Mohammad, et al.
Published: (2026)
by: Rashed, Mohammad, et al.
Published: (2026)
Flow Matching Policy Optimization with Mirror Descent and Entropy Constraints
by: Gao, Ting, et al.
Published: (2026)
by: Gao, Ting, et al.
Published: (2026)
Similar Items
-
DiTaiListener: Controllable High Fidelity Listener Video Generation with Diffusion
by: Siniukov, Maksim, et al.
Published: (2025) -
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026) -
Discrete Facial Encoding: : A Framework for Data-driven Facial Display Discovery
by: Tran, Minh, et al.
Published: (2025) -
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026) -
Dyadic Interaction Modeling for Social Behavior Generation
by: Tran, Minh, et al.
Published: (2024)