VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mazumdar, Amrita, Park, Seonwook, Roy, Rajarshi, Srihari, Nikhil, Wang, Shengze, Zhou, Yuhao, Wang, Julia, Nagano, Koki, De Mello, Shalini |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Instant Expressive Gaussian Head Avatar via 3D-Aware Expression Distillation
von: Jiang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Jiang, Kaiwen, et al.
Veröffentlicht: (2025)
PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026)
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
von: Pan, Shuchang, et al.
Veröffentlicht: (2025)
von: Pan, Shuchang, et al.
Veröffentlicht: (2025)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
von: Lu, Xiangyu, et al.
Veröffentlicht: (2025)
von: Lu, Xiangyu, et al.
Veröffentlicht: (2025)
Coherent 3D Portrait Video Reconstruction via Triplane Fusion
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
Coherent3D: Coherent 3D Portrait Video Reconstruction via Triplane Fusion
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
Avatar Fingerprinting for Authorized Use of Synthetic Talking-Head Videos
von: Prashnani, Ekta, et al.
Veröffentlicht: (2023)
von: Prashnani, Ekta, et al.
Veröffentlicht: (2023)
BLADE: Single-view Body Mesh Learning through Accurate Depth Estimation
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
von: Wang, Shengze, et al.
Veröffentlicht: (2024)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
von: Corvi, Riccardo, et al.
Veröffentlicht: (2025)
von: Corvi, Riccardo, et al.
Veröffentlicht: (2025)
QUEEN: QUantized Efficient ENcoding of Dynamic Gaussians for Streaming Free-viewpoint Videos
von: Girish, Sharath, et al.
Veröffentlicht: (2024)
von: Girish, Sharath, et al.
Veröffentlicht: (2024)
FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
von: Liao, Borui, et al.
Veröffentlicht: (2025)
von: Liao, Borui, et al.
Veröffentlicht: (2025)
COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing
von: Barthel, Florian, et al.
Veröffentlicht: (2026)
von: Barthel, Florian, et al.
Veröffentlicht: (2026)
F-Actor: Controllable Conversational Behaviour in Full-Duplex Models
von: Züfle, Maike, et al.
Veröffentlicht: (2026)
von: Züfle, Maike, et al.
Veröffentlicht: (2026)
MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models
von: Zhang, He, et al.
Veröffentlicht: (2025)
von: Zhang, He, et al.
Veröffentlicht: (2025)
Full-Duplex Strategy for Video Object Segmentation
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
A Unified Approach for Text- and Image-guided 4D Scene Generation
von: Zheng, Yufeng, et al.
Veröffentlicht: (2023)
von: Zheng, Yufeng, et al.
Veröffentlicht: (2023)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
von: Riera, Pablo, et al.
Veröffentlicht: (2026)
von: Riera, Pablo, et al.
Veröffentlicht: (2026)
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
Steering Conversational Large Language Models for Long Emotional Support Conversations
von: Madani, Navid, et al.
Veröffentlicht: (2024)
von: Madani, Navid, et al.
Veröffentlicht: (2024)
GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning
von: Yuan, Ye, et al.
Veröffentlicht: (2023)
von: Yuan, Ye, et al.
Veröffentlicht: (2023)
ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
von: Madani, Navid, et al.
Veröffentlicht: (2025)
von: Madani, Navid, et al.
Veröffentlicht: (2025)
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
Consolidating Strategies for Countering Hate Speech Using Persuasive Dialogues
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
von: Saha, Sougata, et al.
Veröffentlicht: (2024)
Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems
von: Yu, Haoyuan, et al.
Veröffentlicht: (2026)
von: Yu, Haoyuan, et al.
Veröffentlicht: (2026)
Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations
von: Singh, Bhaskar, et al.
Veröffentlicht: (2026)
von: Singh, Bhaskar, et al.
Veröffentlicht: (2026)
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
von: Peng, Yizhou, et al.
Veröffentlicht: (2025)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs
von: Trevithick, Alex, et al.
Veröffentlicht: (2024)
von: Trevithick, Alex, et al.
Veröffentlicht: (2024)
Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
von: Xie, Zhifei, et al.
Veröffentlicht: (2024)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2026)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2026)
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations
von: Pal, Sayantan, et al.
Veröffentlicht: (2024)
von: Pal, Sayantan, et al.
Veröffentlicht: (2024)
VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
von: Bandraupalli, Srihari, et al.
Veröffentlicht: (2025)
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
von: Wu, Donghang, et al.
Veröffentlicht: (2025)
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
von: Cui, Junbo, et al.
Veröffentlicht: (2026)
von: Cui, Junbo, et al.
Veröffentlicht: (2026)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Learning Branching-Time Properties in CTL and ATL via Constraint Solving
von: Bordais, Benjamin, et al.
Veröffentlicht: (2024)
von: Bordais, Benjamin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Instant Expressive Gaussian Head Avatar via 3D-Aware Expression Distillation
von: Jiang, Kaiwen, et al.
Veröffentlicht: (2025) -
PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
von: Roy, Rajarshi, et al.
Veröffentlicht: (2026) -
Enabling Conversational Behavior Reasoning Capabilities in Full-Duplex Speech
von: Pan, Shuchang, et al.
Veröffentlicht: (2025) -
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
von: Lu, Xiangyu, et al.
Veröffentlicht: (2025) -
Coherent 3D Portrait Video Reconstruction via Triplane Fusion
von: Wang, Shengze, et al.
Veröffentlicht: (2024)