OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luo, Cheng, Wang, Jianghui, Li, Bing, Song, Siyang, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
von: Luo, Cheng, et al.
Veröffentlicht: (2023)
von: Luo, Cheng, et al.
Veröffentlicht: (2023)
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
Towards a Multimodal Document-grounded Conversational AI System for Education
von: Taneja, Karan, et al.
Veröffentlicht: (2025)
von: Taneja, Karan, et al.
Veröffentlicht: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024)
Refusal as Silence: Gendered Disparities in Vision-Language Model Responses
von: Luo, Sha, et al.
Veröffentlicht: (2024)
von: Luo, Sha, et al.
Veröffentlicht: (2024)
Regressor-Guided Generative Image Editing Balances User Emotions to Reduce Time Spent Online
von: Gebhardt, Christoph, et al.
Veröffentlicht: (2025)
von: Gebhardt, Christoph, et al.
Veröffentlicht: (2025)
RITA: A Real-time Interactive Talking Avatars Framework
von: Cheng, Wuxinlin, et al.
Veröffentlicht: (2024)
von: Cheng, Wuxinlin, et al.
Veröffentlicht: (2024)
See-Control: A Multimodal Agent Framework for Smartphone Interaction with a Robotic Arm
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
Yume: An Interactive World Generation Model
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
von: Mao, Xiaofeng, et al.
Veröffentlicht: (2025)
Enhancing Online Learning by Integrating Biosensors and Multimodal Learning Analytics for Detecting and Predicting Student Behavior: A Review
von: Becerra, Alvaro, et al.
Veröffentlicht: (2025)
von: Becerra, Alvaro, et al.
Veröffentlicht: (2025)
E3RG: Building Explicit Emotion-driven Empathetic Response Generation System with Multimodal Large Language Model
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
von: Lin, Ronghao, et al.
Veröffentlicht: (2025)
ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
von: Yang, Boyin, et al.
Veröffentlicht: (2025)
von: Yang, Boyin, et al.
Veröffentlicht: (2025)
AI-based Multimodal Biometrics for Detecting Smartphone Distractions: Application to Online Learning
von: Becerra, Alvaro, et al.
Veröffentlicht: (2025)
von: Becerra, Alvaro, et al.
Veröffentlicht: (2025)
Scene-Aware Conversational ADAS with Generative AI for Real-Time Driver Assistance
von: Han, Kyungtae, et al.
Veröffentlicht: (2025)
von: Han, Kyungtae, et al.
Veröffentlicht: (2025)
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
von: Chen, Chieh-Yun, et al.
Veröffentlicht: (2025)
Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
von: Ki, Taekyung, et al.
Veröffentlicht: (2026)
Generative AI for Cel-Animation: A Survey
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
InterFeedback: Unveiling Interactive Intelligence of Large Multimodal Models via Human Feedback
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
von: Zhao, Henry Hengyuan, et al.
Veröffentlicht: (2025)
EEG-based Multimodal Representation Learning for Emotion Recognition
von: Yin, Kang, et al.
Veröffentlicht: (2024)
von: Yin, Kang, et al.
Veröffentlicht: (2024)
Explorer: Robust Collection of Interactable GUI Elements
von: Chaimalas, Iason, et al.
Veröffentlicht: (2025)
von: Chaimalas, Iason, et al.
Veröffentlicht: (2025)
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
von: Rekimoto, Jun
Veröffentlicht: (2025)
von: Rekimoto, Jun
Veröffentlicht: (2025)
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
von: Ouyang, Mingyu, et al.
Veröffentlicht: (2026)
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
von: Li, Po-han, et al.
Veröffentlicht: (2026)
von: Li, Po-han, et al.
Veröffentlicht: (2026)
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition
von: Wang, Zaitian, et al.
Veröffentlicht: (2025)
von: Wang, Zaitian, et al.
Veröffentlicht: (2025)
CG-MER: A Card Game-based Multimodal dataset for Emotion Recognition
von: Farhat, Nessrine, et al.
Veröffentlicht: (2025)
von: Farhat, Nessrine, et al.
Veröffentlicht: (2025)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
von: Cai, Weitong, et al.
Veröffentlicht: (2026)
Seeing Eye to AI: Human Alignment via Gaze-Based Response Rewards for Large Language Models
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
von: Lopez-Cardona, Angela, et al.
Veröffentlicht: (2024)
Simulating Clinical AI Assistance using Multimodal LLMs: A Case Study in Diabetic Retinopathy
von: Barakat, Nadim, et al.
Veröffentlicht: (2025)
von: Barakat, Nadim, et al.
Veröffentlicht: (2025)
StreamAvatar: Streaming Diffusion Models for Real-Time Interactive Human Avatars
von: Sun, Zhiyao, et al.
Veröffentlicht: (2025)
von: Sun, Zhiyao, et al.
Veröffentlicht: (2025)
SkinGEN: an Explainable Dermatology Diagnosis-to-Generation Framework with Interactive Vision-Language Models
von: Lin, Bo, et al.
Veröffentlicht: (2024)
von: Lin, Bo, et al.
Veröffentlicht: (2024)
Achieving Effective Virtual Reality Interactions via Acoustic Gesture Recognition based on Large Language Models
von: Zhang, Xijie, et al.
Veröffentlicht: (2025)
von: Zhang, Xijie, et al.
Veröffentlicht: (2025)
Real-Time Intuitive AI Drawing System for Collaboration: Enhancing Human Creativity through Formal and Contextual Intent Integration
von: Song, Jookyung, et al.
Veröffentlicht: (2025)
von: Song, Jookyung, et al.
Veröffentlicht: (2025)
ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
von: Cheng, Luo, et al.
Veröffentlicht: (2025)
von: Cheng, Luo, et al.
Veröffentlicht: (2025)
Using Text-to-Image Generation for Architectural Design Ideation
von: Paananen, Ville, et al.
Veröffentlicht: (2023)
von: Paananen, Ville, et al.
Veröffentlicht: (2023)
Characterizing Photorealism and Artifacts in Diffusion Model-Generated Images
von: Kamali, Negar, et al.
Veröffentlicht: (2025)
von: Kamali, Negar, et al.
Veröffentlicht: (2025)
Generative Augmented Reality: Paradigms, Technologies, and Future Applications
von: Liang, Chen, et al.
Veröffentlicht: (2025)
von: Liang, Chen, et al.
Veröffentlicht: (2025)
UI-UG: A Unified MLLM for UI Understanding and Generation
von: Yang, Hao, et al.
Veröffentlicht: (2025)
von: Yang, Hao, et al.
Veröffentlicht: (2025)
How to Distinguish AI-Generated Images from Authentic Photographs
von: Kamali, Negar, et al.
Veröffentlicht: (2024)
von: Kamali, Negar, et al.
Veröffentlicht: (2024)
It's a Feature, Not a Bug: Measuring Creative Fluidity in Image Generators
von: Ramaswamy, Aditi, et al.
Veröffentlicht: (2024)
von: Ramaswamy, Aditi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
von: Luo, Cheng, et al.
Veröffentlicht: (2023) -
ReactMotion: Generating Reactive Listener Motions from Speaker Utterance
von: Luo, Cheng, et al.
Veröffentlicht: (2026) -
Towards a Multimodal Document-grounded Conversational AI System for Education
von: Taneja, Karan, et al.
Veröffentlicht: (2025) -
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
von: Kapoor, Raghav, et al.
Veröffentlicht: (2024) -
Refusal as Silence: Gendered Disparities in Vision-Language Model Responses
von: Luo, Sha, et al.
Veröffentlicht: (2024)