EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Junhyeok, Kim, Min Soo, Chung, Jiwan, Cho, Jungbin, Kim, Jisoo, Kim, Sungwoong, Sim, Gyeongbo, Yu, Youngjae |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
di: Kim, Youngmin, et al.
Pubblicazione: (2025)
di: Kim, Youngmin, et al.
Pubblicazione: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
di: Kim, Junhyeok, et al.
Pubblicazione: (2025)
di: Kim, Junhyeok, et al.
Pubblicazione: (2025)
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
di: Kim, Jisoo, et al.
Pubblicazione: (2024)
di: Kim, Jisoo, et al.
Pubblicazione: (2024)
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
di: Cho, Jungbin, et al.
Pubblicazione: (2024)
di: Cho, Jungbin, et al.
Pubblicazione: (2024)
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
di: Cho, Jungbin, et al.
Pubblicazione: (2025)
di: Cho, Jungbin, et al.
Pubblicazione: (2025)
CANVAS: Commonsense-Aware Navigation System for Intuitive Human-Robot Interaction
di: Choi, Suhwan, et al.
Pubblicazione: (2024)
di: Choi, Suhwan, et al.
Pubblicazione: (2024)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
di: Kim, Junho, et al.
Pubblicazione: (2026)
di: Kim, Junho, et al.
Pubblicazione: (2026)
Teaching Metric Distance to Discrete Autoregressive Language Models
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing
di: Hwang, Inwoo, et al.
Pubblicazione: (2026)
di: Hwang, Inwoo, et al.
Pubblicazione: (2026)
EgoX: Egocentric Video Generation from a Single Exocentric Video
di: Kang, Taewoong, et al.
Pubblicazione: (2025)
di: Kang, Taewoong, et al.
Pubblicazione: (2025)
Global Geometry Is Not Enough for Vision Representations
di: Chung, Jiwan, et al.
Pubblicazione: (2026)
di: Chung, Jiwan, et al.
Pubblicazione: (2026)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
di: Kim, Kangsan, et al.
Pubblicazione: (2026)
di: Kim, Kangsan, et al.
Pubblicazione: (2026)
V.I.P. : Iterative Online Preference Distillation for Efficient Video Diffusion Models
di: Kim, Jisoo, et al.
Pubblicazione: (2025)
di: Kim, Jisoo, et al.
Pubblicazione: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
Towards Continuous Sign Language Conversation from Isolated Signs
di: Kim, Youngmin, et al.
Pubblicazione: (2026)
di: Kim, Youngmin, et al.
Pubblicazione: (2026)
OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
di: Oh, Yoonjin, et al.
Pubblicazione: (2025)
di: Oh, Yoonjin, et al.
Pubblicazione: (2025)
EgoCast: Forecasting Egocentric Human Pose in the Wild
di: Escobar, Maria, et al.
Pubblicazione: (2024)
di: Escobar, Maria, et al.
Pubblicazione: (2024)
TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
di: Yan, Jiaqi, et al.
Pubblicazione: (2025)
di: Yan, Jiaqi, et al.
Pubblicazione: (2025)
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
di: Zhang, Yu, et al.
Pubblicazione: (2026)
di: Zhang, Yu, et al.
Pubblicazione: (2026)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
di: Chung, Jiwan, et al.
Pubblicazione: (2024)
di: Chung, Jiwan, et al.
Pubblicazione: (2024)
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
di: Yoon, Taegyoon, et al.
Pubblicazione: (2026)
di: Yoon, Taegyoon, et al.
Pubblicazione: (2026)
Pri4R: Learning World Dynamics for Vision-Language-Action Models with Privileged 4D Representation
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
di: Kim, Jisoo, et al.
Pubblicazione: (2026)
EgoLM: Multi-Modal Language Model of Egocentric Motions
di: Hong, Fangzhou, et al.
Pubblicazione: (2024)
di: Hong, Fangzhou, et al.
Pubblicazione: (2024)
When Vision Speaks for Sound
di: Wen, Xiaofei, et al.
Pubblicazione: (2026)
di: Wen, Xiaofei, et al.
Pubblicazione: (2026)
MASS: Overcoming Language Bias in Image-Text Matching
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
MEVG: Multi-event Video Generation with Text-to-Video Models
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
di: Oh, Gyeongrok, et al.
Pubblicazione: (2023)
Mine-JEPA: In-Domain Self-Supervised Learning for Mine-Like Object Classification in Side-Scan Sonar
di: Kwon, Taeyoun, et al.
Pubblicazione: (2026)
di: Kwon, Taeyoun, et al.
Pubblicazione: (2026)
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
di: Jang, Youngjoon, et al.
Pubblicazione: (2024)
di: Jang, Youngjoon, et al.
Pubblicazione: (2024)
Inlier-Centric Post-Training Quantization for Object Detection Models
di: Kim, Minsu, et al.
Pubblicazione: (2026)
di: Kim, Minsu, et al.
Pubblicazione: (2026)
FastSTAR: Spatiotemporal Token Pruning for Efficient Autoregressive Video Synthesis
di: Yune, Sungwoong, et al.
Pubblicazione: (2026)
di: Yune, Sungwoong, et al.
Pubblicazione: (2026)
A11YN: aligning LLMs for accessible web UI code generation
di: Yoon, Janghan, et al.
Pubblicazione: (2025)
di: Yoon, Janghan, et al.
Pubblicazione: (2025)
VisEscape: A Benchmark for Evaluating Exploration-driven Decision-making in Virtual Escape Rooms
di: Lim, Seungwon, et al.
Pubblicazione: (2025)
di: Lim, Seungwon, et al.
Pubblicazione: (2025)
ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
di: Kim, Jeonghwan, et al.
Pubblicazione: (2024)
Chameleon: A Data-Efficient Generalist for Dense Visual Prediction in the Wild
di: Kim, Donggyun, et al.
Pubblicazione: (2024)
di: Kim, Donggyun, et al.
Pubblicazione: (2024)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
di: Lee, Jongseo, et al.
Pubblicazione: (2025)
CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
di: Kim, Jiwan, et al.
Pubblicazione: (2025)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
di: Shin, Youngwoo, et al.
Pubblicazione: (2026)
di: Shin, Youngwoo, et al.
Pubblicazione: (2026)
Scalp Diagnostic System With Label-Free Segmentation and Training-Free Image Translation
di: Kim, Youngmin, et al.
Pubblicazione: (2024)
di: Kim, Youngmin, et al.
Pubblicazione: (2024)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
di: Nam, Heejeong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
di: Kim, Youngmin, et al.
Pubblicazione: (2025) -
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
di: Chung, Jiwan, et al.
Pubblicazione: (2025) -
GuideDog: A Real-World Egocentric Multimodal Dataset for Blind and Low-Vision Accessibility-Aware Guidance
di: Kim, Junhyeok, et al.
Pubblicazione: (2025) -
DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face Animation
di: Kim, Jisoo, et al.
Pubblicazione: (2024) -
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
di: Cho, Jungbin, et al.
Pubblicazione: (2024)