Maia: A Real-time Non-Verbal Chat for Human-AI Interaction
Fuente:
arXiv
Saved in:
| Main Authors: | Costea, Dragos, Marcu, Alina, Lazar, Cristina, Leordeanu, Marius |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
by: Costea, Dragos, et al.
Published: (2026)
by: Costea, Dragos, et al.
Published: (2026)
A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
by: Costea, Dragos, et al.
Published: (2025)
by: Costea, Dragos, et al.
Published: (2025)
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
by: Constantinescu, Ciprian, et al.
Published: (2025)
by: Constantinescu, Ciprian, et al.
Published: (2025)
Multiple Random Masking Autoencoder Ensembles for Robust Multimodal Semi-supervised Learning
by: Todoran, Alexandru-Raul, et al.
Published: (2024)
by: Todoran, Alexandru-Raul, et al.
Published: (2024)
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)
Quantifying the synthetic and real domain gap in aerial scene understanding
by: Marcu, Alina
Published: (2024)
by: Marcu, Alina
Published: (2024)
Efficient Self-Supervised Neuro-Analytic Visual Servoing for Real-time Quadrotor Control
by: Mocanu, Sebastian, et al.
Published: (2025)
by: Mocanu, Sebastian, et al.
Published: (2025)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
by: Masala, Mihai, et al.
Published: (2025)
by: Masala, Mihai, et al.
Published: (2025)
Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
by: Cudlenco, Nicolae, et al.
Published: (2026)
by: Cudlenco, Nicolae, et al.
Published: (2026)
GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
by: Cudlenco, Nicolae, et al.
Published: (2026)
by: Cudlenco, Nicolae, et al.
Published: (2026)
Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs
by: Mocanu, Sebastian, et al.
Published: (2025)
by: Mocanu, Sebastian, et al.
Published: (2025)
Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
by: Jelea, Andrei, et al.
Published: (2025)
by: Jelea, Andrei, et al.
Published: (2025)
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
by: Jelea, Andrei, et al.
Published: (2025)
by: Jelea, Andrei, et al.
Published: (2025)
Multi-modal video data-pipelines for machine learning with minimal human supervision
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
by: Pîrvu, Mihai-Cristian, et al.
Published: (2025)
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
by: Airinei, Daniel, et al.
Published: (2025)
by: Airinei, Daniel, et al.
Published: (2025)
Iterative Explainability for Weakly Supervised Segmentation in Medical PE Detection
by: Condrea, Florin, et al.
Published: (2024)
by: Condrea, Florin, et al.
Published: (2024)
Learning on the Fly: Replay-Based Continual Object Perception for Indoor Drones
by: Nae, Sebastian-Ion, et al.
Published: (2026)
by: Nae, Sebastian-Ion, et al.
Published: (2026)
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
FloodLense: A Framework for ChatGPT-based Real-time Flood Detection
by: Kumbam, Pranath Reddy, et al.
Published: (2024)
by: Kumbam, Pranath Reddy, et al.
Published: (2024)
ChatPose: Chatting about 3D Human Pose
by: Feng, Yao, et al.
Published: (2023)
by: Feng, Yao, et al.
Published: (2023)
ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
by: Qi, Jinwei, et al.
Published: (2025)
by: Qi, Jinwei, et al.
Published: (2025)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
ChatHuman: Chatting about 3D Humans with Tools
by: Lin, Jing, et al.
Published: (2024)
by: Lin, Jing, et al.
Published: (2024)
HAIFAI: Human-AI Interaction for Mental Face Reconstruction
by: Strohm, Florian, et al.
Published: (2024)
by: Strohm, Florian, et al.
Published: (2024)
Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues
by: Gimeno-Gómez, David, et al.
Published: (2024)
by: Gimeno-Gómez, David, et al.
Published: (2024)
A Real-Time Human Action Recognition Model for Assisted Living
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
SARAH: Spatially Aware Real-time Agentic Humans
by: Ng, Evonne, et al.
Published: (2026)
by: Ng, Evonne, et al.
Published: (2026)
Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cues
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
LongLive: Real-time Interactive Long Video Generation
by: Yang, Shuai, et al.
Published: (2025)
by: Yang, Shuai, et al.
Published: (2025)
ARIG: Autoregressive Interactive Head Generation for Real-time Conversations
by: Guo, Ying, et al.
Published: (2025)
by: Guo, Ying, et al.
Published: (2025)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)
by: Yerukola, Akhila, et al.
Published: (2025)
DiffChat: Learning to Chat with Text-to-Image Synthesis Models for Interactive Image Creation
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customization
by: Song, Quanjian, et al.
Published: (2026)
by: Song, Quanjian, et al.
Published: (2026)
Anatomically aware dual-hop learning for pulmonary embolism detection in CT pulmonary angiograms
by: Condrea, Florin, et al.
Published: (2023)
by: Condrea, Florin, et al.
Published: (2023)
WeGen: A Unified Model for Interactive Multimodal Generation as We Chat
by: Huang, Zhipeng, et al.
Published: (2025)
by: Huang, Zhipeng, et al.
Published: (2025)
ClickTrack: Towards Real-time Interactive Single Object Tracking
by: Wang, Kuiran, et al.
Published: (2024)
by: Wang, Kuiran, et al.
Published: (2024)
ReMoGen: Real-time Human Interaction-to-Reaction Generation via Modular Learning from Diverse Data
by: Ye, Yaoqin, et al.
Published: (2026)
by: Ye, Yaoqin, et al.
Published: (2026)
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
by: He, Chaoqun, et al.
Published: (2026)
by: He, Chaoqun, et al.
Published: (2026)
Similar Items
-
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
by: Costea, Dragos, et al.
Published: (2026) -
A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
by: Costea, Dragos, et al.
Published: (2025) -
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
by: Constantinescu, Ciprian, et al.
Published: (2025) -
Multiple Random Masking Autoencoder Ensembles for Robust Multimodal Semi-supervised Learning
by: Todoran, Alexandru-Raul, et al.
Published: (2024) -
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
by: Mihai-Cristian, Pîrvu, et al.
Published: (2025)