Allo-AVA: A Large-Scale Multimodal Conversational AI Dataset for Allocentric Avatar Gesture Animation
Fuente:
arXiv
Saved in:
| Main Authors: | Punjwani, Saif, Heck, Larry |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Body Language Models
by: Punjwani, Saif, et al.
Published: (2024)
by: Punjwani, Saif, et al.
Published: (2024)
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
by: Punjwani, Saif, et al.
Published: (2025)
by: Punjwani, Saif, et al.
Published: (2025)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
by: Reichman, Benjamin, et al.
Published: (2025)
by: Reichman, Benjamin, et al.
Published: (2025)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
by: Gu, Hexiang, et al.
Published: (2025)
by: Gu, Hexiang, et al.
Published: (2025)
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
by: Wang, Hengyi, et al.
Published: (2026)
by: Wang, Hengyi, et al.
Published: (2026)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
by: Khan, Shaharukh, et al.
Published: (2026)
by: Khan, Shaharukh, et al.
Published: (2026)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
by: Kondic, Jovana, et al.
Published: (2026)
by: Kondic, Jovana, et al.
Published: (2026)
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
by: Yerukola, Akhila, et al.
Published: (2025)
by: Yerukola, Akhila, et al.
Published: (2025)
Advancing Conversational Diagnostic AI with Multimodal Reasoning
by: Saab, Khaled, et al.
Published: (2025)
by: Saab, Khaled, et al.
Published: (2025)
Web2Code: A Large-scale Webpage-to-Code Dataset and Evaluation Framework for Multimodal LLMs
by: Yun, Sukmin, et al.
Published: (2024)
by: Yun, Sukmin, et al.
Published: (2024)
Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars
by: Zhang, Youliang, et al.
Published: (2026)
by: Zhang, Youliang, et al.
Published: (2026)
RingGesture: A Ring-Based Mid-Air Gesture Typing System Powered by a Deep-Learning Word Prediction Framework
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
Fast Registration of Photorealistic Avatars for VR Facial Animation
by: Patel, Chaitanya, et al.
Published: (2024)
by: Patel, Chaitanya, et al.
Published: (2024)
RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
by: Butsanets, Léo, et al.
Published: (2025)
by: Butsanets, Léo, et al.
Published: (2025)
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers
by: Pramanick, Shraman, et al.
Published: (2024)
by: Pramanick, Shraman, et al.
Published: (2024)
Large Multimodal Agents: A Survey
by: Xie, Junlin, et al.
Published: (2024)
by: Xie, Junlin, et al.
Published: (2024)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
by: Han, Songhao, et al.
Published: (2024)
by: Han, Songhao, et al.
Published: (2024)
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
by: Li, Zekun, et al.
Published: (2024)
by: Li, Zekun, et al.
Published: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection
by: Zhang, Huixuan, et al.
Published: (2023)
by: Zhang, Huixuan, et al.
Published: (2023)
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
by: Jang, Jihyoung, et al.
Published: (2025)
by: Jang, Jihyoung, et al.
Published: (2025)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Towards Conversational Medical AI with Eyes, Ears and a Voice
by: Shah, Meet, et al.
Published: (2026)
by: Shah, Meet, et al.
Published: (2026)
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025)
by: Lu, Xingyuan, et al.
Published: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
A Survey on Evaluation of Multimodal Large Language Models
by: Huang, Jiaxing, et al.
Published: (2024)
by: Huang, Jiaxing, et al.
Published: (2024)
A Survey on Agentic Multimodal Large Language Models
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Hierarchy-Aware Multimodal Unlearning for Medical AI
by: Wu, Fengli, et al.
Published: (2025)
by: Wu, Fengli, et al.
Published: (2025)
Graphic Design with Large Multimodal Model
by: Cheng, Yutao, et al.
Published: (2024)
by: Cheng, Yutao, et al.
Published: (2024)
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
by: Zhang, Zilun, et al.
Published: (2023)
by: Zhang, Zilun, et al.
Published: (2023)
Representing Animatable Avatar via Factorized Neural Fields
by: Song, Chunjin, et al.
Published: (2024)
by: Song, Chunjin, et al.
Published: (2024)
Evaluating Multimodal Generative AI with Korean Educational Standards
by: Park, Sanghee, et al.
Published: (2025)
by: Park, Sanghee, et al.
Published: (2025)
Model Composition for Multimodal Large Language Models
by: Chen, Chi, et al.
Published: (2024)
by: Chen, Chi, et al.
Published: (2024)
M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
by: Anugraha, David, et al.
Published: (2025)
by: Anugraha, David, et al.
Published: (2025)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
FaceLLM: A Multimodal Large Language Model for Face Understanding
by: Shahreza, Hatef Otroshi, et al.
Published: (2025)
by: Shahreza, Hatef Otroshi, et al.
Published: (2025)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
by: Li, Wenbin, et al.
Published: (2026)
by: Li, Wenbin, et al.
Published: (2026)
Similar Items
-
Large Body Language Models
by: Punjwani, Saif, et al.
Published: (2024) -
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
by: Punjwani, Saif, et al.
Published: (2025) -
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
by: Reichman, Benjamin, et al.
Published: (2025) -
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
by: Gu, Hexiang, et al.
Published: (2025) -
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
by: Kim, Youngmin, et al.
Published: (2025)