OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Qiang, Luo, Yuanyi, Li, Sirui, Zhang, Wenxiao, Liu, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
di: Cheng, Kanzhi, et al.
Pubblicazione: (2026)
di: Cheng, Kanzhi, et al.
Pubblicazione: (2026)
Veracity: An Open-Source AI Fact-Checking System
di: Curtis, Taylor Lynn, et al.
Pubblicazione: (2025)
di: Curtis, Taylor Lynn, et al.
Pubblicazione: (2025)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
di: Yang, Bufang, et al.
Pubblicazione: (2025)
di: Yang, Bufang, et al.
Pubblicazione: (2025)
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
di: Kim, Taesoo, et al.
Pubblicazione: (2025)
Civil Society in the Loop: Feedback-Driven Adaptation of (L)LM-Assisted Classification in an Open-Source Telegram Monitoring Tool
di: Pustet, Milena, et al.
Pubblicazione: (2025)
di: Pustet, Milena, et al.
Pubblicazione: (2025)
AgentCTG: Harnessing Multi-Agent Collaboration for Fine-Grained Precise Control in Text Generation
di: Zhou, Xinxu, et al.
Pubblicazione: (2025)
di: Zhou, Xinxu, et al.
Pubblicazione: (2025)
CUIfy the XR: An Open-Source Package to Embed LLM-powered Conversational Agents in XR
di: Buldu, Kadir Burak, et al.
Pubblicazione: (2024)
di: Buldu, Kadir Burak, et al.
Pubblicazione: (2024)
Role-Play Zero-Shot Prompting with Large Language Models for Open-Domain Human-Machine Conversation
di: Njifenjou, Ahmed, et al.
Pubblicazione: (2024)
di: Njifenjou, Ahmed, et al.
Pubblicazione: (2024)
Open-Source Conversational AI with SpeechBrain 1.0
di: Ravanelli, Mirco, et al.
Pubblicazione: (2024)
di: Ravanelli, Mirco, et al.
Pubblicazione: (2024)
Hallucinations and Key Information Extraction in Medical Texts: A Comprehensive Assessment of Open-Source Large Language Models
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2025)
di: Das, Anindya Bijoy, et al.
Pubblicazione: (2025)
ReSpAct: Harmonizing Reasoning, Speaking, and Acting Towards Building Large Language Model-Based Conversational AI Agents
di: Dongre, Vardhan, et al.
Pubblicazione: (2024)
di: Dongre, Vardhan, et al.
Pubblicazione: (2024)
Thematic Analysis with Open-Source Generative AI and Machine Learning: A New Method for Inductive Qualitative Codebook Development
di: Katz, Andrew, et al.
Pubblicazione: (2024)
di: Katz, Andrew, et al.
Pubblicazione: (2024)
Multimodal Transformer Models for Turn-taking Prediction: Effects on Conversational Dynamics of Human-Agent Interaction during Cooperative Gameplay
di: Bae, Young-Ho, et al.
Pubblicazione: (2025)
di: Bae, Young-Ho, et al.
Pubblicazione: (2025)
Collaborative Gym: A Framework for Enabling and Evaluating Human-Agent Collaboration
di: Shao, Yijia, et al.
Pubblicazione: (2024)
di: Shao, Yijia, et al.
Pubblicazione: (2024)
Conversational AI Multi-Agent Interoperability, Universal Open APIs for Agentic Natural Language Multimodal Communications
di: Gosmar, Diego, et al.
Pubblicazione: (2024)
di: Gosmar, Diego, et al.
Pubblicazione: (2024)
Are Today's LLMs Ready to Explain Well-Being Concepts?
di: Jiang, Bohan, et al.
Pubblicazione: (2025)
di: Jiang, Bohan, et al.
Pubblicazione: (2025)
Open-Source Large Language Models as Multilingual Crowdworkers: Synthesizing Open-Domain Dialogues in Several Languages With No Examples in Targets and No Machine Translation
di: Njifenjou, Ahmed, et al.
Pubblicazione: (2025)
di: Njifenjou, Ahmed, et al.
Pubblicazione: (2025)
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
di: Liu, Yuhang, et al.
Pubblicazione: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
di: Kapoor, Raghav, et al.
Pubblicazione: (2024)
Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations
di: Stacchio, Lorenzo, et al.
Pubblicazione: (2025)
di: Stacchio, Lorenzo, et al.
Pubblicazione: (2025)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
di: Wang, Kangyu, et al.
Pubblicazione: (2025)
di: Wang, Kangyu, et al.
Pubblicazione: (2025)
Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaboration
di: Yao, Bingsheng, et al.
Pubblicazione: (2025)
di: Yao, Bingsheng, et al.
Pubblicazione: (2025)
Improving Interactive Diagnostic Ability of a Large Language Model Agent Through Clinical Experience Learning
di: Sun, Zhoujian, et al.
Pubblicazione: (2025)
di: Sun, Zhoujian, et al.
Pubblicazione: (2025)
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
di: Wang, Yiyang, et al.
Pubblicazione: (2026)
di: Wang, Yiyang, et al.
Pubblicazione: (2026)
How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses
di: Lin, Jionghao, et al.
Pubblicazione: (2024)
di: Lin, Jionghao, et al.
Pubblicazione: (2024)
Benchmarking LLM Tool-Use in the Wild
di: Yu, Peijie, et al.
Pubblicazione: (2026)
di: Yu, Peijie, et al.
Pubblicazione: (2026)
Towards End-to-End Open Conversational Machine Reading
di: Zhou, Sizhe, et al.
Pubblicazione: (2022)
di: Zhou, Sizhe, et al.
Pubblicazione: (2022)
OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
di: Luo, Cheng, et al.
Pubblicazione: (2025)
di: Luo, Cheng, et al.
Pubblicazione: (2025)
CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation
di: Huq, Faria, et al.
Pubblicazione: (2025)
di: Huq, Faria, et al.
Pubblicazione: (2025)
From Control to Foresight: Simulation as a New Paradigm for Human-Agent Collaboration
di: He, Gaole, et al.
Pubblicazione: (2026)
di: He, Gaole, et al.
Pubblicazione: (2026)
Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
di: La Cava, Lucio, et al.
Pubblicazione: (2024)
Self-Reported Confidence of Large Language Models in Gastroenterology: Analysis of Commercial, Open-Source, and Quantized Models
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
di: Naderi, Nariman, et al.
Pubblicazione: (2025)
The Future of Open Human Feedback
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
di: Zhang, Zhuosheng, et al.
Pubblicazione: (2023)
See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
di: Wu, Zongru, et al.
Pubblicazione: (2025)
di: Wu, Zongru, et al.
Pubblicazione: (2025)
Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
di: Ahmad, Adnan, et al.
Pubblicazione: (2025)
di: Ahmad, Adnan, et al.
Pubblicazione: (2025)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
di: Shin, Jisu, et al.
Pubblicazione: (2025)
di: Shin, Jisu, et al.
Pubblicazione: (2025)
Question Answering for Decisionmaking in Green Building Design: A Multimodal Data Reasoning Method Driven by Large Language Models
di: Li, Yihui, et al.
Pubblicazione: (2024)
di: Li, Yihui, et al.
Pubblicazione: (2024)
Open TutorAI: An Open-source Platform for Personalized and Immersive Learning with Generative AI
di: Hajji, Mohamed El, et al.
Pubblicazione: (2026)
di: Hajji, Mohamed El, et al.
Pubblicazione: (2026)
Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends
di: Wang, Ye, et al.
Pubblicazione: (2026)
di: Wang, Ye, et al.
Pubblicazione: (2026)
Documenti analoghi
-
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
di: Cheng, Kanzhi, et al.
Pubblicazione: (2026) -
Veracity: An Open-Source AI Fact-Checking System
di: Curtis, Taylor Lynn, et al.
Pubblicazione: (2025) -
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
di: Yang, Bufang, et al.
Pubblicazione: (2025) -
Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
di: Kim, Taesoo, et al.
Pubblicazione: (2025) -
Civil Society in the Loop: Feedback-Driven Adaptation of (L)LM-Assisted Classification in an Open-Source Telegram Monitoring Tool
di: Pustet, Milena, et al.
Pubblicazione: (2025)