| _version_ | 1866902054585237504 |
|---|---|
| author | Patel, Manav |
| author_facet | Patel, Manav |
| contents | <p>This paper presents VoiceDoc, an innovative voice-activated intelligent document assistant that<br>leverages Retrieval-Augmented Generation (RAG) to revolutionize information extraction from diverse<br>document formats. The system enables users to interact with documents through natural spoken language,<br>eliminating the need for manual navigation and text-based searches. VoiceDoc integrates advanced speech<br>recognition, large language models (specifically Groq Llama 3.1-70B), and natural text-to-speech synthesis<br>to create a seamless conversational interface. The platform supports multiple document formats<br>including PDF, DOCX, PPTX, and images, utilizing FAISS vector database for efficient semantic search<br>and retrieval. Our implementation demonstrates significant improvements in accessibility, information retrieval<br>speed, and user experience across various professional domains including healthcare, education,<br>legal services, and corporate environments. The system achieves real-time response capabilities while<br>maintaining high accuracy in contextually relevant answer generation, making it a practical solution for<br>knowledge-intensive tasks requiring hands-free operation.</p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_18322577 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | VoiceDoc: A Voice-Activated Intelligent Document Assistant Using Advanced RAG Technology Patel, Manav <p>This paper presents VoiceDoc, an innovative voice-activated intelligent document assistant that<br>leverages Retrieval-Augmented Generation (RAG) to revolutionize information extraction from diverse<br>document formats. The system enables users to interact with documents through natural spoken language,<br>eliminating the need for manual navigation and text-based searches. VoiceDoc integrates advanced speech<br>recognition, large language models (specifically Groq Llama 3.1-70B), and natural text-to-speech synthesis<br>to create a seamless conversational interface. The platform supports multiple document formats<br>including PDF, DOCX, PPTX, and images, utilizing FAISS vector database for efficient semantic search<br>and retrieval. Our implementation demonstrates significant improvements in accessibility, information retrieval<br>speed, and user experience across various professional domains including healthcare, education,<br>legal services, and corporate environments. The system achieves real-time response capabilities while<br>maintaining high accuracy in contextually relevant answer generation, making it a practical solution for<br>knowledge-intensive tasks requiring hands-free operation.</p> |
| title | VoiceDoc: A Voice-Activated Intelligent Document Assistant Using Advanced RAG Technology |
| url | https://doi.org/10.5281/zenodo.18322577 |