Beamforming-LLM: What, Where and When Did I Miss?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Choudhari, Vishal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916939172937728
author Choudhari, Vishal
author_facet Choudhari, Vishal
contents We present Beamforming-LLM, a system that enables users to semantically recall conversations they may have missed in multi-speaker environments. The system combines spatial audio capture using a microphone array with retrieval-augmented generation (RAG) to support natural language queries such as, "What did I miss when I was following the conversation on dogs?" Directional audio streams are separated using beamforming, transcribed with Whisper, and embedded into a vector database using sentence encoders. Upon receiving a user query, semantically relevant segments are retrieved, temporally aligned with non-attended segments, and summarized using a lightweight large language model (GPT-4o-mini). The result is a user-friendly interface that provides contrastive summaries, spatial context, and timestamped audio playback. This work lays the foundation for intelligent auditory memory systems and has broad applications in assistive technology, meeting summarization, and context-aware personal spatial computing.
format Preprint
id arxiv_https___arxiv_org_abs_2509_06221
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beamforming-LLM: What, Where and When Did I Miss?
Choudhari, Vishal
Audio and Speech Processing
Artificial Intelligence
Computation and Language
Human-Computer Interaction
We present Beamforming-LLM, a system that enables users to semantically recall conversations they may have missed in multi-speaker environments. The system combines spatial audio capture using a microphone array with retrieval-augmented generation (RAG) to support natural language queries such as, "What did I miss when I was following the conversation on dogs?" Directional audio streams are separated using beamforming, transcribed with Whisper, and embedded into a vector database using sentence encoders. Upon receiving a user query, semantically relevant segments are retrieved, temporally aligned with non-attended segments, and summarized using a lightweight large language model (GPT-4o-mini). The result is a user-friendly interface that provides contrastive summaries, spatial context, and timestamped audio playback. This work lays the foundation for intelligent auditory memory systems and has broad applications in assistive technology, meeting summarization, and context-aware personal spatial computing.
title Beamforming-LLM: What, Where and When Did I Miss?
topic Audio and Speech Processing
Artificial Intelligence
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2509.06221