FOCAL: A Novel Benchmarking Technique for Multi-modal Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Purwar, Anupam, Choudhary, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
by: Purwar, Anupam, et al.
Published: (2025)
by: Purwar, Anupam, et al.
Published: (2025)
When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
by: Purwar, Anupam, et al.
Published: (2026)
by: Purwar, Anupam, et al.
Published: (2026)
MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings
by: Purwar, Anupam, et al.
Published: (2026)
by: Purwar, Anupam, et al.
Published: (2026)
Toward Noise-Aware Audio Deepfake Detection: Survey, SNR-Benchmarks, and Practical Recipes
by: Sen, Udayon, et al.
Published: (2025)
by: Sen, Udayon, et al.
Published: (2025)
Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
by: Lakshminarasimhan, Aditya
Published: (2024)
by: Lakshminarasimhan, Aditya
Published: (2024)
An LLM Benchmark for Addressee Recognition in Multi-modal Multi-party Dialogue
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
Synaspot: A Lightweight, Streaming Multi-modal Framework for Keyword Spotting with Audio-Text Synergy
by: Li, Kewei, et al.
Published: (2025)
by: Li, Kewei, et al.
Published: (2025)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
by: Zhou, Dingkun, et al.
Published: (2025)
by: Zhou, Dingkun, et al.
Published: (2025)
Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
by: Li, Xueyan, et al.
Published: (2025)
by: Li, Xueyan, et al.
Published: (2025)
Multi-modal expressive personality recognition in data non-ideal audiovisual based on multi-scale feature enhancement and modal augment
by: Kong, Weixuan, et al.
Published: (2025)
by: Kong, Weixuan, et al.
Published: (2025)
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
by: Wang, Xiaopeng, et al.
Published: (2025)
by: Wang, Xiaopeng, et al.
Published: (2025)
Multi-modal Speech Enhancement with Limited Electromyography Channels
by: Feng, Fuyuan, et al.
Published: (2025)
by: Feng, Fuyuan, et al.
Published: (2025)
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
by: Feng, Tiantian, et al.
Published: (2026)
by: Feng, Tiantian, et al.
Published: (2026)
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
by: Maltais, Marie, et al.
Published: (2026)
by: Maltais, Marie, et al.
Published: (2026)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
by: Mu, Bingshen, et al.
Published: (2024)
by: Mu, Bingshen, et al.
Published: (2024)
Shared Multi-modal Embedding Space for Face-Voice Association
by: Simic, Christopher, et al.
Published: (2025)
by: Simic, Christopher, et al.
Published: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
by: Zhou, Yang-Hao, et al.
Published: (2026)
by: Zhou, Yang-Hao, et al.
Published: (2026)
DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
by: Yin, Kang, et al.
Published: (2025)
by: Yin, Kang, et al.
Published: (2025)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
by: Avro, Shamin Bin Habib, et al.
Published: (2025)
by: Avro, Shamin Bin Habib, et al.
Published: (2025)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
TART: A Comprehensive Tool for Technique-Aware Audio-to-Tab Guitar Transcription
by: Gupta, Akshaj, et al.
Published: (2025)
by: Gupta, Akshaj, et al.
Published: (2025)
Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
by: Shrestha, Aayush M., et al.
Published: (2026)
by: Shrestha, Aayush M., et al.
Published: (2026)
Robust Multi-modal Task-oriented Communications with Redundancy-aware Representations
by: Fu, Jingwen, et al.
Published: (2025)
by: Fu, Jingwen, et al.
Published: (2025)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
by: Ray, Soham, et al.
Published: (2026)
by: Ray, Soham, et al.
Published: (2026)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
by: Guan, Wenhao, et al.
Published: (2023)
by: Guan, Wenhao, et al.
Published: (2023)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
by: Janiczek, John, et al.
Published: (2024)
by: Janiczek, John, et al.
Published: (2024)
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents
by: Hu, Jiliang, et al.
Published: (2025)
by: Hu, Jiliang, et al.
Published: (2025)
MSRBench: A Benchmarking Dataset for Music Source Restoration
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Neural Steerer: Novel Steering Vector Synthesis with a Causal Neural Field over Frequency and Source Positions
by: Di Carlo, Diego, et al.
Published: (2023)
by: Di Carlo, Diego, et al.
Published: (2023)
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
by: Ai, Zhiqi, et al.
Published: (2026)
by: Ai, Zhiqi, et al.
Published: (2026)
pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
by: Devaney, Johanna, et al.
Published: (2024)
by: Devaney, Johanna, et al.
Published: (2024)
MMSD-Net: Towards Multi-modal Stuttering Detection
by: Nie, Liangyu, et al.
Published: (2024)
by: Nie, Liangyu, et al.
Published: (2024)
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
by: Xue, Liumeng, et al.
Published: (2026)
by: Xue, Liumeng, et al.
Published: (2026)
A Benchmark for Multi-speaker Anonymization
by: Miao, Xiaoxiao, et al.
Published: (2024)
by: Miao, Xiaoxiao, et al.
Published: (2024)
HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition
by: Dutta, Soumya, et al.
Published: (2023)
by: Dutta, Soumya, et al.
Published: (2023)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
TechSinger: Technique Controllable Multilingual Singing Voice Synthesis via Flow Matching
by: Guo, Wenxiang, et al.
Published: (2025)
by: Guo, Wenxiang, et al.
Published: (2025)
Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
by: Choudhary, Yash, et al.
Published: (2025)
by: Choudhary, Yash, et al.
Published: (2025)
Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
by: Choudhary, Yash, et al.
Published: (2025)
by: Choudhary, Yash, et al.
Published: (2025)
Similar Items
-
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
by: Purwar, Anupam, et al.
Published: (2025) -
When Fine-Tuning Fails and when it Generalises: Role of Data Diversity and Mixed Training in LLM-based TTS
by: Purwar, Anupam, et al.
Published: (2026) -
MM-tau-p$^2$: Persona-Adaptive Prompting for Robust Multi-Modal Agent Evaluation in Dual-Control Settings
by: Purwar, Anupam, et al.
Published: (2026) -
Toward Noise-Aware Audio Deepfake Detection: Survey, SNR-Benchmarks, and Practical Recipes
by: Sen, Udayon, et al.
Published: (2025) -
Musical Chords: A Novel Java Algorithm and App Utility to Enumerate Chord-Progressions Adhering to Music Theory Guidelines
by: Lakshminarasimhan, Aditya
Published: (2024)