PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Yujia, Xue, Liumeng, He, Lei, Chen, Xinyi, Chiu, Aemon Yat Fei, Tian, Wenjie, Zhang, Shaofei, Kong, Qiuqiang, Zhu, Xinfa, Xue, Wei, Lee, Tan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914069758345216
author Xiao, Yujia
Xue, Liumeng
He, Lei
Chen, Xinyi
Chiu, Aemon Yat Fei
Tian, Wenjie
Zhang, Shaofei
Kong, Qiuqiang
Zhu, Xinfa
Xue, Wei
Lee, Tan
author_facet Xiao, Yujia
Xue, Liumeng
He, Lei
Chen, Xinyi
Chiu, Aemon Yat Fei
Tian, Wenjie
Zhang, Shaofei
Kong, Qiuqiang
Zhu, Xinfa
Xue, Wei
Lee, Tan
contents Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited, especially for open-ended long-form content generation. Significant challenges lie in no reference standard answer, no unified evaluation metrics and uncontrollable human judgments. In this work, we take podcast-like audio generation as a starting point and propose PodEval, a comprehensive and well-designed open-source evaluation framework. In this framework: 1) We construct a real-world podcast dataset spanning diverse topics, serving as a reference for human-level creative quality. 2) We introduce a multimodal evaluation strategy and decompose the complex task into three dimensions: text, speech and audio, with different evaluation emphasis on "Content" and "Format". 3) For each modality, we design corresponding evaluation methods, involving both objective metrics and subjective listening test. We leverage representative podcast generation systems (including open-source, close-source, and human-made) in our experiments. The results offer in-depth analysis and insights into podcast generation, demonstrating the effectiveness of PodEval in evaluating open-ended long-form audio. This project is open-source to facilitate public use: https://github.com/yujxx/PodEval.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00485
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
Xiao, Yujia
Xue, Liumeng
He, Lei
Chen, Xinyi
Chiu, Aemon Yat Fei
Tian, Wenjie
Zhang, Shaofei
Kong, Qiuqiang
Zhu, Xinfa
Xue, Wei
Lee, Tan
Sound
Artificial Intelligence
Audio and Speech Processing
Recently, an increasing number of multimodal (text and audio) benchmarks have emerged, primarily focusing on evaluating models' understanding capability. However, exploration into assessing generative capabilities remains limited, especially for open-ended long-form content generation. Significant challenges lie in no reference standard answer, no unified evaluation metrics and uncontrollable human judgments. In this work, we take podcast-like audio generation as a starting point and propose PodEval, a comprehensive and well-designed open-source evaluation framework. In this framework: 1) We construct a real-world podcast dataset spanning diverse topics, serving as a reference for human-level creative quality. 2) We introduce a multimodal evaluation strategy and decompose the complex task into three dimensions: text, speech and audio, with different evaluation emphasis on "Content" and "Format". 3) For each modality, we design corresponding evaluation methods, involving both objective metrics and subjective listening test. We leverage representative podcast generation systems (including open-source, close-source, and human-made) in our experiments. The results offer in-depth analysis and insights into podcast generation, demonstrating the effectiveness of PodEval in evaluating open-ended long-form audio. This project is open-source to facilitate public use: https://github.com/yujxx/PodEval.
title PodEval: A Multimodal Evaluation Framework for Podcast Audio Generation
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2510.00485