SpeechEE: A Novel Benchmark for Speech Event Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Bin, Zhang, Meishan, Fei, Hao, Zhao, Yu, Li, Bobo, Wu, Shengqiong, Ji, Wei, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
von: Zhang, Meishan, et al.
Veröffentlicht: (2024)
von: Zhang, Meishan, et al.
Veröffentlicht: (2024)
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
von: Xing, Fuyu, et al.
Veröffentlicht: (2025)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)
Can We Hear from Events? Generating Speech from Event Camera
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
von: Fang, Jingping, et al.
Veröffentlicht: (2026)
Double Mixture: Towards Continual Event Detection from Speech
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
Grammar Induction from Visual, Speech and Text
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
von: Zhao, Yu, et al.
Veröffentlicht: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
SpeechCraft: A Fine-grained Expressive Speech Dataset with Natural Language Description
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
von: Jin, Zeyu, et al.
Veröffentlicht: (2024)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
MMoFusion: Multi-modal Co-Speech Motion Generation with Diffusion Model
von: Wang, Sen, et al.
Veröffentlicht: (2024)
von: Wang, Sen, et al.
Veröffentlicht: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
von: Lu, Wenhuan, et al.
Veröffentlicht: (2025)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Stepwise Schema-Guided Prompting Framework with Parameter Efficient Instruction Tuning for Multimedia Event Extraction
von: Yuan, Xiang, et al.
Veröffentlicht: (2025)
von: Yuan, Xiang, et al.
Veröffentlicht: (2025)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
von: Hao, Bowen, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
von: Cui, Yang, et al.
Veröffentlicht: (2025)
von: Cui, Yang, et al.
Veröffentlicht: (2025)
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
A Survey on Speech Deepfake Detection
von: Li, Menglu, et al.
Veröffentlicht: (2024)
von: Li, Menglu, et al.
Veröffentlicht: (2024)
COPA: Efficient Vision-Language Pre-training Through Collaborative Object- and Patch-Text Alignment
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
von: Wu, Linzhi, et al.
Veröffentlicht: (2026)
EEmo-Bench: A Benchmark for Multi-modal Large Language Models on Image Evoked Emotion Assessment
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
von: Gao, Lancheng, et al.
Veröffentlicht: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
von: Zhou, Dongliang, et al.
Veröffentlicht: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
von: Wang, Siyu, et al.
Veröffentlicht: (2025)
Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2021)
von: Bhattacharya, Uttaran, et al.
Veröffentlicht: (2021)
Target Speech Diarization with Multimodal Prompts
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
SIDQL: An Efficient Keyframe Extraction and Motion Reconstruction Framework in Motion Capture
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
von: Zhang, Xuling, et al.
Veröffentlicht: (2024)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Recognizing Everything from All Modalities at Once: Grounded Multimodal Universal Information Extraction
von: Zhang, Meishan, et al.
Veröffentlicht: (2024) -
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024) -
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025) -
Benchmarking and Improving LVLMs on Event Extraction from Multimedia Documents
von: Xing, Fuyu, et al.
Veröffentlicht: (2025) -
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
von: Wang, Yongqi, et al.
Veröffentlicht: (2025)