M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhe, Liu, Heyang, Yu, Wenyi, Sun, Guangzhi, Liu, Hongcheng, Wu, Ji, Zhang, Chao, Wang, Yu, Wang, Yanfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
von: Zhao, Zihan, et al.
Veröffentlicht: (2023)
von: Zhao, Zihan, et al.
Veröffentlicht: (2023)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Decoding Linguistic Representations of Human Brain
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
von: Chen, Zhe, et al.
Veröffentlicht: (2025)
Multigenre AI-powered Story Composition
von: de Lima, Edirlei Soares, et al.
Veröffentlicht: (2024)
von: de Lima, Edirlei Soares, et al.
Veröffentlicht: (2024)
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2026)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
LaSR: Context-Aware Speech Recognition via Latent Reasoning
von: Liu, Heyang, et al.
Veröffentlicht: (2026)
von: Liu, Heyang, et al.
Veröffentlicht: (2026)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
von: Xu, Ke, et al.
Veröffentlicht: (2026)
von: Xu, Ke, et al.
Veröffentlicht: (2026)
ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection
von: Khairallah, Ali, et al.
Veröffentlicht: (2025)
von: Khairallah, Ali, et al.
Veröffentlicht: (2025)
AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
von: Gong, Kaixiong, et al.
Veröffentlicht: (2024)
von: Gong, Kaixiong, et al.
Veröffentlicht: (2024)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
von: Wang, Pingjie, et al.
Veröffentlicht: (2025)
Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs
von: Alghallabi, Wafa, et al.
Veröffentlicht: (2025)
von: Alghallabi, Wafa, et al.
Veröffentlicht: (2025)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
Time-Scaling Is What Agents Need Now
von: Liu, Zhi, et al.
Veröffentlicht: (2026)
von: Liu, Zhi, et al.
Veröffentlicht: (2026)
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
von: Yu, Wenyi, et al.
Veröffentlicht: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
MedS$^3$: Towards Medical Slow Thinking with Self-Evolved Soft Dual-sided Process Supervision
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
von: Jiang, Shuyang, et al.
Veröffentlicht: (2025)
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
OCR-Enhanced Multimodal ASR Can Read While Listening
von: Chen, Junli, et al.
Veröffentlicht: (2026)
von: Chen, Junli, et al.
Veröffentlicht: (2026)
SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
von: Yu, Wenyi, et al.
Veröffentlicht: (2024)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
von: Li, Yiqi, et al.
Veröffentlicht: (2025)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
Guided by the Plan: Enhancing Faithful Autoregressive Text-to-Audio Generation with Guided Decoding
von: Wang, Juncheng, et al.
Veröffentlicht: (2026)
von: Wang, Juncheng, et al.
Veröffentlicht: (2026)
Leveraging Diverse Modeling Contexts with Collaborating Learning for Neural Machine Translation
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
von: Liao, Yusheng, et al.
Veröffentlicht: (2024)
SALMONN: Towards Generic Hearing Abilities for Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2023)
von: Tang, Changli, et al.
Veröffentlicht: (2023)
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction
von: Liu, Jiang, et al.
Veröffentlicht: (2024)
von: Liu, Jiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024) -
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024) -
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
von: Liu, Hongcheng, et al.
Veröffentlicht: (2025) -
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
von: Liu, Hongcheng, et al.
Veröffentlicht: (2024) -
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)