MAC-SLU: Multi-Intent Automotive Cabin Spoken Language Understanding Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Peng, Yuezhang, Cai, Chonghao, Liu, Ziang, Fan, Shuai, Jiang, Sheng, Xu, Hua, Liu, Yuxin, Chen, Qiguang, Xu, Kele, Li, Yao, Wang, Sheng, Qin, Libo, Chen, Xie |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
by: Sheng, Zhichao, et al.
Published: (2025)
by: Sheng, Zhichao, et al.
Published: (2025)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
by: Peng, Yuezhang, et al.
Published: (2025)
by: Peng, Yuezhang, et al.
Published: (2025)
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026)
by: Zhou, Qianrui, et al.
Published: (2026)
Dual Dynamic Threshold Adjustment Strategy for Deep Metric Learning
by: Jiang, Xiruo, et al.
Published: (2024)
by: Jiang, Xiruo, et al.
Published: (2024)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
by: Zhang, Hanlei, et al.
Published: (2024)
by: Zhang, Hanlei, et al.
Published: (2024)
Privacy-Preserving Multimedia Mobile Cloud Computing Using Protective Perturbation
by: Tang, Zhongze, et al.
Published: (2024)
by: Tang, Zhongze, et al.
Published: (2024)
Feature Coding in the Era of Large Models: Dataset, Test Conditions, and Benchmark
by: Gao, Changsheng, et al.
Published: (2024)
by: Gao, Changsheng, et al.
Published: (2024)
Evaluating Multimodal Large Language Models on Spoken Sarcasm Understanding
by: Li, Zhu, et al.
Published: (2025)
by: Li, Zhu, et al.
Published: (2025)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
by: He, Jianfeng, et al.
Published: (2023)
by: He, Jianfeng, et al.
Published: (2023)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2023)
by: Zhou, Qianrui, et al.
Published: (2023)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2025)
by: Zhou, Qianrui, et al.
Published: (2025)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
by: Zhang, Lei, et al.
Published: (2025)
by: Zhang, Lei, et al.
Published: (2025)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
by: Su, Yi, et al.
Published: (2025)
by: Su, Yi, et al.
Published: (2025)
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios
by: Yuan, Ziqi, et al.
Published: (2024)
by: Yuan, Ziqi, et al.
Published: (2024)
IBMEA: Exploring Variational Information Bottleneck for Multi-modal Entity Alignment
by: Su, Taoyu, et al.
Published: (2024)
by: Su, Taoyu, et al.
Published: (2024)
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
by: Huang, Yiheng, et al.
Published: (2025)
by: Huang, Yiheng, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
by: Mo, Wentao, et al.
Published: (2025)
by: Mo, Wentao, et al.
Published: (2025)
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
by: Chen, Liangyu, et al.
Published: (2024)
by: Chen, Liangyu, et al.
Published: (2024)
SMTPD: A New Benchmark for Temporal Prediction of Social Media Popularity
by: Xu, Yijie, et al.
Published: (2025)
by: Xu, Yijie, et al.
Published: (2025)
Can We Hear from Events? Generating Speech from Event Camera
by: Fang, Jingping, et al.
Published: (2026)
by: Fang, Jingping, et al.
Published: (2026)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
by: He, Zheqi, et al.
Published: (2024)
by: He, Zheqi, et al.
Published: (2024)
Multimodal LLM-based Query Paraphrasing for Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
by: Gao, Shibo, et al.
Published: (2025)
by: Gao, Shibo, et al.
Published: (2025)
Mitigating Multimodal Inconsistency via Cognitive Dual-Pathway Reasoning for Intent Recognition
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
XY-Cut++: Advanced Layout Ordering via Hierarchical Mask Mechanism on a Novel Benchmark
by: Liu, Shuai, et al.
Published: (2025)
by: Liu, Shuai, et al.
Published: (2025)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
by: Lu, Jiacheng, et al.
Published: (2025)
by: Lu, Jiacheng, et al.
Published: (2025)
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
by: Hao, Fangyu, et al.
Published: (2026)
by: Hao, Fangyu, et al.
Published: (2026)
Universal Organizer of SAM for Unsupervised Semantic Segmentation
by: Li, Tingting, et al.
Published: (2024)
by: Li, Tingting, et al.
Published: (2024)
FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
by: He, Xusheng, et al.
Published: (2025)
by: He, Xusheng, et al.
Published: (2025)
SSTFB: Leveraging self-supervised pretext learning and temporal self-attention with feature branching for real-time video polyp segmentation
by: Xu, Ziang, et al.
Published: (2024)
by: Xu, Ziang, et al.
Published: (2024)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
by: Nguyen, Manh Duong, et al.
Published: (2024)
by: Nguyen, Manh Duong, et al.
Published: (2024)
LoginMEA: Local-to-Global Interaction Network for Multi-modal Entity Alignment
by: Su, Taoyu, et al.
Published: (2024)
by: Su, Taoyu, et al.
Published: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
by: Ji, Shengpeng, et al.
Published: (2025)
by: Ji, Shengpeng, et al.
Published: (2025)
Deep joint source-channel coding for wireless point cloud transmission
by: Zhang, Cixiao, et al.
Published: (2024)
by: Zhang, Cixiao, et al.
Published: (2024)
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
by: Fang, Pengcheng, et al.
Published: (2026)
by: Fang, Pengcheng, et al.
Published: (2026)
Similar Items
-
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
by: Sheng, Zhichao, et al.
Published: (2025) -
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
by: Peng, Yuezhang, et al.
Published: (2025) -
Multimodal Classification and Out-of-distribution Detection for Multimodal Intent Understanding
by: Zhang, Hanlei, et al.
Published: (2024) -
Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2026) -
Dual Dynamic Threshold Adjustment Strategy for Deep Metric Learning
by: Jiang, Xiruo, et al.
Published: (2024)