MiMo-Audio: Audio Language Models are Few-Shot Learners
Fuente:
arXiv
Saved in:
Similar Items
MiMo-VL Technical Report
by: Core Team, et al.
Published: (2025)
by: Core Team, et al.
Published: (2025)
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
by: Xiaomi, LLM-Core, et al.
Published: (2025)
by: Xiaomi, LLM-Core, et al.
Published: (2025)
MiMo-V2-Flash Technical Report
by: Core Team, et al.
Published: (2026)
by: Core Team, et al.
Published: (2026)
Audio ControlNet for Fine-Grained Audio Generation and Editing
by: Zhu, Haina, et al.
Published: (2026)
by: Zhu, Haina, et al.
Published: (2026)
Dual‐Targeted Novel Temozolomide Nanocapsules Encapsulating siPKM2 Inhibit Aerobic Glycolysis to Sensitize Glioblastoma to Chemotherapy
by: Yongkang Zhang, et al.
Published: (2024)
by: Yongkang Zhang, et al.
Published: (2024)
RA-Rec: An Efficient ID Representation Alignment Framework for LLM-based Recommendation
by: Yu, Xiaohan, et al.
Published: (2024)
by: Yu, Xiaohan, et al.
Published: (2024)
Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation
by: Yu, Junjie, et al.
Published: (2026)
by: Yu, Junjie, et al.
Published: (2026)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
by: Wang, Jialing, et al.
Published: (2026)
by: Wang, Jialing, et al.
Published: (2026)
Regulating Crystalline Phase/Plane of Polymer Electrolyte for Rapid Lithium Ion Transfer
by: Su Wang, et al.
Published: (2024)
by: Su Wang, et al.
Published: (2024)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
by: Lin, Jingru, et al.
Published: (2026)
by: Lin, Jingru, et al.
Published: (2026)
Iron regulatory protein from the hard tick Haemaphysalis longicornis : characterization, function and assessment as a protective antigen
by: Duo Wang, et al.
Published: (2024)
by: Duo Wang, et al.
Published: (2024)
Front Cover Image
by: Duo Wang, et al.
Published: (2024)
by: Duo Wang, et al.
Published: (2024)
Review on polymer electrolytes for lithium‐sulfurized polyacrylonitrile batteries
by: Yan Zhang, et al.
Published: (2024)
by: Yan Zhang, et al.
Published: (2024)
Break the ID-Language Barrier: An Adaption Framework for LLM-based Sequential Recommendation
by: Yu, Xiaohan, et al.
Published: (2024)
by: Yu, Xiaohan, et al.
Published: (2024)
AirShot: Efficient Few-Shot Detection for Autonomous Exploration
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
Robust Audio-Visual Segmentation via Audio-Guided Visual Convergent Alignment
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
SMAC-Hard: Enabling Mixed Opponent Strategy Script and Self-play on SMAC
by: Deng, Yue, et al.
Published: (2024)
by: Deng, Yue, et al.
Published: (2024)
Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
by: Chen, Dexia, et al.
Published: (2025)
by: Chen, Dexia, et al.
Published: (2025)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
Integrating Large Language Models into Text Animation: An Intelligent Editing System with Inline and Chat Interaction
by: Zhang, Bao, et al.
Published: (2025)
by: Zhang, Bao, et al.
Published: (2025)
Towards Multimodal Query-Based Spatial Audio Source Extraction
by: Yu, Chenxin, et al.
Published: (2025)
by: Yu, Chenxin, et al.
Published: (2025)
Ultrasonic vibration‐assisted joining of multiform metallic glasses in varied environments
by: Lu‐Yao Li, et al.
Published: (2025)
by: Lu‐Yao Li, et al.
Published: (2025)
MM-STFlowNet: A Transportation Hub-Oriented Multi-Mode Passenger Flow Prediction Method via Spatial-Temporal Dynamic Graph Modeling
by: Zhang, Ronghui, et al.
Published: (2025)
by: Zhang, Ronghui, et al.
Published: (2025)
UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
by: Jiang, Yidi, et al.
Published: (2025)
by: Jiang, Yidi, et al.
Published: (2025)
The HCMV‐encoded miR‐UL36‐3p promotes angiogenesis of endothelial cells by downregulating FOXO3
by: Chen Wang, et al.
Published: (2026)
by: Chen Wang, et al.
Published: (2026)
VA-CDH: A Variance-Aware Method to Optimize Latency for Caching with Delayed Hits
by: Jiang, Bowen, et al.
Published: (2025)
by: Jiang, Bowen, et al.
Published: (2025)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
by: Li, Xiquan, et al.
Published: (2026)
by: Li, Xiquan, et al.
Published: (2026)
Flavor Nernst effects in quantum paramagnets
by: Lu, Bowen, et al.
Published: (2024)
by: Lu, Bowen, et al.
Published: (2024)
nnSAM: Plug-and-play Segment Anything Model Improves nnUNet Performance
by: Li, Yunxiang, et al.
Published: (2023)
by: Li, Yunxiang, et al.
Published: (2023)
Plug‐and‐play segment anything model improves nnUNet performance
by: Yunxiang Li, et al.
Published: (2024)
by: Yunxiang Li, et al.
Published: (2024)
FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
by: Li, Li, et al.
Published: (2021)
by: Li, Li, et al.
Published: (2021)
CSAFL: A Clustered Semi-Asynchronous Federated Learning Framework
by: Zhang, Yu, et al.
Published: (2021)
by: Zhang, Yu, et al.
Published: (2021)
Vertically Concentrated Quantum Wells Enabling Highly Efficient Deep‐Blue Perovskite Light‐Emitting Diodes
by: Yu Xia, et al.
Published: (2024)
by: Yu Xia, et al.
Published: (2024)
Vertically Concentrated Quantum Wells Enabling Highly Efficient Deep‐Blue Perovskite Light‐Emitting Diodes
by: Yu Xia, et al.
Published: (2024)
by: Yu Xia, et al.
Published: (2024)
A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
by: Wang, Pingjie, et al.
Published: (2024)
by: Wang, Pingjie, et al.
Published: (2024)
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection
by: Chen, Meng, et al.
Published: (2026)
by: Chen, Meng, et al.
Published: (2026)
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
by: Wang, Le, et al.
Published: (2025)
by: Wang, Le, et al.
Published: (2025)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
by: Song, Zirui, et al.
Published: (2025)
by: Song, Zirui, et al.
Published: (2025)
Similar Items
-
MiMo-VL Technical Report
by: Core Team, et al.
Published: (2025) -
MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
by: Xiaomi, LLM-Core, et al.
Published: (2025) -
MiMo-V2-Flash Technical Report
by: Core Team, et al.
Published: (2026) -
Audio ControlNet for Fine-Grained Audio Generation and Editing
by: Zhu, Haina, et al.
Published: (2026) -
Dual‐Targeted Novel Temozolomide Nanocapsules Encapsulating siPKM2 Inhibit Aerobic Glycolysis to Sensitize Glioblastoma to Chemotherapy
by: Yongkang Zhang, et al.
Published: (2024)