Saved in:
| Main Authors: | Guo, Wenxiang, Pan, Changhao, Zhu, Zhiyuan, Hu, Xintong, Zhang, Yu, Tang, Li, Yang, Rui, Wang, Han, Zhang, Zongbao, Wang, Yuhan, Chen, Yixuan, Xu, Hankun, Xu, Ke, Fan, Pengfei, Chen, Zhetao, Yu, Yanhao, Huang, Qiange, Wu, Fei, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.10396 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
by: Guo, Wenxiang, et al.
Published: (2025)
by: Guo, Wenxiang, et al.
Published: (2025)
ASAudio: A Survey of Advanced Spatial Audio Research
by: Zhu, Zhiyuan, et al.
Published: (2025)
by: Zhu, Zhiyuan, et al.
Published: (2025)
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
by: Hu, Xintong, et al.
Published: (2025)
by: Hu, Xintong, et al.
Published: (2025)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
by: Lei, Ke, et al.
Published: (2026)
by: Lei, Ke, et al.
Published: (2026)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Refining CLIP's Spatial Awareness: A Visual-Centric Perspective
by: Qiu, Congpei, et al.
Published: (2025)
by: Qiu, Congpei, et al.
Published: (2025)
AcOrch: Accelerating Sampling-based GNN Training under CPU-NPU Heterogeneous Environments
by: Chen, Kefu, et al.
Published: (2026)
by: Chen, Kefu, et al.
Published: (2026)
Towards Autonomous Graph Data Analytics with Analytics-Augmented Generation
by: Wang, Qiange, et al.
Published: (2026)
by: Wang, Qiange, et al.
Published: (2026)
Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
by: Pan, Changhao, et al.
Published: (2026)
by: Pan, Changhao, et al.
Published: (2026)
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
by: Guo, Yiwei, et al.
Published: (2025)
by: Guo, Yiwei, et al.
Published: (2025)
EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation
by: Fu, Zhenbo, et al.
Published: (2026)
by: Fu, Zhenbo, et al.
Published: (2026)
Characterization of shale pore heterogeneity and its controlling factors: A case study of the Longmaxi Formation in Western Hubei, China
by: Zongbao Diao, et al.
Published: (2024)
by: Zongbao Diao, et al.
Published: (2024)
Comprehensive Evaluation of GNN Training Systems: A Data Management Perspective
by: Yuan, Hao, et al.
Published: (2023)
by: Yuan, Hao, et al.
Published: (2023)
SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Revealing the Immune and Inflammatory Mechanisms of Electroacupuncture in Male IBS Rats Through Multi‐Omics Analysis
by: Lijun Wang, et al.
Published: (2025)
by: Lijun Wang, et al.
Published: (2025)
ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality
by: Luo, Yu-Xiang, et al.
Published: (2025)
by: Luo, Yu-Xiang, et al.
Published: (2025)
R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation
by: Xu, Xiuwei, et al.
Published: (2025)
by: Xu, Xiuwei, et al.
Published: (2025)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
Exploration of Two‐Dimensional Conductance Along LaAlO 3 /SrTiO 3 Interface for Pressure Sensing
by: Yiwen Shi, et al.
Published: (2025)
by: Yiwen Shi, et al.
Published: (2025)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
by: Li, Bohan, et al.
Published: (2025)
by: Li, Bohan, et al.
Published: (2025)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
by: Chen, Wenxi, et al.
Published: (2024)
by: Chen, Wenxi, et al.
Published: (2024)
What is "Spatial" about Spatial Computing?
by: Wang, Yibo, et al.
Published: (2025)
by: Wang, Yibo, et al.
Published: (2025)
Spatial analysis of tails of air pollution PDFs in Europe
by: He, Hankun, et al.
Published: (2024)
by: He, Hankun, et al.
Published: (2024)
UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
by: Qiu, Congpei, et al.
Published: (2026)
by: Qiu, Congpei, et al.
Published: (2026)
Versatile Framework for Song Generation with Prompt-based Control
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
by: Wang, Hankun, et al.
Published: (2024)
by: Wang, Hankun, et al.
Published: (2024)
Towards Multimodal Query-Based Spatial Audio Source Extraction
by: Yu, Chenxin, et al.
Published: (2025)
by: Yu, Chenxin, et al.
Published: (2025)
Spatially resolved analysis of Stellar Populations in NGC 2992: Impact of AGN feedback
by: Xu, Xiaoyu, et al.
Published: (2024)
by: Xu, Xiaoyu, et al.
Published: (2024)
A Survey on Speech Large Language Models for Understanding
by: Peng, Jing, et al.
Published: (2024)
by: Peng, Jing, et al.
Published: (2024)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
by: Zhang, Situo, et al.
Published: (2026)
by: Zhang, Situo, et al.
Published: (2026)
RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
by: Yang, Bing, et al.
Published: (2024)
by: Yang, Bing, et al.
Published: (2024)
AdaEAGLE: Optimizing Speculative Decoding via Explicit Modeling of Adaptive Draft Structures
by: Zhang, Situo, et al.
Published: (2024)
by: Zhang, Situo, et al.
Published: (2024)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
Can Large Language Models Understand Spatial Audio?
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
GNSS Real-Time Kinematic Positioning
by: Li, Bofeng, et al.
Published: (2025)
by: Li, Bofeng, et al.
Published: (2025)
Efficient exact sequential lifting algorithm for binary knapsack set
by: Wang, Xintong, et al.
Published: (2026)
by: Wang, Xintong, et al.
Published: (2026)
Janus Tough Adhesive Ionotronics for High‐Fidelity Biomechanical and Electrophysiological Signal Recording
by: Xiao‐Xue Wang, et al.
Published: (2026)
by: Xiao‐Xue Wang, et al.
Published: (2026)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
Similar Items
-
STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation
by: Guo, Wenxiang, et al.
Published: (2025) -
ASAudio: A Survey of Advanced Spatial Audio Research
by: Zhu, Zhiyuan, et al.
Published: (2025) -
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
by: Hu, Xintong, et al.
Published: (2025) -
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
by: Lei, Ke, et al.
Published: (2026) -
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
by: Zhang, Yu, et al.
Published: (2025)