IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Changheon, Sim, Yuseop, Jung, Hoin, Lee, Jiho, Lee, Hojun, Kang, Yun Seok, Woo, Sucheol, Kim, Garam, Park, Hyung Wook, Jun, Martin Byung-Guk |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
di: Han, Changheon, et al.
Pubblicazione: (2025)
di: Han, Changheon, et al.
Pubblicazione: (2025)
Inverse Nonlinearity Compensation of Hyperelastic Deformation in Dielectric Elastomer for Acoustic Actuation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
Track Role Prediction of Single-Instrumental Sequences
di: Han, Changheon, et al.
Pubblicazione: (2024)
di: Han, Changheon, et al.
Pubblicazione: (2024)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
di: Yeon, Inmo, et al.
Pubblicazione: (2023)
di: Yeon, Inmo, et al.
Pubblicazione: (2023)
DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis
di: Lee, Dongheon, et al.
Pubblicazione: (2025)
di: Lee, Dongheon, et al.
Pubblicazione: (2025)
Inter-channel Conv-TasNet for multichannel speech enhancement
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
di: Lee, Dongheon, et al.
Pubblicazione: (2021)
A Novel Deep Learning Framework for Efficient Multichannel Acoustic Feedback Control
di: Wu, Yuan-Kuei, et al.
Pubblicazione: (2025)
di: Wu, Yuan-Kuei, et al.
Pubblicazione: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
di: Chung, Soo-Whan, et al.
Pubblicazione: (2025)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
di: Kong, Jungil, et al.
Pubblicazione: (2023)
di: Kong, Jungil, et al.
Pubblicazione: (2023)
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
di: Choi, Dayun, et al.
Pubblicazione: (2025)
di: Choi, Dayun, et al.
Pubblicazione: (2025)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2024)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2024)
SAM: A Mamba-2 State-Space Audio-Language Model
di: Lee, Taehan, et al.
Pubblicazione: (2025)
di: Lee, Taehan, et al.
Pubblicazione: (2025)
WoW-Bench: Evaluating Fine-Grained Acoustic Perception in Audio-Language Models via Marine Mammal Vocalizations
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2025)
CoPlay: Audio-agnostic Cognitive Scaling for Acoustic Sensing
di: Li, Yin, et al.
Pubblicazione: (2024)
di: Li, Yin, et al.
Pubblicazione: (2024)
3D Room Geometry Inference from Multichannel Room Impulse Response using Deep Neural Network
di: Yeon, Inmo, et al.
Pubblicazione: (2024)
di: Yeon, Inmo, et al.
Pubblicazione: (2024)
DISPATCH: Distilling Selective Patches for Speech Enhancement
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
di: Kim, Dohwan, et al.
Pubblicazione: (2025)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
Variational Bayesian Adaptive Learning of Deep Latent Variables for Acoustic Knowledge Transfer
di: Hu, Hu, et al.
Pubblicazione: (2025)
di: Hu, Hu, et al.
Pubblicazione: (2025)
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
di: Choi, Joon-Seung, et al.
Pubblicazione: (2025)
di: Choi, Joon-Seung, et al.
Pubblicazione: (2025)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
di: Cha, Jun-Hyeok, et al.
Pubblicazione: (2025)
di: Cha, Jun-Hyeok, et al.
Pubblicazione: (2025)
Improving Acoustic Scene Classification in Low-Resource Conditions
di: Chen, Zhi, et al.
Pubblicazione: (2024)
di: Chen, Zhi, et al.
Pubblicazione: (2024)
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
di: Huang, Kuan-Po, et al.
Pubblicazione: (2025)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2025)
Adversarial Multi-Task Learning for Disentangling Timbre and Pitch in Singing Voice Synthesis
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
di: Kim, Tae-Woo, et al.
Pubblicazione: (2022)
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
di: Jin, Hojun, et al.
Pubblicazione: (2025)
di: Jin, Hojun, et al.
Pubblicazione: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
Differentiable Acoustic Radiance Transfer
di: Lee, Sungho, et al.
Pubblicazione: (2025)
di: Lee, Sungho, et al.
Pubblicazione: (2025)
Adversarial Domain Adaptation for Metal Cutting Sound Detection: Leveraging Abundant Lab Data for Scarce Industry Data
di: Mostafiz, Mir Imtiaz, et al.
Pubblicazione: (2024)
di: Mostafiz, Mir Imtiaz, et al.
Pubblicazione: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
di: Xiao, Yang, et al.
Pubblicazione: (2026)
di: Xiao, Yang, et al.
Pubblicazione: (2026)
Evaluation of Virtual Acoustic Environments with Different Acoustic Level of Detail
di: Fichna, Stefan, et al.
Pubblicazione: (2023)
di: Fichna, Stefan, et al.
Pubblicazione: (2023)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
di: Heo, Hyun-Jun, et al.
Pubblicazione: (2023)
di: Heo, Hyun-Jun, et al.
Pubblicazione: (2023)
Speech-Aware Neural Diarization with Encoder-Decoder Attractor Guided by Attention Constraints
di: Lee, PeiYing, et al.
Pubblicazione: (2024)
di: Lee, PeiYing, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
di: Han, Changheon, et al.
Pubblicazione: (2025) -
Inverse Nonlinearity Compensation of Hyperelastic Deformation in Dielectric Elastomer for Acoustic Actuation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024) -
Track Role Prediction of Single-Instrumental Sequences
di: Han, Changheon, et al.
Pubblicazione: (2024) -
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
di: Lee, Dongheon, et al.
Pubblicazione: (2024) -
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)