Wavelet GPT: Wavelet Inspired Large Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Verma, Prateek |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Whisper-GPT: A Hybrid Representation Audio Large Language Model
von: Verma, Prateek
Veröffentlicht: (2024)
von: Verma, Prateek
Veröffentlicht: (2024)
Towards Signal Processing In Large Language Models
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
Adaptive Large Language Models By Layerwise Attention Shortcuts
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
von: Verma, Prateek, et al.
Veröffentlicht: (2024)
Large Language Models Implicitly Learn to See and Hear Just By Reading
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)
von: Li, Xu, et al.
Veröffentlicht: (2024)
Thinking While Listening: Simple Test Time Scaling For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
von: Verma, Prateek, et al.
Veröffentlicht: (2025)
VoxKnesset: A Large-Scale Longitudinal Hebrew Speech Dataset for Aging Speaker Modeling
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
von: Marmor, Yanir, et al.
Veröffentlicht: (2026)
Diverse Audio Embeddings -- Bringing Features Back Outperforms CLAP!
von: Verma, Prateek
Veröffentlicht: (2023)
von: Verma, Prateek
Veröffentlicht: (2023)
A Language Model With Million Context Length For Raw Audio
von: Verma, Prateek
Veröffentlicht: (2022)
von: Verma, Prateek
Veröffentlicht: (2022)
Is Attention always needed? A Case Study on Language Identification from Speech
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
von: Mandal, Atanu, et al.
Veröffentlicht: (2021)
An Explainable Proxy Model for Multiabel Audio Segmentation
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
von: Mariotte, Théo, et al.
Veröffentlicht: (2024)
Model as Loss: A Self-Consistent Training Paradigm
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
von: Phaye, Saisamarth Rajesh, et al.
Veröffentlicht: (2025)
Spoken Language Intelligence of Large Language Models for Language Learning
von: Peng, Linkai, et al.
Veröffentlicht: (2023)
von: Peng, Linkai, et al.
Veröffentlicht: (2023)
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
von: Chung, Yoonjin, et al.
Veröffentlicht: (2024)
TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
von: Huang, Zhaolan, et al.
Veröffentlicht: (2024)
Large Language Models' Internal Perception of Symbolic Music
von: Shin, Andrew, et al.
Veröffentlicht: (2025)
von: Shin, Andrew, et al.
Veröffentlicht: (2025)
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
von: Sakai, Yusuke, et al.
Veröffentlicht: (2024)
Audio Transformers
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
von: Verma, Prateek, et al.
Veröffentlicht: (2021)
Content Adaptive Front End For Audio Classification
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
von: Verma, Prateek, et al.
Veröffentlicht: (2023)
RIFT: Entropy-Optimised Fractional Wavelet Constellations for Ideal Time-Frequency Estimation
von: Cozens, James M., et al.
Veröffentlicht: (2025)
von: Cozens, James M., et al.
Veröffentlicht: (2025)
Wavelet-Based Time-Frequency Fingerprinting for Feature Extraction of Traditional Irish Music
von: Shore, Noah
Veröffentlicht: (2025)
von: Shore, Noah
Veröffentlicht: (2025)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2024)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
von: Guo, Z., et al.
Veröffentlicht: (2022)
von: Guo, Z., et al.
Veröffentlicht: (2022)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
von: Wang, Jinhan, et al.
Veröffentlicht: (2024)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
von: Parker, Julian D, et al.
Veröffentlicht: (2024)
LMAC-TD: Producing Time Domain Explanations for Audio Classifiers
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
von: Mancini, Eleonora, et al.
Veröffentlicht: (2024)
Gull: A Generative Multifunctional Audio Codec
von: Luo, Yi, et al.
Veröffentlicht: (2024)
von: Luo, Yi, et al.
Veröffentlicht: (2024)
Expressive Acoustic Guitar Sound Synthesis with an Instrument-Specific Input Representation and Diffusion Outpainting
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
von: Kim, Hounsu, et al.
Veröffentlicht: (2024)
ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
von: Kim, Daewoong, et al.
Veröffentlicht: (2024)
BUET Multi-disease Heart Sound Dataset: A Comprehensive Auscultation Dataset for Developing Computer-Aided Diagnostic Systems
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
von: Ali, Shams Nafisa, et al.
Veröffentlicht: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
von: Lee, Sang-Hoon, et al.
Veröffentlicht: (2024)
Real-time Timbre Remapping with Differentiable DSP
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
von: Shier, Jordie, et al.
Veröffentlicht: (2024)
Deep Active Speech Cancellation with Mamba-Masking Network
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
von: Mishaly, Yehuda, et al.
Veröffentlicht: (2025)
Lightweight Self-Supervised Detection of Fundamental Frequency and Accurate Probability of Voicing in Monophonic Music
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
von: Bitra, Venkat Suprabath, et al.
Veröffentlicht: (2026)
Design Of Rubble Analyzer Probe Using ML For Earthquake
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
von: Sebastian, Abhishek, et al.
Veröffentlicht: (2023)
ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
von: Ackva, Valentin, et al.
Veröffentlicht: (2025)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Whisper-GPT: A Hybrid Representation Audio Large Language Model
von: Verma, Prateek
Veröffentlicht: (2024) -
Towards Signal Processing In Large Language Models
von: Verma, Prateek, et al.
Veröffentlicht: (2024) -
Adaptive Large Language Models By Layerwise Attention Shortcuts
von: Verma, Prateek, et al.
Veröffentlicht: (2024) -
Large Language Models Implicitly Learn to See and Hear Just By Reading
von: Verma, Prateek, et al.
Veröffentlicht: (2025) -
MaskSR: Masked Language Model for Full-band Speech Restoration
von: Li, Xu, et al.
Veröffentlicht: (2024)