W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Babaei, Hossein, White, Mel, Baraniuk, Richard G. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WaLRUS: Wavelets for Long-range Representation Using SSMs
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
SaFARi: State-Space Models for Frame-Agnostic Representation
by: Babaei, Hossein, et al.
Published: (2025)
by: Babaei, Hossein, et al.
Published: (2025)
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
by: Premananth, Gowtham, et al.
Published: (2025)
by: Premananth, Gowtham, et al.
Published: (2025)
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
by: Ono, Nobutaka
Published: (2026)
by: Ono, Nobutaka
Published: (2026)
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
ToS: A Team of Specialists ensemble framework for Stereo Sound Event Localization and Detection with distance estimation in Video
by: Berghi, Davide, et al.
Published: (2026)
by: Berghi, Davide, et al.
Published: (2026)
Livestock feeding behaviour: A review on automated systems for ruminant monitoring
by: Chelotti, José, et al.
Published: (2023)
by: Chelotti, José, et al.
Published: (2023)
LiVeAction: a Lightweight, Versatile, and Asymmetric Neural Codec Design for Real-time Operation
by: Jacobellis, Dan, et al.
Published: (2026)
by: Jacobellis, Dan, et al.
Published: (2026)
Multimodal sensor fusion for real-time location-dependent defect detection in laser-directed energy deposition
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
Integrating Spatial and Semantic Embeddings for Stereo Sound Event Localization in Videos
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
by: Wang, Zanxu, et al.
Published: (2025)
by: Wang, Zanxu, et al.
Published: (2025)
RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior
by: Lee, Ching-Hua, et al.
Published: (2025)
by: Lee, Ching-Hua, et al.
Published: (2025)
Leveraging Unlabeled Audio-Visual Data in Speech Emotion Recognition using Knowledge Distillation
by: Pendyala, Varsha, et al.
Published: (2025)
by: Pendyala, Varsha, et al.
Published: (2025)
Learned Compression for Compressed Learning
by: Jacobellis, Dan, et al.
Published: (2024)
by: Jacobellis, Dan, et al.
Published: (2024)
Efficient Test-Time Adaptation through Latent Subspace Coefficients Search
by: Luo, Xinyu, et al.
Published: (2025)
by: Luo, Xinyu, et al.
Published: (2025)
TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization
by: Malard, Hugo, et al.
Published: (2024)
by: Malard, Hugo, et al.
Published: (2024)
A Smart-Glasses for Emergency Medical Services via Multimodal Multitask Learning
by: Jin, Liuyi, et al.
Published: (2025)
by: Jin, Liuyi, et al.
Published: (2025)
SoundSil-DS: Deep Denoising and Segmentation of Sound-field Images with Silhouettes
by: Tanigawa, Risako, et al.
Published: (2024)
by: Tanigawa, Risako, et al.
Published: (2024)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
Leveraging Reverberation and Visual Depth Cues for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection
by: Islam, Md Shofiqul, et al.
Published: (2025)
by: Islam, Md Shofiqul, et al.
Published: (2025)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
DLIOS: An LLM-Augmented Real-Time Multi-Modal Interactive Enhancement Overlay System for Douyin Live Streaming
by: Wen, Shuide, et al.
Published: (2026)
by: Wen, Shuide, et al.
Published: (2026)
Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training
by: Tao, Ruijie, et al.
Published: (2024)
by: Tao, Ruijie, et al.
Published: (2024)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
by: Eliav, Amit, et al.
Published: (2024)
by: Eliav, Amit, et al.
Published: (2024)
BUT System Description for CHiME-9 MCoRec Challenge
by: Klement, Dominik, et al.
Published: (2026)
by: Klement, Dominik, et al.
Published: (2026)
Listening for "You": Enhancing Speech Image Retrieval via Target Speaker Extraction
by: Yang, Wenhao, et al.
Published: (2025)
by: Yang, Wenhao, et al.
Published: (2025)
Bounds on Agreement between Subjective and Objective Measurements
by: Pieper, Jaden, et al.
Published: (2026)
by: Pieper, Jaden, et al.
Published: (2026)
Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
by: Park, Jay, et al.
Published: (2025)
by: Park, Jay, et al.
Published: (2025)
SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet
by: Zhong, Zhi, et al.
Published: (2025)
by: Zhong, Zhi, et al.
Published: (2025)
Multimodal Machine Learning Can Predict Videoconference Fluidity and Enjoyment
by: Chang, Andrew, et al.
Published: (2025)
by: Chang, Andrew, et al.
Published: (2025)
TITAN: Bringing The Deep Image Prior to Implicit Representations
by: Luzi, Lorenzo, et al.
Published: (2022)
by: Luzi, Lorenzo, et al.
Published: (2022)
Joint Source-Environment Adaptation of Data-Driven Underwater Acoustic Source Ranging Based on Model Uncertainty
by: Kari, Dariush, et al.
Published: (2025)
by: Kari, Dariush, et al.
Published: (2025)
KunquDB: An Attempt for Speaker Verification in the Chinese Opera Scenario
by: Zhou, Huali, et al.
Published: (2024)
by: Zhou, Huali, et al.
Published: (2024)
Attentive AV-FusionNet: Audio-Visual Quality Prediction with Hybrid Attention
by: Salaj, Ina, et al.
Published: (2025)
by: Salaj, Ina, et al.
Published: (2025)
Improvement Of Audiovisual Quality Estimation Using A Nonlinear Autoregressive Exogenous Neural Network And Bitstream Parameters
by: Kossi, Koffi, et al.
Published: (2024)
by: Kossi, Koffi, et al.
Published: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
Efficient Face Detection with Audio-Based Region Proposals for Human-Robot Interactions
by: Aris, William, et al.
Published: (2023)
by: Aris, William, et al.
Published: (2023)
Joint Source-Environment Adaptation for Deep Learning-Based Underwater Acoustic Source Ranging
by: Kari, Dariush, et al.
Published: (2025)
by: Kari, Dariush, et al.
Published: (2025)
Similar Items
-
WaLRUS: Wavelets for Long-range Representation Using SSMs
by: Babaei, Hossein, et al.
Published: (2025) -
SaFARi: State-Space Models for Frame-Agnostic Representation
by: Babaei, Hossein, et al.
Published: (2025) -
Multimodal Biomarkers for Schizophrenia: Towards Individual Symptom Severity Estimation
by: Premananth, Gowtham, et al.
Published: (2025) -
Fast Swap-Based Element Selection for Multiplication-Free Dimension Reduction
by: Ono, Nobutaka
Published: (2026) -
Spatial and Semantic Embedding Integration for Stereo Sound Event Localization and Detection in Regular Videos
by: Berghi, Davide, et al.
Published: (2025)