WDMIR: Wavelet-Driven Multimodal Intent Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Weiyin, Zhang, Kai, Zhang, Yanghai, Liu, Qi, Sun, Xinjie, Lu, Junyu, Zhu, Linbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Full-reference Point Cloud Quality Assessment Using Spectral Graph Wavelets
by: Watanabe, Ryosuke, et al.
Published: (2024)
by: Watanabe, Ryosuke, et al.
Published: (2024)
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
Communicate Less, Synthesize the Rest: Latency-aware Intent-based Generative Semantic Multicasting with Diffusion Models
by: Liu, Xinkai, et al.
Published: (2024)
by: Liu, Xinkai, et al.
Published: (2024)
Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications
by: Qiao, Li, et al.
Published: (2025)
by: Qiao, Li, et al.
Published: (2025)
Enhanced Radar Perception via Multi-Task Learning: Towards Refined Data for Sensor Fusion Applications
by: Sun, Huawei, et al.
Published: (2024)
by: Sun, Huawei, et al.
Published: (2024)
WiMANS: A Benchmark Dataset for WiFi-based Multi-user Activity Sensing
by: Huang, Shuokang, et al.
Published: (2024)
by: Huang, Shuokang, et al.
Published: (2024)
Frequency-Spatial Interaction Driven Network for Low-Light Image Enhancement
by: Tao, Yunhong, et al.
Published: (2025)
by: Tao, Yunhong, et al.
Published: (2025)
Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature Fusion
by: Lim, DongHoon, et al.
Published: (2025)
by: Lim, DongHoon, et al.
Published: (2025)
GenHPE: Generative Counterfactuals for 3D Human Pose Estimation with Radio Frequency Signals
by: Huang, Shuokang, et al.
Published: (2025)
by: Huang, Shuokang, et al.
Published: (2025)
Latency-Aware Generative Semantic Communications with Pre-Trained Diffusion Models
by: Qiao, Li, et al.
Published: (2024)
by: Qiao, Li, et al.
Published: (2024)
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
by: Yang, Dejie, et al.
Published: (2025)
by: Yang, Dejie, et al.
Published: (2025)
Shared-kernel Wavelet Neural Networks for Poisson Image Reconstruction
by: Gong, Yuanhao, et al.
Published: (2026)
by: Gong, Yuanhao, et al.
Published: (2026)
Tackle CSM in JPEG Steganalysis with Data Adaptation
by: Abecidan, Rony, et al.
Published: (2026)
by: Abecidan, Rony, et al.
Published: (2026)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
by: Zhu, Haodong, et al.
Published: (2025)
by: Zhu, Haodong, et al.
Published: (2025)
DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion Recognition
by: Jiang, Peiyuan, et al.
Published: (2025)
by: Jiang, Peiyuan, et al.
Published: (2025)
R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
Wavelet-Decoupling Contrastive Enhancement Network for Fine-Grained Skeleton-Based Action Recognition
by: Chang, Haochen, et al.
Published: (2024)
by: Chang, Haochen, et al.
Published: (2024)
Task-Oriented Communication for Edge Video Analytics
by: Shao, Jiawei, et al.
Published: (2022)
by: Shao, Jiawei, et al.
Published: (2022)
AsyReC: A Multimodal Graph-based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification
by: Tang, Wang, et al.
Published: (2025)
by: Tang, Wang, et al.
Published: (2025)
SZTU-CMU at MER2024: Improving Emotion-LLaMA with Conv-Attention for Multimodal Emotion Recognition
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
Digit Recognition using Multimodal Spiking Neural Networks
by: Bjorndahl, William, et al.
Published: (2024)
by: Bjorndahl, William, et al.
Published: (2024)
Multi-scale Bottleneck Transformer for Weakly Supervised Multimodal Violence Detection
by: Sun, Shengyang, et al.
Published: (2024)
by: Sun, Shengyang, et al.
Published: (2024)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
by: Koo, Inyong, et al.
Published: (2026)
by: Koo, Inyong, et al.
Published: (2026)
Towards Robust and Realible Multimodal Misinformation Recognition with Incomplete Modality
by: Zhou, Hengyang, et al.
Published: (2025)
by: Zhou, Hengyang, et al.
Published: (2025)
MCIHN: A Hybrid Network Model Based on Multi-path Cross-modal Interaction for Multimodal Emotion Recognition
by: Zhang, Haoyang, et al.
Published: (2025)
by: Zhang, Haoyang, et al.
Published: (2025)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
by: Zhang, Liyun
Published: (2024)
by: Zhang, Liyun
Published: (2024)
Exploration of Learned Lifting-Based Transform Structures for Fully Scalable and Accessible Wavelet-Like Image Compression
by: Li, Xinyue, et al.
Published: (2024)
by: Li, Xinyue, et al.
Published: (2024)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
by: Zhao, Fuzheng, et al.
Published: (2024)
by: Zhao, Fuzheng, et al.
Published: (2024)
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
by: Skoularikis, Anastasios, et al.
Published: (2025)
by: Skoularikis, Anastasios, et al.
Published: (2025)
Releasing the Parameter Latency of Neural Representation for High-Efficiency Video Compression
by: Zhang, Gai, et al.
Published: (2024)
by: Zhang, Gai, et al.
Published: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
by: Chi, Xiaowei, et al.
Published: (2024)
by: Chi, Xiaowei, et al.
Published: (2024)
Proxy-Tuning: Tailoring Multimodal Autoregressive Models for Subject-Driven Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
Multifaceted Evaluation of Audio-Visual Capability for MLLMs: Effectiveness, Efficiency, Generalizability and Robustness
by: Zhao, Yusheng, et al.
Published: (2025)
by: Zhao, Yusheng, et al.
Published: (2025)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
by: Yang, Jasmine, et al.
Published: (2026)
by: Yang, Jasmine, et al.
Published: (2026)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
by: Zhang, Xinjie, et al.
Published: (2024)
by: Zhang, Xinjie, et al.
Published: (2024)
Similar Items
-
Full-reference Point Cloud Quality Assessment Using Spectral Graph Wavelets
by: Watanabe, Ryosuke, et al.
Published: (2024) -
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025) -
Communicate Less, Synthesize the Rest: Latency-aware Intent-based Generative Semantic Multicasting with Diffusion Models
by: Liu, Xinkai, et al.
Published: (2024) -
Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications
by: Qiao, Li, et al.
Published: (2025) -
Enhanced Radar Perception via Multi-Task Learning: Towards Refined Data for Sensor Fusion Applications
by: Sun, Huawei, et al.
Published: (2024)