Bi-modal Prediction and Transformation Coding for Compressing Complex Human Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Hoang, Huong, Suzuki, Keito, Nguyen, Truong, Cosman, Pamela |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025)
by: Hoang, Huong, et al.
Published: (2025)
Weyl-Heisenberg Transform Capabilities in JPEG Compression Standard
by: Asiryan, V., et al.
Published: (2025)
by: Asiryan, V., et al.
Published: (2025)
DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
by: Chi, Kaiyi, et al.
Published: (2025)
by: Chi, Kaiyi, et al.
Published: (2025)
Video Quality Evaluation Methodology and Result of AV2 Compression Performance
by: Lei, Zhijun, et al.
Published: (2026)
by: Lei, Zhijun, et al.
Published: (2026)
Lossless 4:2:0 Screen Content Coding Using Luma-Guided Soft Context Formation
by: Och, Hannah, et al.
Published: (2025)
by: Och, Hannah, et al.
Published: (2025)
A multi-modal approach for identifying schizophrenia using cross-modal attention
by: Premananth, Gowtham, et al.
Published: (2023)
by: Premananth, Gowtham, et al.
Published: (2023)
Over-the-Air Learning-based Geometry Point Cloud Transmission
by: Bian, Chenghong, et al.
Published: (2023)
by: Bian, Chenghong, et al.
Published: (2023)
Federated Multi-Agent DRL for Radio Resource Management in Industrial 6G in-X subnetworks
by: Madsen, Bjarke, et al.
Published: (2024)
by: Madsen, Bjarke, et al.
Published: (2024)
A multimodal stress detection dataset with facial expressions and physiological signals
by: Hosseini, Majid, et al.
Published: (2022)
by: Hosseini, Majid, et al.
Published: (2022)
D$^2$-JSCC: Digital Deep Joint Source-channel Coding for Semantic Communications
by: Huang, Jianhao, et al.
Published: (2024)
by: Huang, Jianhao, et al.
Published: (2024)
A Survey on Multimodal Wearable Sensor-based Human Action Recognition
by: Ni, Jianyuan, et al.
Published: (2024)
by: Ni, Jianyuan, et al.
Published: (2024)
Prediction, Communication, and Computing Duration Optimization for VR Video Streaming
by: Wei, Xing, et al.
Published: (2019)
by: Wei, Xing, et al.
Published: (2019)
Low-complexity 8-point DCT Approximation Based on Angle Similarity for Image and Video Coding
by: Oliveira, R. S., et al.
Published: (2018)
by: Oliveira, R. S., et al.
Published: (2018)
Real-Time Interactive Hybrid Ocean: Spectrum-Consistent Wave Particle-FFT Coupling
by: Xue, Shengze, et al.
Published: (2025)
by: Xue, Shengze, et al.
Published: (2025)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
by: Koo, Inyong, et al.
Published: (2026)
by: Koo, Inyong, et al.
Published: (2026)
WildFusion: Multimodal Implicit 3D Reconstructions in the Wild
by: Liu, Yanbaihui, et al.
Published: (2024)
by: Liu, Yanbaihui, et al.
Published: (2024)
Graph-based Scalable Sampling of 3D Point Cloud Attributes
by: Sridhara, Shashank N., et al.
Published: (2024)
by: Sridhara, Shashank N., et al.
Published: (2024)
LATENTPATCH: A Non-Parametric Approach for Face Generation and Editing
by: Samuth, Benjamin, et al.
Published: (2024)
by: Samuth, Benjamin, et al.
Published: (2024)
Deep Joint Semantic Coding and Beamforming for Near-Space Airship-Borne Massive MIMO Network
by: Wu, Minghui, et al.
Published: (2024)
by: Wu, Minghui, et al.
Published: (2024)
Data-independent Low-complexity KLT Approximations for Image and Video Coding
by: Radünz, A. P., et al.
Published: (2021)
by: Radünz, A. P., et al.
Published: (2021)
NeuSaver: Neural Adaptive Power Consumption Optimization for Mobile Video Streaming
by: Park, Kyoungjun, et al.
Published: (2021)
by: Park, Kyoungjun, et al.
Published: (2021)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
by: Xu, Zhongweiyang, et al.
Published: (2024)
by: Xu, Zhongweiyang, et al.
Published: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
by: Chou, Huang-Cheng, et al.
Published: (2024)
by: Chou, Huang-Cheng, et al.
Published: (2024)
Token Communications: A Large Model-Driven Framework for Cross-modal Context-aware Semantic Communications
by: Qiao, Li, et al.
Published: (2025)
by: Qiao, Li, et al.
Published: (2025)
Real-time Extended Reality Video Transmission Optimization Based on Frame-priority Scheduling
by: Pan, Guangjin, et al.
Published: (2024)
by: Pan, Guangjin, et al.
Published: (2024)
A Subjective Quality Evaluation of 3D Mesh with Dynamic Level of Detail in Virtual Reality
by: Nguyen, Duc, et al.
Published: (2024)
by: Nguyen, Duc, et al.
Published: (2024)
On the Parameter Estimation of Sinusoidal Models for Speech and Audio Signals
by: Kafentzis, George P.
Published: (2024)
by: Kafentzis, George P.
Published: (2024)
Dynamic and Super-Personalized Media Ecosystem Driven by Generative AI: Unpredictable Plays Never Repeating The Same
by: Ahn, Sungjun, et al.
Published: (2024)
by: Ahn, Sungjun, et al.
Published: (2024)
Subjective Quality Assessment of Dynamic 3D Meshes in Virtual Reality Environment
by: Nguyen, Duc V., et al.
Published: (2026)
by: Nguyen, Duc V., et al.
Published: (2026)
GenHPE: Generative Counterfactuals for 3D Human Pose Estimation with Radio Frequency Signals
by: Huang, Shuokang, et al.
Published: (2025)
by: Huang, Shuokang, et al.
Published: (2025)
Generative AI Meets 6G and Beyond: Diffusion Models for Semantic Communications
by: Qin, Hai-Long, et al.
Published: (2025)
by: Qin, Hai-Long, et al.
Published: (2025)
Point Cloud in the Air
by: Shao, Yulin, et al.
Published: (2024)
by: Shao, Yulin, et al.
Published: (2024)
Full-reference Point Cloud Quality Assessment Using Spectral Graph Wavelets
by: Watanabe, Ryosuke, et al.
Published: (2024)
by: Watanabe, Ryosuke, et al.
Published: (2024)
Investigating the Generalizability of Physiological Characteristics of Anxiety
by: Zhou, Emily, et al.
Published: (2024)
by: Zhou, Emily, et al.
Published: (2024)
ArrayDPS: Unsupervised Blind Speech Separation with a Diffusion Prior
by: Xu, Zhongweiyang, et al.
Published: (2025)
by: Xu, Zhongweiyang, et al.
Published: (2025)
Recent Advances in Discrete Speech Tokens: A Review
by: Guo, Yiwei, et al.
Published: (2025)
by: Guo, Yiwei, et al.
Published: (2025)
Analyzable Chain-of-Musical-Thought Prompting for High-Fidelity Music Generation
by: Lam, Max W. Y., et al.
Published: (2025)
by: Lam, Max W. Y., et al.
Published: (2025)
Learning Temporal Resolution in Spectrogram for Audio Classification
by: Liu, Haohe, et al.
Published: (2022)
by: Liu, Haohe, et al.
Published: (2022)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
by: Liu, Haohe, et al.
Published: (2023)
by: Liu, Haohe, et al.
Published: (2023)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
by: Liu, Haohe, et al.
Published: (2024)
by: Liu, Haohe, et al.
Published: (2024)
Similar Items
-
KeyNode-Driven Geometry Coding for Real-World Scanned Human Dynamic Mesh Compression
by: Hoang, Huong, et al.
Published: (2025) -
Weyl-Heisenberg Transform Capabilities in JPEG Compression Standard
by: Asiryan, V., et al.
Published: (2025) -
DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
by: Chi, Kaiyi, et al.
Published: (2025) -
Video Quality Evaluation Methodology and Result of AV2 Compression Performance
by: Lei, Zhijun, et al.
Published: (2026) -
Lossless 4:2:0 Screen Content Coding Using Luma-Guided Soft Context Formation
by: Och, Hannah, et al.
Published: (2025)