Bi-modal Prediction and Transformation Coding for Compressing Complex Human Dynamics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hoang, Huong, Suzuki, Keito, Nguyen, Truong, Cosman, Pamela
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915504804855808
author Hoang, Huong
Suzuki, Keito
Nguyen, Truong
Cosman, Pamela
author_facet Hoang, Huong
Suzuki, Keito
Nguyen, Truong
Cosman, Pamela
contents For dynamic human motion sequences, the original KeyNode-Driven codec often struggles to retain compression efficiency when confronted with rapid movements or strong non-rigid deformations. This paper proposes a novel Bi-modal coding framework that enhances the flexibility of motion representation by integrating semantic segmentation and region-specific transformation modeling. The rigid transformation model (rotation & translation) is extended with a hybrid scheme that selectively applies affine transformations-rotation, translation, scaling, and shearing-only to deformation-rich regions (e.g., the torso, where loose clothing induces high variability), while retaining rigid models elsewhere. The affine model is decomposed into minimal parameter sets for efficient coding and combined through a component selection strategy guided by a Lagrangian Rate-Distortion optimization. The results show that the Bi-modal method achieves more accurate mesh deformation, especially in sequences involving complex non-rigid motion, without compromising compression efficiency in simpler regions, with an average bit-rate saving of 33.81% compared to the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bi-modal Prediction and Transformation Coding for Compressing Complex Human Dynamics
Hoang, Huong
Suzuki, Keito
Nguyen, Truong
Cosman, Pamela
Signal Processing
Multimedia
For dynamic human motion sequences, the original KeyNode-Driven codec often struggles to retain compression efficiency when confronted with rapid movements or strong non-rigid deformations. This paper proposes a novel Bi-modal coding framework that enhances the flexibility of motion representation by integrating semantic segmentation and region-specific transformation modeling. The rigid transformation model (rotation & translation) is extended with a hybrid scheme that selectively applies affine transformations-rotation, translation, scaling, and shearing-only to deformation-rich regions (e.g., the torso, where loose clothing induces high variability), while retaining rigid models elsewhere. The affine model is decomposed into minimal parameter sets for efficient coding and combined through a component selection strategy guided by a Lagrangian Rate-Distortion optimization. The results show that the Bi-modal method achieves more accurate mesh deformation, especially in sequences involving complex non-rigid motion, without compromising compression efficiency in simpler regions, with an average bit-rate saving of 33.81% compared to the baseline.
title Bi-modal Prediction and Transformation Coding for Compressing Complex Human Dynamics
topic Signal Processing
Multimedia
url https://arxiv.org/abs/2509.16919