Controlla: Learning Controllability via Graph-Constrained Latent Geometry

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Murthy, Jamuna S., Monsefi, Amin Karimi, Ramnath, Rajiv
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911690547789824
author Murthy, Jamuna S.
Monsefi, Amin Karimi
Ramnath, Rajiv
author_facet Murthy, Jamuna S.
Monsefi, Amin Karimi
Ramnath, Rajiv
contents Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While effective, such approaches do not explicitly structure how semantic attributes evolve, which can lead to identity drift and inconsistent cross-modal behavior. We propose Controlla, a modular factorized-control framework that treats controllability as a property of structured latent geometry. Controlla learns identity and attribute factors from multimodal inputs and aligns them with graph priors using graph-constrained optimal transport, encouraging attributes to follow graph-consistent trajectories while preserving reference identity. To evaluate this setting, we construct AffectHuman-43K, a leakage-aware multimodal benchmark for reference-grounded affective control, and introduce geometry-aware metrics for trajectory consistency and latent disentanglement. Experiments show consistent improvements in controllability, identity preservation, and cross-modal alignment, with additional analyses on graph sensitivity, extensibility, and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2605_16603
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Controlla: Learning Controllability via Graph-Constrained Latent Geometry
Murthy, Jamuna S.
Monsefi, Amin Karimi
Ramnath, Rajiv
Computer Vision and Pattern Recognition
Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While effective, such approaches do not explicitly structure how semantic attributes evolve, which can lead to identity drift and inconsistent cross-modal behavior. We propose Controlla, a modular factorized-control framework that treats controllability as a property of structured latent geometry. Controlla learns identity and attribute factors from multimodal inputs and aligns them with graph priors using graph-constrained optimal transport, encouraging attributes to follow graph-consistent trajectories while preserving reference identity. To evaluate this setting, we construct AffectHuman-43K, a leakage-aware multimodal benchmark for reference-grounded affective control, and introduce geometry-aware metrics for trajectory consistency and latent disentanglement. Experiments show consistent improvements in controllability, identity preservation, and cross-modal alignment, with additional analyses on graph sensitivity, extensibility, and robustness.
title Controlla: Learning Controllability via Graph-Constrained Latent Geometry
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.16603