AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Pengfei, Xie, Tianxin, Yang, Minghao, Liu, Li
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917541631229952
author Zhang, Pengfei
Xie, Tianxin
Yang, Minghao
Liu, Li
author_facet Zhang, Pengfei
Xie, Tianxin
Yang, Minghao
Liu, Li
contents REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.
format Preprint
id arxiv_https___arxiv_org_abs_2603_01006
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
Zhang, Pengfei
Xie, Tianxin
Yang, Minghao
Liu, Li
Sound
Artificial Intelligence
Machine Learning
Multimedia
68T07, 62M45
I.2.6; I.5.4; G.3
REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth. In this work, we introduce Attribution-Guided REPresentation Alignment (AG-REPA), a novel causal layer selection strategy for representation alignment in audio Flow Matching. Firstly, we find that layers that best store semantic/acoustic information (high teacher-space similarity) are not necessarily the layers that contribute most to the velocity field that drives generation, and we call it Store-Contribute Dissociation (SCD). To turn this insight into an actionable training guidance, we propose a forward-only gate ablation (FoG-A) that quantifies each layer's causal contribution via the induced change in the predicted velocity field, enabling sparse layer selection and adaptive weighting for alignment. Across unified speech and general-audio training (LibriSpeech + AudioSet) under different token-conditioning topologies, AG-REPA consistently outperforms REPA baselines. Overall, our results show that alignment is most effective when applied to the causally dominant layers that drive the velocity field, rather than to layers that are representationally rich but functionally passive.
title AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
topic Sound
Artificial Intelligence
Machine Learning
Multimedia
68T07, 62M45
I.2.6; I.5.4; G.3
url https://arxiv.org/abs/2603.01006