Representation Learning with Conditional Information Flow Maximization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Dou, Wei, Lingwei, Zhou, Wei, Hu, Songlin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911983599616000
author Hu, Dou
Wei, Lingwei
Zhou, Wei
Hu, Songlin
author_facet Hu, Dou
Wei, Lingwei
Zhou, Wei
Hu, Songlin
contents This paper proposes an information-theoretic representation learning framework, named conditional information flow maximization, to extract noise-invariant sufficient representations for the input data and target task. It promotes the learned representations have good feature uniformity and sufficient predictive ability, which can enhance the generalization of pre-trained language models (PLMs) for the target task. Firstly, an information flow maximization principle is proposed to learn more sufficient representations for the input and target by simultaneously maximizing both input-representation and representation-label mutual information. Unlike the information bottleneck, we handle the input-representation information in an opposite way to avoid the over-compression issue of latent representations. Besides, to mitigate the negative effect of potential redundant features from the input, we design a conditional information minimization principle to eliminate negative redundant features while preserve noise-invariant features. Experiments on 13 language understanding benchmarks demonstrate that our method effectively improves the performance of PLMs for classification and regression. Extensive experiments show that the learned representations are more sufficient, robust and transferable.
format Preprint
id arxiv_https___arxiv_org_abs_2406_05510
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Representation Learning with Conditional Information Flow Maximization
Hu, Dou
Wei, Lingwei
Zhou, Wei
Hu, Songlin
Machine Learning
Computation and Language
This paper proposes an information-theoretic representation learning framework, named conditional information flow maximization, to extract noise-invariant sufficient representations for the input data and target task. It promotes the learned representations have good feature uniformity and sufficient predictive ability, which can enhance the generalization of pre-trained language models (PLMs) for the target task. Firstly, an information flow maximization principle is proposed to learn more sufficient representations for the input and target by simultaneously maximizing both input-representation and representation-label mutual information. Unlike the information bottleneck, we handle the input-representation information in an opposite way to avoid the over-compression issue of latent representations. Besides, to mitigate the negative effect of potential redundant features from the input, we design a conditional information minimization principle to eliminate negative redundant features while preserve noise-invariant features. Experiments on 13 language understanding benchmarks demonstrate that our method effectively improves the performance of PLMs for classification and regression. Extensive experiments show that the learned representations are more sufficient, robust and transferable.
title Representation Learning with Conditional Information Flow Maximization
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2406.05510