Pixel Sentence Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Chenghao, Huang, Zhuoxu, Chen, Danlu, Hudson, G Thomas, Li, Yizhi, Duan, Haoran, Lin, Chenghua, Fu, Jie, Han, Jungong, Moubayed, Noura Al
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911776214351872
author Xiao, Chenghao
Huang, Zhuoxu
Chen, Danlu
Hudson, G Thomas
Li, Yizhi
Duan, Haoran
Lin, Chenghua
Fu, Jie
Han, Jungong
Moubayed, Noura Al
author_facet Xiao, Chenghao
Huang, Zhuoxu
Chen, Danlu
Hudson, G Thomas
Li, Yizhi
Duan, Haoran
Lin, Chenghua
Fu, Jie
Han, Jungong
Moubayed, Noura Al
contents Pretrained language models are long known to be subpar in capturing sentence and document-level semantics. Though heavily investigated, transferring perturbation-based methods from unsupervised visual representation learning to NLP remains an unsolved problem. This is largely due to the discreteness of subword units brought by tokenization of language models, limiting small perturbations of inputs to form semantics-preserved positive pairs. In this work, we conceptualize the learning of sentence-level textual semantics as a visual representation learning process. Drawing from cognitive and linguistic sciences, we introduce an unsupervised visual sentence representation learning framework, employing visually-grounded text perturbation methods like typos and word order shuffling, resonating with human cognitive patterns, and enabling perturbation to texts to be perceived as continuous. Our approach is further bolstered by large-scale unsupervised topical alignment training and natural language inference supervision, achieving comparable performance in semantic textual similarity (STS) to existing state-of-the-art NLP methods. Additionally, we unveil our method's inherent zero-shot cross-lingual transferability and a unique leapfrogging pattern across languages during iterative training. To our knowledge, this is the first representation learning method devoid of traditional language models for understanding sentence and document semantics, marking a stride closer to human-like textual comprehension. Our code is available at https://github.com/gowitheflow-1998/Pixel-Linguist
format Preprint
id arxiv_https___arxiv_org_abs_2402_08183
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Pixel Sentence Representation Learning
Xiao, Chenghao
Huang, Zhuoxu
Chen, Danlu
Hudson, G Thomas
Li, Yizhi
Duan, Haoran
Lin, Chenghua
Fu, Jie
Han, Jungong
Moubayed, Noura Al
Computation and Language
Computer Vision and Pattern Recognition
Pretrained language models are long known to be subpar in capturing sentence and document-level semantics. Though heavily investigated, transferring perturbation-based methods from unsupervised visual representation learning to NLP remains an unsolved problem. This is largely due to the discreteness of subword units brought by tokenization of language models, limiting small perturbations of inputs to form semantics-preserved positive pairs. In this work, we conceptualize the learning of sentence-level textual semantics as a visual representation learning process. Drawing from cognitive and linguistic sciences, we introduce an unsupervised visual sentence representation learning framework, employing visually-grounded text perturbation methods like typos and word order shuffling, resonating with human cognitive patterns, and enabling perturbation to texts to be perceived as continuous. Our approach is further bolstered by large-scale unsupervised topical alignment training and natural language inference supervision, achieving comparable performance in semantic textual similarity (STS) to existing state-of-the-art NLP methods. Additionally, we unveil our method's inherent zero-shot cross-lingual transferability and a unique leapfrogging pattern across languages during iterative training. To our knowledge, this is the first representation learning method devoid of traditional language models for understanding sentence and document semantics, marking a stride closer to human-like textual comprehension. Our code is available at https://github.com/gowitheflow-1998/Pixel-Linguist
title Pixel Sentence Representation Learning
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2402.08183