NARAIM: Native Aspect Ratio Autoregressive Image Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fernández, Daniel Gallo, van der Klis, Robert, Matişan, Răzvan-Andrei, Partyka, Janusz, Gavves, Efstratios, Papa, Samuele, Lippe, Phillip
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913598836572160
author Fernández, Daniel Gallo
van der Klis, Robert
Matişan, Răzvan-Andrei
Partyka, Janusz
Gavves, Efstratios
Papa, Samuele
Lippe, Phillip
author_facet Fernández, Daniel Gallo
van der Klis, Robert
Matişan, Răzvan-Andrei
Partyka, Janusz
Gavves, Efstratios
Papa, Samuele
Lippe, Phillip
contents While vision transformers are able to solve a wide variety of computer vision tasks, no pre-training method has yet demonstrated the same scaling laws as observed in language models. Autoregressive models show promising results, but are commonly trained on images that are cropped or transformed into square images, which distorts or destroys information present in the input. To overcome this limitation, we propose NARAIM, a vision model pre-trained with an autoregressive objective that uses images in their native aspect ratio. By maintaining the native aspect ratio, we preserve the original spatial context, thereby enhancing the model's ability to interpret visual information. In our experiments, we show that maintaining the aspect ratio improves performance on a downstream classification task.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10012
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NARAIM: Native Aspect Ratio Autoregressive Image Models
Fernández, Daniel Gallo
van der Klis, Robert
Matişan, Răzvan-Andrei
Partyka, Janusz
Gavves, Efstratios
Papa, Samuele
Lippe, Phillip
Computer Vision and Pattern Recognition
While vision transformers are able to solve a wide variety of computer vision tasks, no pre-training method has yet demonstrated the same scaling laws as observed in language models. Autoregressive models show promising results, but are commonly trained on images that are cropped or transformed into square images, which distorts or destroys information present in the input. To overcome this limitation, we propose NARAIM, a vision model pre-trained with an autoregressive objective that uses images in their native aspect ratio. By maintaining the native aspect ratio, we preserve the original spatial context, thereby enhancing the model's ability to interpret visual information. In our experiments, we show that maintaining the aspect ratio improves performance on a downstream classification task.
title NARAIM: Native Aspect Ratio Autoregressive Image Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.10012