Saved in:
Bibliographic Details
Main Authors: Crulis, Ben, De Runz, Cyril, Serres, Barthelemy, Venturini, Gilles
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2504.06298
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910906911293440
author Crulis, Ben
De Runz, Cyril
Serres, Barthelemy
Venturini, Gilles
author_facet Crulis, Ben
De Runz, Cyril
Serres, Barthelemy
Venturini, Gilles
contents We propose a process to compress a pre-trained Vision Language Model into a ternary version of itself instead of training a ternary model from scratch. A new initialization scheme from pre-trained weights based on the k-means algorithm is proposed to reduce the ternarization time. We implement different custom operators for executing the ternary model on the TensorFlow Lite Engine. We compare the original model with its ternary and binary versions in terms of memory consumption, inference speed and perplexity. We find that the ternary model using our custom ternary matrix multiplication operator provides a good compromise in term of memory usage and perplexity, while having the fastest token generation speed.
format Preprint
id arxiv_https___arxiv_org_abs_2504_06298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ternarization of Vision Language Models for use on edge devices
Crulis, Ben
De Runz, Cyril
Serres, Barthelemy
Venturini, Gilles
Computer Vision and Pattern Recognition
Machine Learning
We propose a process to compress a pre-trained Vision Language Model into a ternary version of itself instead of training a ternary model from scratch. A new initialization scheme from pre-trained weights based on the k-means algorithm is proposed to reduce the ternarization time. We implement different custom operators for executing the ternary model on the TensorFlow Lite Engine. We compare the original model with its ternary and binary versions in terms of memory consumption, inference speed and perplexity. We find that the ternary model using our custom ternary matrix multiplication operator provides a good compromise in term of memory usage and perplexity, while having the fastest token generation speed.
title Ternarization of Vision Language Models for use on edge devices
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2504.06298