Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sultan, Oren, Khasin, Alex, Shiran, Guy, Greenstein-Messica, Asnat, Shahaf, Dafna
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909343987793920
author Sultan, Oren
Khasin, Alex
Shiran, Guy
Greenstein-Messica, Asnat
Shahaf, Dafna
author_facet Sultan, Oren
Khasin, Alex
Shiran, Guy
Greenstein-Messica, Asnat
Shahaf, Dafna
contents We present a practical distillation approach to fine-tune LLMs for invoking tools in real-time applications. We focus on visual editing tasks; specifically, we modify images and videos by interpreting user stylistic requests, specified in natural language ("golden hour"), using an LLM to select the appropriate tools and their parameters to achieve the desired visual effect. We found that proprietary LLMs such as GPT-3.5-Turbo show potential in this task, but their high cost and latency make them unsuitable for real-time applications. In our approach, we fine-tune a (smaller) student LLM with guidance from a (larger) teacher LLM and behavioral signals. We introduce offline metrics to evaluate student LLMs. Both online and offline experiments show that our student models manage to match the performance of our teacher model (GPT-3.5-Turbo), significantly reducing costs and latency. Lastly, we show that fine-tuning was improved by 25% in low-data regimes using augmentation.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02952
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications
Sultan, Oren
Khasin, Alex
Shiran, Guy
Greenstein-Messica, Asnat
Shahaf, Dafna
Computation and Language
Artificial Intelligence
We present a practical distillation approach to fine-tune LLMs for invoking tools in real-time applications. We focus on visual editing tasks; specifically, we modify images and videos by interpreting user stylistic requests, specified in natural language ("golden hour"), using an LLM to select the appropriate tools and their parameters to achieve the desired visual effect. We found that proprietary LLMs such as GPT-3.5-Turbo show potential in this task, but their high cost and latency make them unsuitable for real-time applications. In our approach, we fine-tune a (smaller) student LLM with guidance from a (larger) teacher LLM and behavioral signals. We introduce offline metrics to evaluate student LLMs. Both online and offline experiments show that our student models manage to match the performance of our teacher model (GPT-3.5-Turbo), significantly reducing costs and latency. Lastly, we show that fine-tuning was improved by 25% in low-data regimes using augmentation.
title Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2410.02952