Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Yu, Yongsheng, Luo, Jiebo
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914810317242368
author Yu, Yongsheng
Luo, Jiebo
author_facet Yu, Yongsheng
Luo, Jiebo
contents Conventional demographic inference methods have predominantly operated under the supervision of accurately labeled data, yet struggle to adapt to shifting social landscapes and diverse cultural contexts, leading to narrow specialization and limited accuracy in applications. Recently, the emergence of large multimodal models (LMMs) has shown transformative potential across various research tasks, such as visual comprehension and description. In this study, we explore the application of LMMs to demographic inference and introduce a benchmark for both quantitative and qualitative evaluation. Our findings indicate that LMMs possess advantages in zero-shot learning, interpretability, and handling uncurated 'in-the-wild' inputs, albeit with a propensity for off-target predictions. To enhance LMM performance and achieve comparability with supervised learning baselines, we propose a Chain-of-Thought augmented prompting approach, which effectively mitigates the off-target prediction issue.
format Preprint
id arxiv_https___arxiv_org_abs_2405_15687
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
Yu, Yongsheng
Luo, Jiebo
Computer Vision and Pattern Recognition
Machine Learning
Conventional demographic inference methods have predominantly operated under the supervision of accurately labeled data, yet struggle to adapt to shifting social landscapes and diverse cultural contexts, leading to narrow specialization and limited accuracy in applications. Recently, the emergence of large multimodal models (LMMs) has shown transformative potential across various research tasks, such as visual comprehension and description. In this study, we explore the application of LMMs to demographic inference and introduce a benchmark for both quantitative and qualitative evaluation. Our findings indicate that LMMs possess advantages in zero-shot learning, interpretability, and handling uncurated 'in-the-wild' inputs, albeit with a propensity for off-target predictions. To enhance LMM performance and achieve comparability with supervised learning baselines, we propose a Chain-of-Thought augmented prompting approach, which effectively mitigates the off-target prediction issue.
title Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2405.15687