Evaluating Large Language Models for Causal Modeling

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Razouk, Houssam, Benischke, Leonie, Niess, Georg, Kern, Roman
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913584794042368
author Razouk, Houssam
Benischke, Leonie
Niess, Georg
Kern, Roman
author_facet Razouk, Houssam
Benischke, Leonie
Niess, Georg
Kern, Roman
contents In this paper, we consider the process of transforming causal domain knowledge into a representation that aligns more closely with guidelines from causal data science. To this end, we introduce two novel tasks related to distilling causal domain knowledge into causal variables and detecting interaction entities using LLMs. We have determined that contemporary LLMs are helpful tools for conducting causal modeling tasks in collaboration with human experts, as they can provide a wider perspective. Specifically, LLMs, such as GPT-4-turbo and Llama3-70b, perform better in distilling causal domain knowledge into causal variables compared to sparse expert models, such as Mixtral-8x22b. On the contrary, sparse expert models such as Mixtral-8x22b stand out as the most effective in identifying interaction entities. Finally, we highlight the dependency between the domain where the entities are generated and the performance of the chosen LLM for causal modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15888
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Evaluating Large Language Models for Causal Modeling
Razouk, Houssam
Benischke, Leonie
Niess, Georg
Kern, Roman
Computation and Language
I.2.0
I.2.0
In this paper, we consider the process of transforming causal domain knowledge into a representation that aligns more closely with guidelines from causal data science. To this end, we introduce two novel tasks related to distilling causal domain knowledge into causal variables and detecting interaction entities using LLMs. We have determined that contemporary LLMs are helpful tools for conducting causal modeling tasks in collaboration with human experts, as they can provide a wider perspective. Specifically, LLMs, such as GPT-4-turbo and Llama3-70b, perform better in distilling causal domain knowledge into causal variables compared to sparse expert models, such as Mixtral-8x22b. On the contrary, sparse expert models such as Mixtral-8x22b stand out as the most effective in identifying interaction entities. Finally, we highlight the dependency between the domain where the entities are generated and the performance of the chosen LLM for causal modeling.
title Evaluating Large Language Models for Causal Modeling
topic Computation and Language
I.2.0
I.2.0
url https://arxiv.org/abs/2411.15888