ExioNAICS: Enterprises Level Emission Estimation Dataset with Large Language Models

Fuente: Zenodo
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Yanming, Ma, jin
Natura: Recurso digital
Lingua:inglese
Pubblicazione: Zenodo 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866902189016875008
author Guo, Yanming
Ma, jin
author_facet Guo, Yanming
Ma, jin
contents <h1>Introduction</h1> <p><strong>ExioNAICS</strong> is the first enterprise-level ML-ready benchmark dataset tailored for <strong>GHG emission estimation</strong>, bridging <strong>sector classification</strong> with carbon intensity analysis. In contrast to broad sectoral databases like <strong>ExioML</strong>, which offer global coverage of 163 sectors across 49 regions, ExioNAICS focuses on <strong>enterprise granularity</strong> by providing <strong>20,850 textual descriptions</strong> mapped to <strong>validated NAICS codes</strong> and augmented with <strong>166 sectoral carbon intensity factors</strong>. This design enables the automation of Scope 3 emission estimates (e.g., from purchased goods and services) at the <strong>firm level</strong>, a critical yet often overlooked component of supply chain emissions.</p> <p>ExioNAICS is derived from the high-quality EE-MRIO dataset, ensuring robust economic and environmental data. By integrating firm-specific text descriptions, <strong>NAICS industry labels</strong>, and <strong>ExioML-based</strong> carbon intensity factors, ExioNAICS <strong>overcomes key data bottleneck</strong>s in enterprise-level GHG accounting. It significantly <strong>lowers the entry barrier</strong> for smaller firms and researchers by standardizing data formats and linking them to a recognized classification framework.</p> <p>In demonstrating its usability, we formulate a <strong>NAICS classification</strong> and subsequent <strong>emission estimation</strong> pipeline using <strong>contrastive learning</strong> (Sentence-BERT). Our results showcase near state-of-the-art retrieval accuracy, paving the way for more <strong>accessible</strong>, <strong>cost-effective</strong>, and <strong>scalable</strong> approaches to corporate carbon accounting. ExioNAICS thus facilitates synergy between machine learning and climate research, fostering the **integration** of advanced NLP techniques in eco-economic studies at the enterprise scale.</p> <h1>Dataset</h1> <p>ExioNAICS serves as a <strong>hybrid textual and numeric</strong> dataset, capturing both <strong>enterprise descriptions</strong> (text modality) and <strong>sectoral carbon intensity factors</strong> (numeric modality). These data components are linked through <strong>NAICS codes</strong>, allowing end-to-end modeling of how enterprise descriptions map to sector emission intensities. Key dataset features include:</p> <p>- Enterprise Description<br>- NAICS Description<br>- Sectoral Emission Factor<br>- Over 20,000 textual entries<br>- Hierarchical coverage: NAICS 2–6 digit codes (20 to 1,114 categories)</p> <h1>NAICS Classification</h1> <p>NAICS Classification is a fundamental component of <strong>enterprise-level GHG emission estimation</strong>. By assigning each firm to the appropriate sector category, practitioners can reference the corresponding carbon intensity factors, facilitating more accurate reporting. ExioNAICS adopts a <strong>natural language processing</strong> approach to <strong>NAICS classification</strong>, treating the task as an information retrieval problem. </p> <p>Each <strong>enterprise description</strong> (query) is encoded separately, and matched against <strong>NAICS descriptions</strong> (corpus) based on the cosine similarity of their embeddings. This methodology leverages a dual-tower architecture, wherein the first tower processes the query (enterprise text) and the second tower processes NAICS descriptions. </p> <p>We apply <strong>machine learning</strong> to fine-tune a pre-trained Sentence-BERT model. Zero-shot SBERT models may achieve only around <strong>20% Top-1 accuracy</strong> on the <strong>1000 classes sector classification task</strong>, whereas <strong>contrastive fine-tuning raises this to over 75%</strong>. Further <strong>preprocessing exceeding 77% Top-1 accuracy</strong>, such as lowercasing and URL removal, can add incremental gains, leading to state-of-the-art results.</p> <h1>Versions</h1> <p>Version 1 using ExioML as Emission Factor, Version 2 using EPA as Emission Factor.</p> <h1>Citation</h1> <pre>@article{guo2025group, title={Group Reasoning Emission Estimation Networks}, author={Guo, Yanming and Qian, Xiao and Credit, Kevin and Ma, Jin}, journal={arXiv preprint arXiv:2502.06874}, year={2025} }</pre>
format Recurso digital
id zenodo_https___doi_org_10_5281_zenodo_15010461
institution Zenodo
language eng
publishDate 2025
publisher Zenodo
record_format zenodo
spellingShingle ExioNAICS: Enterprises Level Emission Estimation Dataset with Large Language Models
Guo, Yanming
Ma, jin
Machine learning
Emission estimation
Life cycle
Environmental sustainability
Natural language processing
<h1>Introduction</h1> <p><strong>ExioNAICS</strong> is the first enterprise-level ML-ready benchmark dataset tailored for <strong>GHG emission estimation</strong>, bridging <strong>sector classification</strong> with carbon intensity analysis. In contrast to broad sectoral databases like <strong>ExioML</strong>, which offer global coverage of 163 sectors across 49 regions, ExioNAICS focuses on <strong>enterprise granularity</strong> by providing <strong>20,850 textual descriptions</strong> mapped to <strong>validated NAICS codes</strong> and augmented with <strong>166 sectoral carbon intensity factors</strong>. This design enables the automation of Scope 3 emission estimates (e.g., from purchased goods and services) at the <strong>firm level</strong>, a critical yet often overlooked component of supply chain emissions.</p> <p>ExioNAICS is derived from the high-quality EE-MRIO dataset, ensuring robust economic and environmental data. By integrating firm-specific text descriptions, <strong>NAICS industry labels</strong>, and <strong>ExioML-based</strong> carbon intensity factors, ExioNAICS <strong>overcomes key data bottleneck</strong>s in enterprise-level GHG accounting. It significantly <strong>lowers the entry barrier</strong> for smaller firms and researchers by standardizing data formats and linking them to a recognized classification framework.</p> <p>In demonstrating its usability, we formulate a <strong>NAICS classification</strong> and subsequent <strong>emission estimation</strong> pipeline using <strong>contrastive learning</strong> (Sentence-BERT). Our results showcase near state-of-the-art retrieval accuracy, paving the way for more <strong>accessible</strong>, <strong>cost-effective</strong>, and <strong>scalable</strong> approaches to corporate carbon accounting. ExioNAICS thus facilitates synergy between machine learning and climate research, fostering the **integration** of advanced NLP techniques in eco-economic studies at the enterprise scale.</p> <h1>Dataset</h1> <p>ExioNAICS serves as a <strong>hybrid textual and numeric</strong> dataset, capturing both <strong>enterprise descriptions</strong> (text modality) and <strong>sectoral carbon intensity factors</strong> (numeric modality). These data components are linked through <strong>NAICS codes</strong>, allowing end-to-end modeling of how enterprise descriptions map to sector emission intensities. Key dataset features include:</p> <p>- Enterprise Description<br>- NAICS Description<br>- Sectoral Emission Factor<br>- Over 20,000 textual entries<br>- Hierarchical coverage: NAICS 2–6 digit codes (20 to 1,114 categories)</p> <h1>NAICS Classification</h1> <p>NAICS Classification is a fundamental component of <strong>enterprise-level GHG emission estimation</strong>. By assigning each firm to the appropriate sector category, practitioners can reference the corresponding carbon intensity factors, facilitating more accurate reporting. ExioNAICS adopts a <strong>natural language processing</strong> approach to <strong>NAICS classification</strong>, treating the task as an information retrieval problem. </p> <p>Each <strong>enterprise description</strong> (query) is encoded separately, and matched against <strong>NAICS descriptions</strong> (corpus) based on the cosine similarity of their embeddings. This methodology leverages a dual-tower architecture, wherein the first tower processes the query (enterprise text) and the second tower processes NAICS descriptions. </p> <p>We apply <strong>machine learning</strong> to fine-tune a pre-trained Sentence-BERT model. Zero-shot SBERT models may achieve only around <strong>20% Top-1 accuracy</strong> on the <strong>1000 classes sector classification task</strong>, whereas <strong>contrastive fine-tuning raises this to over 75%</strong>. Further <strong>preprocessing exceeding 77% Top-1 accuracy</strong>, such as lowercasing and URL removal, can add incremental gains, leading to state-of-the-art results.</p> <h1>Versions</h1> <p>Version 1 using ExioML as Emission Factor, Version 2 using EPA as Emission Factor.</p> <h1>Citation</h1> <pre>@article{guo2025group, title={Group Reasoning Emission Estimation Networks}, author={Guo, Yanming and Qian, Xiao and Credit, Kevin and Ma, Jin}, journal={arXiv preprint arXiv:2502.06874}, year={2025} }</pre>
title ExioNAICS: Enterprises Level Emission Estimation Dataset with Large Language Models
topic Machine learning
Emission estimation
Life cycle
Environmental sustainability
Natural language processing
url https://doi.org/10.5281/zenodo.15010461