CoCoHD: Congress Committee Hearing Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hiray, Arnav, Liu, Yunsong, Song, Mingxiao, Shah, Agam, Chava, Sudheer
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910632358445056
author Hiray, Arnav
Liu, Yunsong
Song, Mingxiao
Shah, Agam
Chava, Sudheer
author_facet Hiray, Arnav
Liu, Yunsong
Song, Mingxiao
Shah, Agam
Chava, Sudheer
contents U.S. congressional hearings significantly influence the national economy and social fabric, impacting individual lives. Despite their importance, there is a lack of comprehensive datasets for analyzing these discourses. To address this, we propose the Congress Committee Hearing Dataset (CoCoHD), covering hearings from 1997 to 2024 across 86 committees, with 32,697 records. This dataset enables researchers to study policy language on critical issues like healthcare, LGBTQ+ rights, and climate justice. We demonstrate its potential with a case study on 1,000 energy-related sentences, analyzing the Energy and Commerce Committee's stance on fossil fuel consumption. By fine-tuning pre-trained language models, we create energy-relevant measures for each hearing. Our market analysis shows that natural language analysis using CoCoHD can predict and highlight trends in the energy sector.
format Preprint
id arxiv_https___arxiv_org_abs_2410_03099
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CoCoHD: Congress Committee Hearing Dataset
Hiray, Arnav
Liu, Yunsong
Song, Mingxiao
Shah, Agam
Chava, Sudheer
Computation and Language
U.S. congressional hearings significantly influence the national economy and social fabric, impacting individual lives. Despite their importance, there is a lack of comprehensive datasets for analyzing these discourses. To address this, we propose the Congress Committee Hearing Dataset (CoCoHD), covering hearings from 1997 to 2024 across 86 committees, with 32,697 records. This dataset enables researchers to study policy language on critical issues like healthcare, LGBTQ+ rights, and climate justice. We demonstrate its potential with a case study on 1,000 energy-related sentences, analyzing the Energy and Commerce Committee's stance on fossil fuel consumption. By fine-tuning pre-trained language models, we create energy-relevant measures for each hearing. Our market analysis shows that natural language analysis using CoCoHD can predict and highlight trends in the energy sector.
title CoCoHD: Congress Committee Hearing Dataset
topic Computation and Language
url https://arxiv.org/abs/2410.03099