Tasks and Roles in Legal AI: Data Curation, Annotation, and Verification
Fuente:
arXiv
Saved in:
| Main Authors: | Koenecke, Allison, Stiglitz, Jed, Mimno, David, Wilkens, Matthew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automate or Assist? The Role of Computational Models in Identifying Gendered Discourse in US Capital Trial Transcripts
by: Wen-Yi, Andrea W, et al.
Published: (2024)
by: Wen-Yi, Andrea W, et al.
Published: (2024)
[Lions: 1] and [Tigers: 2] and [Bears: 3], Oh My! Literary Coreference Annotation with LLMs
by: Hicke, Rebecca M. M., et al.
Published: (2024)
by: Hicke, Rebecca M. M., et al.
Published: (2024)
The Afterlives of Shakespeare and Company in Online Social Readership
by: Antoniak, Maria, et al.
Published: (2024)
by: Antoniak, Maria, et al.
Published: (2024)
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
A City of Millions: Mapping Literary Social Networks At Scale
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
AI as a Tool for Simulation-Based Experiments in Literary Studies
by: Wilkens, Matthew
Published: (2026)
by: Wilkens, Matthew
Published: (2026)
Lost in Space: Finding the Right Tokens for Structured Output
by: Hamilton, Sil, et al.
Published: (2025)
by: Hamilton, Sil, et al.
Published: (2025)
Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese
by: Lyu, Hanjia, et al.
Published: (2025)
by: Lyu, Hanjia, et al.
Published: (2025)
Are You There God? Lightweight Narrative Annotation of Christian Fiction with LMs
by: Hicke, Rebecca M. M., et al.
Published: (2025)
by: Hicke, Rebecca M. M., et al.
Published: (2025)
Stronger Random Baselines for In-Context Learning
by: Yauney, Gregory, et al.
Published: (2024)
by: Yauney, Gregory, et al.
Published: (2024)
Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories
by: Hamilton, Sil, et al.
Published: (2026)
by: Hamilton, Sil, et al.
Published: (2026)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
Into the Unknown: Accounting for Missing Demographic Data when Mitigating Ad Delivery Skew
by: Corpus, Isabel, et al.
Published: (2026)
by: Corpus, Isabel, et al.
Published: (2026)
Looking for the Inner Music: Probing LLMs' Understanding of Literary Style
by: Hicke, Rebecca M. M., et al.
Published: (2025)
by: Hicke, Rebecca M. M., et al.
Published: (2025)
Show or Tell? Modeling the evolution of request-making in Human-LLM conversations
by: Zhu, Shengqi, et al.
Published: (2025)
by: Zhu, Shengqi, et al.
Published: (2025)
Priming, Path-dependence, and Plasticity: Understanding the molding of user-LLM interaction and its implications from (many) chat logs in the wild
by: Zhu, Shengqi, et al.
Published: (2026)
by: Zhu, Shengqi, et al.
Published: (2026)
Analyzing Dialectical Biases in LLMs for Knowledge and Reasoning Benchmarks
by: Pan, Eileen, et al.
Published: (2025)
by: Pan, Eileen, et al.
Published: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
Evaluating the Role of Verifiers in Test-Time Scaling for Legal Reasoning Tasks
by: Romano, Davide, et al.
Published: (2025)
by: Romano, Davide, et al.
Published: (2025)
The Zero Body Problem: Probing LLM Use of Sensory Language
by: Hicke, Rebecca M. M., et al.
Published: (2025)
by: Hicke, Rebecca M. M., et al.
Published: (2025)
Leveraging Machine Learning to Detect Data Curation Activities
by: Lafia, Sara, et al.
Published: (2021)
by: Lafia, Sara, et al.
Published: (2021)
Challenges and Considerations in Annotating Legal Data: A Comprehensive Overview
by: Darji, Harshil, et al.
Published: (2024)
by: Darji, Harshil, et al.
Published: (2024)
Digital Gatekeepers: Google's Role in Curating Hashtags and Subreddits
by: Poudel, Amrit, et al.
Published: (2025)
by: Poudel, Amrit, et al.
Published: (2025)
GDC Cohort Copilot: An AI Copilot for Curating Cohorts from the Genomic Data Commons
by: Song, Steven, et al.
Published: (2025)
by: Song, Steven, et al.
Published: (2025)
Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic Perspectives
by: Xu, Yinuo, et al.
Published: (2025)
by: Xu, Yinuo, et al.
Published: (2025)
Curating Grounded Synthetic Data with Global Perspectives for Equitable AI
by: Törnquist, Elin, et al.
Published: (2024)
by: Törnquist, Elin, et al.
Published: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
by: Mei, Katelyn Xiaoying, et al.
Published: (2025)
by: Mei, Katelyn Xiaoying, et al.
Published: (2025)
ACORD: An Expert-Annotated Retrieval Dataset for Legal Contract Drafting
by: Wang, Steven H., et al.
Published: (2025)
by: Wang, Steven H., et al.
Published: (2025)
Do Chinese models speak Chinese languages?
by: Wen-Yi, Andrea W, et al.
Published: (2025)
by: Wen-Yi, Andrea W, et al.
Published: (2025)
LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
by: Fujita, Shogo, et al.
Published: (2025)
by: Fujita, Shogo, et al.
Published: (2025)
Lawma: The Power of Specialization for Legal Annotation
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
by: Dominguez-Olmedo, Ricardo, et al.
Published: (2024)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
by: Eisenstein, Jacob, et al.
Published: (2022)
by: Eisenstein, Jacob, et al.
Published: (2022)
Automated Verification of Monotonic Data Structure Traversals in C
by: Sotoudeh, Matthew
Published: (2025)
by: Sotoudeh, Matthew
Published: (2025)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
by: Idahl, Maximilian, et al.
Published: (2026)
by: Idahl, Maximilian, et al.
Published: (2026)
Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools
by: Magesh, Varun, et al.
Published: (2024)
by: Magesh, Varun, et al.
Published: (2024)
How Chinese are Chinese Language Models? The Puzzling Lack of Language Policy in China's LLMs
by: Wen-Yi, Andrea W, et al.
Published: (2024)
by: Wen-Yi, Andrea W, et al.
Published: (2024)
Toxicity of the Commons: Curating Open-Source Pre-Training Data
by: Arnett, Catherine, et al.
Published: (2024)
by: Arnett, Catherine, et al.
Published: (2024)
Automated Data Curation for Robust Language Model Fine-Tuning
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
On Crowdsourcing Task Design for Discourse Relation Annotation
by: Yung, Frances, et al.
Published: (2024)
by: Yung, Frances, et al.
Published: (2024)
Similar Items
-
Automate or Assist? The Role of Computational Models in Identifying Gendered Discourse in US Capital Trial Transcripts
by: Wen-Yi, Andrea W, et al.
Published: (2024) -
[Lions: 1] and [Tigers: 2] and [Bears: 3], Oh My! Literary Coreference Annotation with LLMs
by: Hicke, Rebecca M. M., et al.
Published: (2024) -
The Afterlives of Shakespeare and Company in Online Social Readership
by: Antoniak, Maria, et al.
Published: (2024) -
Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
by: Hamilton, Sil, et al.
Published: (2025) -
A City of Millions: Mapping Literary Social Networks At Scale
by: Hamilton, Sil, et al.
Published: (2025)