Beyond Text: Characterizing Domain Expert Needs in Document Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gururaja, Sireesh, Gandhi, Nupoor, Milbauer, Jeremiah, Strubell, Emma
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917987631497216
author Gururaja, Sireesh
Gandhi, Nupoor
Milbauer, Jeremiah
Strubell, Emma
author_facet Gururaja, Sireesh
Gandhi, Nupoor
Milbauer, Jeremiah
Strubell, Emma
contents Working with documents is a key part of almost any knowledge work, from contextualizing research in a literature review to reviewing legal precedent. Recently, as their capabilities have expanded, primarily text-based NLP systems have often been billed as able to assist or even automate this kind of work. But to what extent are these systems able to model these tasks as experts conceptualize and perform them now? In this study, we interview sixteen domain experts across two domains to understand their processes of document research, and compare it to the current state of NLP systems. We find that our participants processes are idiosyncratic, iterative, and rely extensively on the social context of a document in addition its content; existing approaches in NLP and adjacent fields that explicitly center the document as an object, rather than as merely a container for text, tend to better reflect our participants' priorities, though they are often less accessible outside their research communities. We call on the NLP community to more carefully consider the role of the document in building useful tools that are accessible, personalizable, iterative, and socially aware.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Beyond Text: Characterizing Domain Expert Needs in Document Research
Gururaja, Sireesh
Gandhi, Nupoor
Milbauer, Jeremiah
Strubell, Emma
Computation and Language
Computers and Society
Working with documents is a key part of almost any knowledge work, from contextualizing research in a literature review to reviewing legal precedent. Recently, as their capabilities have expanded, primarily text-based NLP systems have often been billed as able to assist or even automate this kind of work. But to what extent are these systems able to model these tasks as experts conceptualize and perform them now? In this study, we interview sixteen domain experts across two domains to understand their processes of document research, and compare it to the current state of NLP systems. We find that our participants processes are idiosyncratic, iterative, and rely extensively on the social context of a document in addition its content; existing approaches in NLP and adjacent fields that explicitly center the document as an object, rather than as merely a container for text, tend to better reflect our participants' priorities, though they are often less accessible outside their research communities. We call on the NLP community to more carefully consider the role of the document in building useful tools that are accessible, personalizable, iterative, and socially aware.
title Beyond Text: Characterizing Domain Expert Needs in Document Research
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2504.12495