DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909771270979584 |
|---|---|
| author | Venkit, Pranav Narayanan Laban, Philippe Zhou, Yilun Huang, Kung-Hsiang Mao, Yixin Wu, Chien-Sheng |
| author_facet | Venkit, Pranav Narayanan Laban, Philippe Zhou, Yilun Huang, Kung-Hsiang Mao, Yixin Wu, Chien-Sheng |
| contents | Generative search engines and deep research LLM agents promise trustworthy, source-grounded synthesis, yet users regularly encounter overconfidence, weak sourcing, and confusing citation practices. We introduce DeepTRACE, a novel sociotechnically grounded audit framework that turns prior community-identified failure cases into eight measurable dimensions spanning answer text, sources, and citations. DeepTRACE uses statement-level analysis (decomposition, confidence scoring) and builds citation and factual-support matrices to audit how systems reason with and attribute evidence end-to-end. Using automated extraction pipelines for popular public models (e.g., GPT-4.5/5, You.com, Perplexity, Copilot/Bing, Gemini) and an LLM-judge with validated agreement to human raters, we evaluate both web-search engines and deep-research configurations. Our findings show that generative search engines and deep research agents frequently produce one-sided, highly confident responses on debate queries and include large fractions of statements unsupported by their own listed sources. Deep-research configurations reduce overconfidence and can attain high citation thoroughness, but they remain highly one-sided on debate queries and still exhibit large fractions of unsupported statements, with citation accuracy ranging from 40--80% across systems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_04499 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence Venkit, Pranav Narayanan Laban, Philippe Zhou, Yilun Huang, Kung-Hsiang Mao, Yixin Wu, Chien-Sheng Computation and Language Artificial Intelligence Generative search engines and deep research LLM agents promise trustworthy, source-grounded synthesis, yet users regularly encounter overconfidence, weak sourcing, and confusing citation practices. We introduce DeepTRACE, a novel sociotechnically grounded audit framework that turns prior community-identified failure cases into eight measurable dimensions spanning answer text, sources, and citations. DeepTRACE uses statement-level analysis (decomposition, confidence scoring) and builds citation and factual-support matrices to audit how systems reason with and attribute evidence end-to-end. Using automated extraction pipelines for popular public models (e.g., GPT-4.5/5, You.com, Perplexity, Copilot/Bing, Gemini) and an LLM-judge with validated agreement to human raters, we evaluate both web-search engines and deep-research configurations. Our findings show that generative search engines and deep research agents frequently produce one-sided, highly confident responses on debate queries and include large fractions of statements unsupported by their own listed sources. Deep-research configurations reduce overconfidence and can attain high citation thoroughness, but they remain highly one-sided on debate queries and still exhibit large fractions of unsupported statements, with citation accuracy ranging from 40--80% across systems. |
| title | DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2509.04499 |