DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence
Fuente:
arXiv
Saved in:
| Main Authors: | Venkit, Pranav Narayanan, Laban, Philippe, Zhou, Yilun, Huang, Kung-Hsiang, Mao, Yixin, Wu, Chien-Sheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
by: Choubey, Prafulla Kumar, et al.
Published: (2026)
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
by: Qiu, Haoyi, et al.
Published: (2025)
by: Qiu, Haoyi, et al.
Published: (2025)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
by: Venkit, Pranav Narayanan, et al.
Published: (2025)
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
The Need for a Socially-Grounded Persona Framework for User Simulation
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
by: Venkit, Pranav Narayanan, et al.
Published: (2026)
From Melting Pots to Misrepresentations: Exploring Harms in Generative AI
by: Gautam, Sanjana, et al.
Published: (2024)
by: Gautam, Sanjana, et al.
Published: (2024)
CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
by: Huang, Kung-Hsiang, et al.
Published: (2024)
by: Huang, Kung-Hsiang, et al.
Published: (2024)
TRACE: Tourism Recommendation with Accountable Citation Evidence
by: Zhao, Zixu, et al.
Published: (2026)
by: Zhao, Zixu, et al.
Published: (2026)
Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
by: Chakrabarty, Tuhin, et al.
Published: (2024)
by: Chakrabarty, Tuhin, et al.
Published: (2024)
AI-Slop to AI-Polish? Aligning Language Models through Edit-Based Writing Rewards and Test-time Computation
by: Chakrabarty, Tuhin, et al.
Published: (2025)
by: Chakrabarty, Tuhin, et al.
Published: (2025)
What if AI systems weren't chatbots?
by: Ghosh, Sourojit, et al.
Published: (2026)
by: Ghosh, Sourojit, et al.
Published: (2026)
Do Generative AI Models Output Harm while Representing Non-Western Cultures: Evidence from A Community-Centered Approach
by: Ghosh, Sourojit, et al.
Published: (2024)
by: Ghosh, Sourojit, et al.
Published: (2024)
Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge
by: Chakravorti, Tatiana, et al.
Published: (2026)
by: Chakravorti, Tatiana, et al.
Published: (2026)
Social Scientists on the Role of AI in Research
by: Chakravorti, Tatiana, et al.
Published: (2025)
by: Chakravorti, Tatiana, et al.
Published: (2025)
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
An Audit on the Perspectives and Challenges of Hallucinations in NLP
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
Attribution Gradients: Incrementally Unfolding Citations for Critical Examination of Attributed AI Answers
by: Kambhamettu, Hita, et al.
Published: (2025)
by: Kambhamettu, Hita, et al.
Published: (2025)
SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits
by: Thorat, Onkar, et al.
Published: (2024)
by: Thorat, Onkar, et al.
Published: (2024)
BingoGuard: LLM Content Moderation Tools with Risk Levels
by: Yin, Fan, et al.
Published: (2025)
by: Yin, Fan, et al.
Published: (2025)
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles
by: Huang, Kung-Hsiang, et al.
Published: (2023)
by: Huang, Kung-Hsiang, et al.
Published: (2023)
CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
Benchmarking Deep Search over Heterogeneous Enterprise Data
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
by: Choubey, Prafulla Kumar, et al.
Published: (2025)
Summary of a Haystack: A Challenge to Long-Context LLMs and RAG Systems
by: Laban, Philippe, et al.
Published: (2024)
by: Laban, Philippe, et al.
Published: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
The Unappreciated Role of Intent in Algorithmic Moderation of Social Media Content
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Hölder regularity of solutions of the steady Boltzmann equation with soft potentials
by: Wu, Kung-Chien, et al.
Published: (2024)
by: Wu, Kung-Chien, et al.
Published: (2024)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
by: Chen, Yanyu, et al.
Published: (2026)
by: Chen, Yanyu, et al.
Published: (2026)
Do RAG Systems Cover What Matters? Evaluating and Optimizing Responses with Sub-Question Coverage
by: Xie, Kaige, et al.
Published: (2024)
by: Xie, Kaige, et al.
Published: (2024)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
by: Laban, Philippe, et al.
Published: (2023)
by: Laban, Philippe, et al.
Published: (2023)
Race and Privacy in Broadcast Police Communications
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
by: Venkit, Pranav Narayanan, et al.
Published: (2024)
Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
by: Allaham, Mowafak, et al.
Published: (2026)
by: Allaham, Mowafak, et al.
Published: (2026)
Art or Artifice? Large Language Models and the False Promise of Creativity
by: Chakrabarty, Tuhin, et al.
Published: (2023)
by: Chakrabarty, Tuhin, et al.
Published: (2023)
Can Third-parties Read Our Emotions?
by: Li, Jiayi, et al.
Published: (2025)
by: Li, Jiayi, et al.
Published: (2025)
Documenting Patterns of Exoticism of Marginalized Populations within Text-to-Image Generators
by: Ghosh, Sourojit, et al.
Published: (2025)
by: Ghosh, Sourojit, et al.
Published: (2025)
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
by: Huang, Kung-Hsiang, et al.
Published: (2025)
by: Huang, Kung-Hsiang, et al.
Published: (2025)
BiblioAudit: Automated Citation Integrity & Verification System
by: Tiwari, Satyam
Published: (2026)
by: Tiwari, Satyam
Published: (2026)
Can DeepFake Speech be Reliably Detected?
by: Liu, Hongbin, et al.
Published: (2024)
by: Liu, Hongbin, et al.
Published: (2024)
Similar Items
-
Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses
by: Venkit, Pranav Narayanan, et al.
Published: (2024) -
Dont Stop Early: Scalable Enterprise Deep Research with Controlled Information Flow and Evidence-Aware Termination
by: Choubey, Prafulla Kumar, et al.
Published: (2026) -
MMPersuade: A Dataset and Evaluation Framework for Multimodal Persuasion
by: Qiu, Haoyi, et al.
Published: (2025) -
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
by: Venkit, Pranav Narayanan, et al.
Published: (2025) -
InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation
by: Li, Yu, et al.
Published: (2026)