Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Chandrahasan, Prahaladh, Jin, Jiahe, Zhang, Zhihan, Wang, Tevin, Tang, Andy, Mo, Lucy, Ziyadi, Morteza, Ribeiro, Leonardo F. R., Qiu, Zimeng, Dreyer, Markus, Asai, Akari, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)
by: Wang, Tevin, et al.
Published: (2025)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024)
by: Wang, Tevin, et al.
Published: (2024)
Interpret and Control Dense Retrieval with Sparse Latent Features
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
by: Huang, Yukun, et al.
Published: (2026)
by: Huang, Yukun, et al.
Published: (2026)
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
by: Chen, Zijian, et al.
Published: (2026)
by: Chen, Zijian, et al.
Published: (2026)
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
by: Coelho, João, et al.
Published: (2025)
by: Coelho, João, et al.
Published: (2025)
Fine-grained Hallucination Detection and Editing for Language Models
by: Mishra, Abhika, et al.
Published: (2024)
by: Mishra, Abhika, et al.
Published: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
by: Jin, Jiahe, et al.
Published: (2025)
by: Jin, Jiahe, et al.
Published: (2025)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews
by: Liu, Mengqiao, et al.
Published: (2025)
by: Liu, Mengqiao, et al.
Published: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
by: Zeng, Yixiao, et al.
Published: (2025)
by: Zeng, Yixiao, et al.
Published: (2025)
Control Force Characteristics and Seismic Control Performance Produced by Deep Reinforcement Learning
by: Takehiko Asai
Published: (2026)
by: Takehiko Asai
Published: (2026)
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
by: Shao, Rulin, et al.
Published: (2025)
by: Shao, Rulin, et al.
Published: (2025)
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
by: Prabhakar, Akshara, et al.
Published: (2025)
by: Prabhakar, Akshara, et al.
Published: (2025)
Combating the Elsagate phenomenon: Deep learning architectures for disturbing cartoons
by: Ishikawa, Akari, et al.
Published: (2019)
by: Ishikawa, Akari, et al.
Published: (2019)
C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
by: Gao, Rui, et al.
Published: (2026)
by: Gao, Rui, et al.
Published: (2026)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
by: Xu, Sirui, et al.
Published: (2026)
by: Xu, Sirui, et al.
Published: (2026)
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
by: Ning, Jingjie, et al.
Published: (2026)
by: Ning, Jingjie, et al.
Published: (2026)
Spermidine regulates wheat grain weight at high planting density by promoting the synthesis of sucrose and starch in inferior grains
by: Juan Li, et al.
Published: (2024)
by: Juan Li, et al.
Published: (2024)
Out of Style: RAG's Fragility to Linguistic Variation
by: Cao, Tianyu, et al.
Published: (2025)
by: Cao, Tianyu, et al.
Published: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
by: Chen, Tong, et al.
Published: (2025)
by: Chen, Tong, et al.
Published: (2025)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
A High-Fidelity Surrogate Framework for Rapid Solvation of Gas-Phase and Interfacial Reaction Potential Energy Surfaces
by: Junwei Lucas, Bao, et al.
Published: (2026)
by: Junwei Lucas, Bao, et al.
Published: (2026)
Certifying Counterfactual Bias in LLMs
by: Chaudhary, Isha, et al.
Published: (2024)
by: Chaudhary, Isha, et al.
Published: (2024)
Deep Learning-based Approaches for State Space Models: A Selective Review
by: Lin, Jiahe, et al.
Published: (2024)
by: Lin, Jiahe, et al.
Published: (2024)
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
by: Guo, Zhihan, et al.
Published: (2025)
by: Guo, Zhihan, et al.
Published: (2025)
Fine-grained deterministic hardness of the shortest vector problem
by: Hittmeir, Markus
Published: (2025)
by: Hittmeir, Markus
Published: (2025)
Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
by: Deng, Yong, et al.
Published: (2025)
by: Deng, Yong, et al.
Published: (2025)
Diagnosis of Fuel Cell Health Status with Deep Sparse Auto-Encoder Neural Network
by: Fei, Chenyan, et al.
Published: (2025)
by: Fei, Chenyan, et al.
Published: (2025)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
by: Habash, Nizar, et al.
Published: (2024)
by: Habash, Nizar, et al.
Published: (2024)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
by: Yin, Zhihan, et al.
Published: (2026)
by: Yin, Zhihan, et al.
Published: (2026)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
by: Du, Mingxuan, et al.
Published: (2025)
by: Du, Mingxuan, et al.
Published: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
NeoQA: Evidence-based Question Answering with Generated News Events
by: Glockner, Max, et al.
Published: (2025)
by: Glockner, Max, et al.
Published: (2025)
Interpretable Fine‐Grained Phenotypes of Subcellular Dynamics via Unsupervised Deep Learning
by: Chuangqi Wang, et al.
Published: (2024)
by: Chuangqi Wang, et al.
Published: (2024)
Deep Learning for Crime Forecasting: The Role of Mobility at Fine-grained Spatiotemporal Scales
by: Zumel, Ariadna Albors, et al.
Published: (2025)
by: Zumel, Ariadna Albors, et al.
Published: (2025)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
by: Zhu, Fengbin, et al.
Published: (2025)
by: Zhu, Fengbin, et al.
Published: (2025)
Similar Items
-
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025) -
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024) -
Interpret and Control Dense Retrieval with Sparse Latent Features
by: Kang, Hao, et al.
Published: (2024) -
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
by: Huang, Yukun, et al.
Published: (2026) -
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
by: Chen, Zijian, et al.
Published: (2026)