Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Chandrahasan, Prahaladh, Jin, Jiahe, Zhang, Zhihan, Wang, Tevin, Tang, Andy, Mo, Lucy, Ziyadi, Morteza, Ribeiro, Leonardo F. R., Qiu, Zimeng, Dreyer, Markus, Asai, Akari, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
par: Wang, Tevin, et autres
Publié: (2025)
par: Wang, Tevin, et autres
Publié: (2025)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
par: Wang, Tevin, et autres
Publié: (2024)
par: Wang, Tevin, et autres
Publié: (2024)
Interpret and Control Dense Retrieval with Sparse Latent Features
par: Kang, Hao, et autres
Publié: (2024)
par: Kang, Hao, et autres
Publié: (2024)
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
par: Huang, Yukun, et autres
Publié: (2026)
par: Huang, Yukun, et autres
Publié: (2026)
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
par: Chen, Zijian, et autres
Publié: (2026)
par: Chen, Zijian, et autres
Publié: (2026)
DeepResearchGym: A Free, Transparent, and Reproducible Evaluation Sandbox for Deep Research
par: Coelho, João, et autres
Publié: (2025)
par: Coelho, João, et autres
Publié: (2025)
Fine-grained Hallucination Detection and Editing for Language Models
par: Mishra, Abhika, et autres
Publié: (2024)
par: Mishra, Abhika, et autres
Publié: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
par: Kang, Hao, et autres
Publié: (2025)
par: Kang, Hao, et autres
Publié: (2025)
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
par: Jin, Jiahe, et autres
Publié: (2025)
par: Jin, Jiahe, et autres
Publié: (2025)
ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
par: Kang, Hao, et autres
Publié: (2024)
par: Kang, Hao, et autres
Publié: (2024)
Understand User Opinions of Large Language Models via LLM-Powered In-the-Moment User Experience Interviews
par: Liu, Mengqiao, et autres
Publié: (2025)
par: Liu, Mengqiao, et autres
Publié: (2025)
RARE: Retrieval-Aware Robustness Evaluation for Retrieval-Augmented Generation Systems
par: Zeng, Yixiao, et autres
Publié: (2025)
par: Zeng, Yixiao, et autres
Publié: (2025)
Control Force Characteristics and Seismic Control Performance Produced by Deep Reinforcement Learning
par: Takehiko Asai
Publié: (2026)
par: Takehiko Asai
Publié: (2026)
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
par: Shao, Rulin, et autres
Publié: (2025)
par: Shao, Rulin, et autres
Publié: (2025)
SkillLearnBench: Benchmarking Continual Learning Methods for Agent Skill Generation on Real-World Tasks
par: Zhong, Shanshan, et autres
Publié: (2026)
par: Zhong, Shanshan, et autres
Publié: (2026)
Enterprise Deep Research: Steerable Multi-Agent Deep Research for Enterprise Analytics
par: Prabhakar, Akshara, et autres
Publié: (2025)
par: Prabhakar, Akshara, et autres
Publié: (2025)
Combating the Elsagate phenomenon: Deep learning architectures for disturbing cartoons
par: Ishikawa, Akari, et autres
Publié: (2019)
par: Ishikawa, Akari, et autres
Publié: (2019)
C-MORAL: Controllable Multi-Objective Molecular Optimization with Reinforcement Alignment for LLMs
par: Gao, Rui, et autres
Publié: (2026)
par: Gao, Rui, et autres
Publié: (2026)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
par: Xu, Sirui, et autres
Publié: (2026)
par: Xu, Sirui, et autres
Publié: (2026)
Auto Research with Specialist Agents Develops Effective and Non-Trivial Training Recipes
par: Ning, Jingjie, et autres
Publié: (2026)
par: Ning, Jingjie, et autres
Publié: (2026)
Spermidine regulates wheat grain weight at high planting density by promoting the synthesis of sucrose and starch in inferior grains
par: Juan Li, et autres
Publié: (2024)
par: Juan Li, et autres
Publié: (2024)
Out of Style: RAG's Fragility to Linguistic Variation
par: Cao, Tianyu, et autres
Publié: (2025)
par: Cao, Tianyu, et autres
Publié: (2025)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
par: Chen, Tong, et autres
Publié: (2025)
par: Chen, Tong, et autres
Publié: (2025)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
par: Yu, Zichun, et autres
Publié: (2025)
par: Yu, Zichun, et autres
Publié: (2025)
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
par: Yu, Zichun, et autres
Publié: (2026)
par: Yu, Zichun, et autres
Publié: (2026)
A High-Fidelity Surrogate Framework for Rapid Solvation of Gas-Phase and Interfacial Reaction Potential Energy Surfaces
par: Junwei Lucas, Bao, et autres
Publié: (2026)
par: Junwei Lucas, Bao, et autres
Publié: (2026)
Certifying Counterfactual Bias in LLMs
par: Chaudhary, Isha, et autres
Publié: (2024)
par: Chaudhary, Isha, et autres
Publié: (2024)
Deep Learning-based Approaches for State Space Models: A Selective Review
par: Lin, Jiahe, et autres
Publié: (2024)
par: Lin, Jiahe, et autres
Publié: (2024)
ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents
par: Guo, Zhihan, et autres
Publié: (2025)
par: Guo, Zhihan, et autres
Publié: (2025)
Fine-grained deterministic hardness of the shortest vector problem
par: Hittmeir, Markus
Publié: (2025)
par: Hittmeir, Markus
Publié: (2025)
Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward
par: Deng, Yong, et autres
Publié: (2025)
par: Deng, Yong, et autres
Publié: (2025)
Diagnosis of Fuel Cell Health Status with Deep Sparse Auto-Encoder Neural Network
par: Fei, Chenyan, et autres
Publié: (2025)
par: Fei, Chenyan, et autres
Publié: (2025)
Guidelines for Fine-grained Sentence-level Arabic Readability Annotation
par: Habash, Nizar, et autres
Publié: (2024)
par: Habash, Nizar, et autres
Publié: (2024)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
par: Yin, Zhihan, et autres
Publié: (2026)
par: Yin, Zhihan, et autres
Publié: (2026)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
par: Du, Mingxuan, et autres
Publié: (2025)
par: Du, Mingxuan, et autres
Publié: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
par: Muhamed, Aashiq, et autres
Publié: (2025)
par: Muhamed, Aashiq, et autres
Publié: (2025)
NeoQA: Evidence-based Question Answering with Generated News Events
par: Glockner, Max, et autres
Publié: (2025)
par: Glockner, Max, et autres
Publié: (2025)
Interpretable Fine‐Grained Phenotypes of Subcellular Dynamics via Unsupervised Deep Learning
par: Chuangqi Wang, et autres
Publié: (2024)
par: Chuangqi Wang, et autres
Publié: (2024)
Deep Learning for Crime Forecasting: The Role of Mobility at Fine-grained Spatiotemporal Scales
par: Zumel, Ariadna Albors, et autres
Publié: (2025)
par: Zumel, Ariadna Albors, et autres
Publié: (2025)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
par: Zhu, Fengbin, et autres
Publié: (2025)
par: Zhu, Fengbin, et autres
Publié: (2025)
Documents similaires
-
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
par: Wang, Tevin, et autres
Publié: (2025) -
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
par: Wang, Tevin, et autres
Publié: (2024) -
Interpret and Control Dense Retrieval with Sparse Latent Features
par: Kang, Hao, et autres
Publié: (2024) -
DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
par: Huang, Yukun, et autres
Publié: (2026) -
AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
par: Chen, Zijian, et autres
Publié: (2026)