DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Janghoon, Kim, Heegyu, Lee, Changho, Lee, Dahm, Park, Min Hyung, Song, Hosung, Choi, Stanley Jungkyu, Lee, Moontae, Lee, Honglak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024)
by: Han, Janghoon, et al.
Published: (2024)
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Spanning Tree Autoregressive Visual Generation
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024)
by: Lee, Changho, et al.
Published: (2024)
HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images
by: Choi, Sungik, et al.
Published: (2024)
by: Choi, Sungik, et al.
Published: (2024)
LGAI-EMBEDDING-Preview Technical Report
by: Choi, Jooyoung, et al.
Published: (2025)
by: Choi, Jooyoung, et al.
Published: (2025)
Mitigating Biases for Instruction-following Language Models via Bias Neurons Elimination
by: Yang, Nakyeong, et al.
Published: (2023)
by: Yang, Nakyeong, et al.
Published: (2023)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
by: Khalifa, Muhammad, et al.
Published: (2023)
by: Khalifa, Muhammad, et al.
Published: (2023)
Towards Diverse Evaluation of Class Incremental Learning: A Representation Learning Perspective
by: Cha, Sungmin, et al.
Published: (2022)
by: Cha, Sungmin, et al.
Published: (2022)
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
by: Khalifa, Muhammad, et al.
Published: (2026)
by: Khalifa, Muhammad, et al.
Published: (2026)
Training-free Detection of AI-generated images via Cropping Robustness
by: Choi, Sungik, et al.
Published: (2025)
by: Choi, Sungik, et al.
Published: (2025)
MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
by: Zhang, Yunxiang, et al.
Published: (2025)
by: Zhang, Yunxiang, et al.
Published: (2025)
To Predict or Not To Predict? Proportionally Masked Autoencoders for Tabular Data Imputation
by: Kim, Jungkyu, et al.
Published: (2024)
by: Kim, Jungkyu, et al.
Published: (2024)
FLEX: Expert-level False-Less EXecution Metric for Reliable Text-to-SQL Benchmark
by: Kim, Heegyu, et al.
Published: (2024)
by: Kim, Heegyu, et al.
Published: (2024)
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
by: Shen, Siqi, et al.
Published: (2025)
by: Shen, Siqi, et al.
Published: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
by: Shen, Siqi, et al.
Published: (2024)
by: Shen, Siqi, et al.
Published: (2024)
Learning to Unlearn: Instance-wise Unlearning for Pre-trained Classifiers
by: Cha, Sungmin, et al.
Published: (2023)
by: Cha, Sungmin, et al.
Published: (2023)
Small Language Models Need Strong Verifiers to Self-Correct Reasoning
by: Zhang, Yunxiang, et al.
Published: (2024)
by: Zhang, Yunxiang, et al.
Published: (2024)
When Is Enough Not Enough? Illusory Completion in Search Agents
by: Ko, Dayoon, et al.
Published: (2026)
by: Ko, Dayoon, et al.
Published: (2026)
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
by: Kim, Geon-Hyeong, et al.
Published: (2025)
by: Kim, Geon-Hyeong, et al.
Published: (2025)
B4DL: A Benchmark for 4D LiDAR LLM in Spatio-Temporal Understanding
by: Choi, Changho, et al.
Published: (2025)
by: Choi, Changho, et al.
Published: (2025)
Partial-Multivariate Model for Forecasting
by: Lee, Jaehoon, et al.
Published: (2024)
by: Lee, Jaehoon, et al.
Published: (2024)
Shifting from Ranking to Set Selection for Retrieval Augmented Generation
by: Lee, Dahyun, et al.
Published: (2025)
by: Lee, Dahyun, et al.
Published: (2025)
Hybrid Deep Searcher: Scalable Parallel and Sequential Search Reasoning
by: Ko, Dayoon, et al.
Published: (2025)
by: Ko, Dayoon, et al.
Published: (2025)
Process Reward Models That Think
by: Khalifa, Muhammad, et al.
Published: (2025)
by: Khalifa, Muhammad, et al.
Published: (2025)
Enhancing Mixture-of-Experts Specialization via Cluster-Aware Upcycling
by: Chu, Sanghyeok, et al.
Published: (2026)
by: Chu, Sanghyeok, et al.
Published: (2026)
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
by: Kim, Sejin, et al.
Published: (2024)
by: Kim, Sejin, et al.
Published: (2024)
Magnetic-field dependent VB- spin decoherence in hexagonal boron nitrides: A first-principles study
by: Lee, Jaewook, et al.
Published: (2025)
by: Lee, Jaewook, et al.
Published: (2025)
Magnetic‐Field Dependent V B − Spin Decoherence in Hexagonal Boron Nitrides: A First‐Principles Study
by: Jaewook Lee, et al.
Published: (2025)
by: Jaewook Lee, et al.
Published: (2025)
Scaling Web Agent Training through Automatic Data Generation and Fine-grained Evaluation
by: Logeswaran, Lajanugen, et al.
Published: (2026)
by: Logeswaran, Lajanugen, et al.
Published: (2026)
Counterfactual Voting Adjustment for Quality Assessment and Fairer Voting in Online Platforms with Helpfulness Evaluation
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
Learning to Generate Unit Test via Adversarial Reinforcement Learning
by: Lee, Dongjun, et al.
Published: (2025)
by: Lee, Dongjun, et al.
Published: (2025)
EXAONE Deep: Reasoning Enhanced Language Models
by: Bae, Kyunghoon, et al.
Published: (2025)
by: Bae, Kyunghoon, et al.
Published: (2025)
Coding-Free and Privacy-Preserving Agentic Framework for Data-Driven Clinical Research
by: Kim, Taehun, et al.
Published: (2026)
by: Kim, Taehun, et al.
Published: (2026)
Bigness of the tangent bundle of a Fano threefold with Picard number two
by: Kim, Hosung, et al.
Published: (2022)
by: Kim, Hosung, et al.
Published: (2022)
Positivity of the tangent bundle of rational surfaces with nef anticanonical divisor
by: Kim, Hosung, et al.
Published: (2024)
by: Kim, Hosung, et al.
Published: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
by: Lee, Kyungjae, et al.
Published: (2024)
by: Lee, Kyungjae, et al.
Published: (2024)
KVoiceBench, KOpenAudioBench, and KMMAU: Agent-Driven Korean Speech Benchmarks for Evaluating SpeechLMs
by: Kim, Haechan, et al.
Published: (2026)
by: Kim, Haechan, et al.
Published: (2026)
From Documents to Segments: A Contextual Reformulation for Topic Assignment
by: Yoon, Hoonsang, et al.
Published: (2026)
by: Yoon, Hoonsang, et al.
Published: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
by: Kim, Soyeon, et al.
Published: (2026)
by: Kim, Soyeon, et al.
Published: (2026)
Similar Items
-
Deep Exploration of Cross-Lingual Zero-Shot Generalization in Instruction Tuning
by: Han, Janghoon, et al.
Published: (2024) -
KL Penalty Control via Perturbation for Direct Preference Optimization
by: Lee, Sangkyu, et al.
Published: (2025) -
Spanning Tree Autoregressive Visual Generation
by: Lee, Sangkyu, et al.
Published: (2025) -
Instruction Matters: A Simple yet Effective Task Selection for Optimized Instruction Tuning of Specific Tasks
by: Lee, Changho, et al.
Published: (2024) -
HFI: A unified framework for training-free detection and implicit watermarking of latent diffusion model generated images
by: Choi, Sungik, et al.
Published: (2024)