From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Shuoling, Tan, Zhiquan, Yi, Kun, Wu, Hui, Li, Yihan, Yan, Jiangpeng, Chen, Liyuan, Chen, Kai, Yang, Qiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News
von: Chen, Liyuan, et al.
Veröffentlicht: (2026)
von: Chen, Liyuan, et al.
Veröffentlicht: (2026)
The Label Horizon Paradox: Rethinking Supervision Targets in Financial Forecasting
von: Song, Chen-Hui, et al.
Veröffentlicht: (2026)
von: Song, Chen-Hui, et al.
Veröffentlicht: (2026)
Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment
von: Liu, Shuoling, et al.
Veröffentlicht: (2024)
von: Liu, Shuoling, et al.
Veröffentlicht: (2024)
Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges
von: Chen, Liyuan, et al.
Veröffentlicht: (2025)
von: Chen, Liyuan, et al.
Veröffentlicht: (2025)
Federated Co-tuning Framework for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)
von: Fan, Tao, et al.
Veröffentlicht: (2024)
IntentScore: Intent-Conditioned Action Evaluation for Computer-Use Agents
von: Chen, Rongqian, et al.
Veröffentlicht: (2026)
von: Chen, Rongqian, et al.
Veröffentlicht: (2026)
From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
Argus: Evidence Assembly for Scalable Deep Research Agents
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
von: Hong, Yinrong, et al.
Veröffentlicht: (2025)
Multimodal Fusion of Skeleton Dynamics and Clinical Gait Features for Video-Based Cerebral Palsy Severity Assessment
von: Yang, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Yang, Kaiyuan, et al.
Veröffentlicht: (2026)
LLM-Powered Intent-Based Categorization of Phishing Emails
von: Eilertsen, Even, et al.
Veröffentlicht: (2025)
von: Eilertsen, Even, et al.
Veröffentlicht: (2025)
Categorical Continuous Symmetry
von: Jia, Qiang, et al.
Veröffentlicht: (2025)
von: Jia, Qiang, et al.
Veröffentlicht: (2025)
From Shallow to Deep: Pinning Semantic Intent via Causal GRPO
von: Zhou, Shuyi, et al.
Veröffentlicht: (2026)
von: Zhou, Shuyi, et al.
Veröffentlicht: (2026)
SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation
von: Chen, Zhiquan, et al.
Veröffentlicht: (2026)
von: Chen, Zhiquan, et al.
Veröffentlicht: (2026)
ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research
von: Xia, Yihan, et al.
Veröffentlicht: (2026)
von: Xia, Yihan, et al.
Veröffentlicht: (2026)
Statistical Mechanics and Categorical Entropy
von: Wu, Haiqi, et al.
Veröffentlicht: (2025)
von: Wu, Haiqi, et al.
Veröffentlicht: (2025)
DeepThink: Aligning Language Models with Domain-Specific User Intents
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Know Your Intent: An Autonomous Multi-Perspective LLM Agent Framework for DeFi User Transaction Intent Mining
von: Mao, Qian'ang, et al.
Veröffentlicht: (2025)
von: Mao, Qian'ang, et al.
Veröffentlicht: (2025)
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
von: Yao, Yang, et al.
Veröffentlicht: (2025)
von: Yao, Yang, et al.
Veröffentlicht: (2025)
Categorical Symmetries via Operator Algebras
von: Jia, Qiang, et al.
Veröffentlicht: (2026)
von: Jia, Qiang, et al.
Veröffentlicht: (2026)
DeepPsy-Agent: A Stage-Aware and Deep-Thinking Emotional Support Agent System
von: Chen, Kai, et al.
Veröffentlicht: (2025)
von: Chen, Kai, et al.
Veröffentlicht: (2025)
A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
von: Tan, Zhiquan, et al.
Veröffentlicht: (2025)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2025)
PAINT: Partial-Solution Adaptive Interpolated Training for Self-Distilled Reasoners
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
Understanding Grokking Through A Robustness Viewpoint
von: Tan, Zhiquan, et al.
Veröffentlicht: (2023)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2023)
Self-Supervised On-Policy Distillation for Reasoning Language Models
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2026)
Information-Theoretic Perspectives on Optimizers
von: Tan, Zhiquan, et al.
Veröffentlicht: (2025)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2025)
Scientific Algorithm Discovery by Augmenting AlphaEvolve with Deep Research
von: Liu, Gang, et al.
Veröffentlicht: (2025)
von: Liu, Gang, et al.
Veröffentlicht: (2025)
Deep Learning Approaches for Multimodal Intent Recognition: A Survey
von: Zhao, Jingwei, et al.
Veröffentlicht: (2025)
von: Zhao, Jingwei, et al.
Veröffentlicht: (2025)
Exploring Information-Theoretic Metrics Associated with Neural Collapse in Supervised Training
von: Song, Kun, et al.
Veröffentlicht: (2024)
von: Song, Kun, et al.
Veröffentlicht: (2024)
Exploring the Interplay of Motivation, Engagement and Critical Thinking Among EFL Learners: Evidence From Structural Equation Modelling
von: Jijia Yang, et al.
Veröffentlicht: (2025)
von: Jijia Yang, et al.
Veröffentlicht: (2025)
You Only Anonymize What Is Not Intent-Relevant: Suppressing Non-Intent Privacy Evidence
von: Shen, Weihao, et al.
Veröffentlicht: (2026)
von: Shen, Weihao, et al.
Veröffentlicht: (2026)
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2025)
Enhancing Zeroth-order Fine-tuning for Language Models with Low-rank Structures
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
Evaluating Deep Clustering Algorithms on Non-Categorical 3D CAD Models
von: Xiang, Siyuan, et al.
Veröffentlicht: (2024)
von: Xiang, Siyuan, et al.
Veröffentlicht: (2024)
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
von: Lian, Shijie, et al.
Veröffentlicht: (2026)
von: Lian, Shijie, et al.
Veröffentlicht: (2026)
From Evidence to Impact: Bridging the Implementation Gap in Geriatric Deprescribing
von: Songhe Chen, et al.
Veröffentlicht: (2025)
von: Songhe Chen, et al.
Veröffentlicht: (2025)
From Spots to Pixels: Dense Spatial Gene Expression Prediction from Histology Images
von: Zhang, Ruikun, et al.
Veröffentlicht: (2025)
von: Zhang, Ruikun, et al.
Veröffentlicht: (2025)
Dimension-Level Intent Fidelity Evaluation for Large Language Models: Evidence from Structured Prompt Ablation
von: Peng, GAng
Veröffentlicht: (2026)
von: Peng, GAng
Veröffentlicht: (2026)
TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents
von: Chen, Yanyu, et al.
Veröffentlicht: (2026)
von: Chen, Yanyu, et al.
Veröffentlicht: (2026)
DeepShop: A Benchmark for Deep Research Shopping Agents
von: Lyu, Yougang, et al.
Veröffentlicht: (2025)
von: Lyu, Yougang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News
von: Chen, Liyuan, et al.
Veröffentlicht: (2026) -
The Label Horizon Paradox: Rethinking Supervision Targets in Financial Forecasting
von: Song, Chen-Hui, et al.
Veröffentlicht: (2026) -
Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment
von: Liu, Shuoling, et al.
Veröffentlicht: (2024) -
Advancing Financial Engineering with Foundation Models: Progress, Applications, and Challenges
von: Chen, Liyuan, et al.
Veröffentlicht: (2025) -
Federated Co-tuning Framework for Large and Small Language Models
von: Fan, Tao, et al.
Veröffentlicht: (2024)