SciNet: Evaluating AI Agents in Relation-Aware Scientific Literature Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shao, Chenyang, Xu, Fengli, Li, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
von: Shao, Chenyang, et al.
Veröffentlicht: (2026)
von: Shao, Chenyang, et al.
Veröffentlicht: (2026)
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
Mixture of Knowledge Minigraph Agents for Literature Review Generation
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
von: Zhang, Zhi, et al.
Veröffentlicht: (2024)
A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs
von: Zhu, Andy, et al.
Veröffentlicht: (2025)
von: Zhu, Andy, et al.
Veröffentlicht: (2025)
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
von: Wu, Weimin, et al.
Veröffentlicht: (2025)
FinGEAR: Financial Mapping-Guided Enhanced Answer Retrieval
von: Li, Ying, et al.
Veröffentlicht: (2025)
von: Li, Ying, et al.
Veröffentlicht: (2025)
MatSciRE: Leveraging Pointer Networks to Automate Entity and Relation Extraction for Material Science Knowledge-base Construction
von: Mullick, Ankan, et al.
Veröffentlicht: (2024)
von: Mullick, Ankan, et al.
Veröffentlicht: (2024)
SciNav: A General Agent Framework for Scientific Coding Tasks
von: Zhang, Tianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Tianshu, et al.
Veröffentlicht: (2026)
Golden Touchstone: A Comprehensive Bilingual Benchmark for Evaluating Financial Large Language Models
von: Wu, Xiaojun, et al.
Veröffentlicht: (2024)
von: Wu, Xiaojun, et al.
Veröffentlicht: (2024)
All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection
von: Jiang, Yuechen, et al.
Veröffentlicht: (2026)
von: Jiang, Yuechen, et al.
Veröffentlicht: (2026)
DIALECTIC: A Multi-Agent System for Startup Evaluation
von: Bae, Jae Yoon, et al.
Veröffentlicht: (2026)
von: Bae, Jae Yoon, et al.
Veröffentlicht: (2026)
AI PB: A Grounded Generative Agent for Personalized Investment Insights
von: Park, Daewoo, et al.
Veröffentlicht: (2025)
von: Park, Daewoo, et al.
Veröffentlicht: (2025)
EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
von: Shi, Yao, et al.
Veröffentlicht: (2025)
von: Shi, Yao, et al.
Veröffentlicht: (2025)
SenseAI: A Human-in-the-Loop Dataset for RLHF-Aligned Financial Sentiment Reasoning
von: Kabalisa, Berny
Veröffentlicht: (2026)
von: Kabalisa, Berny
Veröffentlicht: (2026)
ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
von: Liu, Yujie, et al.
Veröffentlicht: (2025)
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
von: Huang, Jimin, et al.
Veröffentlicht: (2024)
Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation
von: Walia, Jaskaran Singh, et al.
Veröffentlicht: (2025)
von: Walia, Jaskaran Singh, et al.
Veröffentlicht: (2025)
Aethorix v1.0: An Integrated Scientific AI Agent for Scalable Inorganic Materials Innovation and Industrial Implementation
von: Shi, Yingjie, et al.
Veröffentlicht: (2025)
von: Shi, Yingjie, et al.
Veröffentlicht: (2025)
The CLEF-2026 FinMMEval Lab: Multilingual and Multimodal Evaluation of Financial AI Systems
von: Xie, Zhuohan, et al.
Veröffentlicht: (2026)
von: Xie, Zhuohan, et al.
Veröffentlicht: (2026)
UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
von: Yang, Yuzhe, et al.
Veröffentlicht: (2024)
MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrier
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
von: Yang, Zonglin, et al.
Veröffentlicht: (2026)
PolyGnosis 2.0: Enhancing LLM Reasoning via Agentic Harness Engineering for Polymarket and OSINT Insight Extraction
von: Wang, Daren, et al.
Veröffentlicht: (2026)
von: Wang, Daren, et al.
Veröffentlicht: (2026)
NumLLM: Numeric-Sensitive Large Language Model for Chinese Finance
von: Su, Huan-Yi, et al.
Veröffentlicht: (2024)
von: Su, Huan-Yi, et al.
Veröffentlicht: (2024)
RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements
von: Zhan, Zaifu, et al.
Veröffentlicht: (2024)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2024)
AEL: Agent Evolving Learning for Open-Ended Environments
von: Xu, Wujiang, et al.
Veröffentlicht: (2026)
von: Xu, Wujiang, et al.
Veröffentlicht: (2026)
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
von: Fan, Tianyu, et al.
Veröffentlicht: (2025)
von: Fan, Tianyu, et al.
Veröffentlicht: (2025)
Automating MD simulations for Proteins using Large language Models: NAMD-Agent
von: Chandrasekhar, Achuth, et al.
Veröffentlicht: (2025)
von: Chandrasekhar, Achuth, et al.
Veröffentlicht: (2025)
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
von: Ye, Yuxuan, et al.
Veröffentlicht: (2026)
von: Ye, Yuxuan, et al.
Veröffentlicht: (2026)
SciNets: Graph-Constrained Multi-Hop Reasoning for Scientific Literature Synthesis
von: Dubey, Sauhard
Veröffentlicht: (2025)
von: Dubey, Sauhard
Veröffentlicht: (2025)
BeamformNet: Deep Learning-Based Beamforming Method for DoA Estimation via Implicit Spatial Signal Focusing and Noise Suppression
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
von: Deng, Xuyao, et al.
Veröffentlicht: (2025)
GrifFinNet: A Graph-Relation Integrated Transformer for Financial Predictions
von: Dai, Chenlanhui, et al.
Veröffentlicht: (2025)
von: Dai, Chenlanhui, et al.
Veröffentlicht: (2025)
FinBoardBench: Benchmarking Dynamic Wealth Management and Strategic Financial Reasoning of LLMs via Board Game Simulations
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
von: Hu, Xuesi, et al.
Veröffentlicht: (2026)
Cross-Asset Risk Management: Integrating LLMs for Real-Time Monitoring of Equity, Fixed Income, and Currency Markets
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
Dynamic Hedging Strategies in Derivatives Markets with LLM-Driven Sentiment and News Analytics
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
The Hype Index: an NLP-driven Measure of Market News Attention
von: Cao, Zheng, et al.
Veröffentlicht: (2025)
von: Cao, Zheng, et al.
Veröffentlicht: (2025)
BERTopic-Driven Stock Market Predictions: Unraveling Sentiment Insights
von: Zhu, Enmin, et al.
Veröffentlicht: (2024)
von: Zhu, Enmin, et al.
Veröffentlicht: (2024)
SilverSight: A Multi-Task Chinese Financial Large Language Model Based on Adaptive Semantic Space Learning
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2024)
Program of Thoughts for Financial Reasoning: Leveraging Dynamic In-Context Examples and Generative Retrieval
von: Khatuya, Subhendu, et al.
Veröffentlicht: (2025)
von: Khatuya, Subhendu, et al.
Veröffentlicht: (2025)
Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation
von: Oestreich, Julian, et al.
Veröffentlicht: (2026)
von: Oestreich, Julian, et al.
Veröffentlicht: (2026)
Reverse Physician-AI Relationship: Full-process Clinical Diagnosis Driven by a Large Language Model
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
von: Xu, Shicheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
von: Shao, Chenyang, et al.
Veröffentlicht: (2026) -
OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists
von: Shao, Chenyang, et al.
Veröffentlicht: (2025) -
Mixture of Knowledge Minigraph Agents for Literature Review Generation
von: Zhang, Zhi, et al.
Veröffentlicht: (2024) -
A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs
von: Zhu, Andy, et al.
Veröffentlicht: (2025) -
Sci2Pol: Evaluating and Fine-tuning LLMs on Scientific-to-Policy Brief Generation
von: Wu, Weimin, et al.
Veröffentlicht: (2025)