Gespeichert in:
| Hauptverfasser: | Wang, Shuting, Tan, Jiejun, Dou, Zhicheng, Wen, Ji-Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.13018 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
von: Zhang, Yiman, et al.
Veröffentlicht: (2025)
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
von: Tan, Jiejun, et al.
Veröffentlicht: (2025)
von: Tan, Jiejun, et al.
Veröffentlicht: (2025)
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)
Memory Matters More: Event-Centric Memory as a Logic Map for Agent Searching and Reasoning
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
von: Hu, Yuyang, et al.
Veröffentlicht: (2026)
RichRAG: Crafting Rich Responses for Multi-faceted Queries in Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
OmniGAIA: Towards Native Omni-Modal AI Agents
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
von: Li, Xiaoxi, et al.
Veröffentlicht: (2026)
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models
von: Guo, Xin, et al.
Veröffentlicht: (2023)
von: Guo, Xin, et al.
Veröffentlicht: (2023)
One Token Can Help! Learning Scalable and Pluggable Virtual Tokens for Retrieval-Augmented Large Language Models
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
von: Zhu, Yutao, et al.
Veröffentlicht: (2024)
FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning
von: Dou, Shaoyu, et al.
Veröffentlicht: (2025)
von: Dou, Shaoyu, et al.
Veröffentlicht: (2025)
Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
von: Wang, Shuting, et al.
Veröffentlicht: (2025)
QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models
von: Kang, Zhaolu, et al.
Veröffentlicht: (2026)
von: Kang, Zhaolu, et al.
Veröffentlicht: (2026)
Query-oriented Data Augmentation for Session Search
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
von: Chen, Haonan, et al.
Veröffentlicht: (2024)
UFO: a Unified and Flexible Framework for Evaluating Factuality of Large Language Models
von: Huang, Zhaoheng, et al.
Veröffentlicht: (2024)
von: Huang, Zhaoheng, et al.
Veröffentlicht: (2024)
ProRAG: Process-Supervised Reinforcement Learning for Retrieval-Augmented Generation
von: Wang, Zhao, et al.
Veröffentlicht: (2026)
von: Wang, Zhao, et al.
Veröffentlicht: (2026)
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
von: Jin, Jiajie, et al.
Veröffentlicht: (2024)
Large Language Models for Information Retrieval: A Survey
von: Zhu, Yutao, et al.
Veröffentlicht: (2023)
von: Zhu, Yutao, et al.
Veröffentlicht: (2023)
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
AssistRAG: Boosting the Potential of Large Language Models with an Intelligent Information Assistant
von: Zhou, Yujia, et al.
Veröffentlicht: (2024)
von: Zhou, Yujia, et al.
Veröffentlicht: (2024)
SMARTFinRAG: Interactive Modularized Financial RAG Benchmark
von: Zha, Yiwei
Veröffentlicht: (2025)
von: Zha, Yiwei
Veröffentlicht: (2025)
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
von: Cheng, Yiruo, et al.
Veröffentlicht: (2024)
von: Cheng, Yiruo, et al.
Veröffentlicht: (2024)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
FinS-Pilot: A Benchmark for Online Financial RAG System
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
AutoMedEval: Harnessing Language Models for Automatic Medical Capability Evaluation
von: Zhang, Xiechi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiechi, et al.
Veröffentlicht: (2025)
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2026)
HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants
von: Gritta, Milan, et al.
Veröffentlicht: (2024)
von: Gritta, Milan, et al.
Veröffentlicht: (2024)
From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
von: Deng, Zhirui, et al.
Veröffentlicht: (2024)
von: Deng, Zhirui, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2024)
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
von: Wang, Wanying, et al.
Veröffentlicht: (2024)
OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models
von: Kim, Seunghee, et al.
Veröffentlicht: (2026)
von: Kim, Seunghee, et al.
Veröffentlicht: (2026)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
von: Guo, Beichen, et al.
Veröffentlicht: (2025)
von: Guo, Beichen, et al.
Veröffentlicht: (2025)
HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks
von: Qian, Hongjin, et al.
Veröffentlicht: (2025)
von: Qian, Hongjin, et al.
Veröffentlicht: (2025)
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
von: Kong, Zicheng, et al.
Veröffentlicht: (2026)
von: Kong, Zicheng, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024) -
Small Models, Big Insights: Leveraging Slim Proxy Models To Decide When and What to Retrieve for LLMs
von: Tan, Jiejun, et al.
Veröffentlicht: (2024) -
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
von: Zhang, Yiman, et al.
Veröffentlicht: (2025) -
HierSearch: A Hierarchical Enterprise Deep Search Framework Integrating Local and Web Searches
von: Tan, Jiejun, et al.
Veröffentlicht: (2025) -
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
von: Tan, Jiejun, et al.
Veröffentlicht: (2024)