LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhou, Heng, Yu, Ao, Fan, Yuchen, Shi, Jianing, Kang, Li, Geng, Hejia, Zhang, Yongting, Fan, Yutao, Wu, Yuhao, He, Tiancheng, Qin, Yiran, Bai, Lei, Yin, Zhenfei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
par: Zhou, Heng, et autres
Publié: (2026)
par: Zhou, Heng, et autres
Publié: (2026)
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
par: Zhou, Heng, et autres
Publié: (2025)
par: Zhou, Heng, et autres
Publié: (2025)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
par: Tan, Zelin, et autres
Publié: (2025)
par: Tan, Zelin, et autres
Publié: (2025)
LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval
par: ai, Gensmo., et autres
Publié: (2026)
par: ai, Gensmo., et autres
Publié: (2026)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
par: Zhou, Heng, et autres
Publié: (2026)
par: Zhou, Heng, et autres
Publié: (2026)
EffiBench: Benchmarking the Efficiency of Automatically Generated Code
par: Huang, Dong, et autres
Publié: (2024)
par: Huang, Dong, et autres
Publié: (2024)
HoSNN: Adversarially-Robust Homeostatic Spiking Neural Networks with Adaptive Firing Thresholds
par: Geng, Hejia, et autres
Publié: (2023)
par: Geng, Hejia, et autres
Publié: (2023)
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
par: Wu, Xiaobao, et autres
Publié: (2024)
par: Wu, Xiaobao, et autres
Publié: (2024)
AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
par: Zha, Jirong, et autres
Publié: (2025)
par: Zha, Jirong, et autres
Publié: (2025)
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
par: He, Linyang, et autres
Publié: (2026)
par: He, Linyang, et autres
Publié: (2026)
SGR-Bench: Benchmarking Search Agents on State-Gated Retrieval
par: Li, Ningyuan, et autres
Publié: (2026)
par: Li, Ningyuan, et autres
Publié: (2026)
Agentic-R: Learning to Retrieve for Agentic Search
par: Liu, Wenhan, et autres
Publié: (2026)
par: Liu, Wenhan, et autres
Publié: (2026)
Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents
par: Zhou, Heng, et autres
Publié: (2026)
par: Zhou, Heng, et autres
Publié: (2026)
DeonticBench: A Benchmark for Reasoning over Rules
par: Dou, Guangyao, et autres
Publié: (2026)
par: Dou, Guangyao, et autres
Publié: (2026)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
par: Thakur, Nandan, et autres
Publié: (2024)
par: Thakur, Nandan, et autres
Publié: (2024)
TRACE: Timely Retrieval and Alignment for Cybersecurity Knowledge Graph Construction and Expansion
par: Xu, Zijing, et autres
Publié: (2026)
par: Xu, Zijing, et autres
Publié: (2026)
Unlocking compositional design and mechanism for mechanical/thermal properties of high‐entropy rare‐earth disilicates
par: Yun Fan, et autres
Publié: (2025)
par: Yun Fan, et autres
Publié: (2025)
Retrieval-Infused Reasoning Sandbox: A Benchmark for Decoupling Retrieval and Reasoning Capabilities
par: Ying, Shuangshuang, et autres
Publié: (2026)
par: Ying, Shuangshuang, et autres
Publié: (2026)
Composing Recurrent Spiking Neural Networks using Locally-Recurrent Motifs and Risk-Mitigating Architectural Optimization
par: Zhang, Wenrui, et autres
Publié: (2021)
par: Zhang, Wenrui, et autres
Publié: (2021)
Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
par: Shen, Yiqing, et autres
Publié: (2025)
par: Shen, Yiqing, et autres
Publié: (2025)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
par: Zhang, Zhengbo, et autres
Publié: (2026)
par: Zhang, Zhengbo, et autres
Publié: (2026)
Deja Vu in Plots: Leveraging Cross-Session Evidence with Retrieval-Augmented LLMs for Live Streaming Risk Assessment
par: Qiao, Yiran, et autres
Publié: (2026)
par: Qiao, Yiran, et autres
Publié: (2026)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
par: Jing, Huihao, et autres
Publié: (2025)
par: Jing, Huihao, et autres
Publié: (2025)
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
par: Tan, Zelin, et autres
Publié: (2026)
par: Tan, Zelin, et autres
Publié: (2026)
Knowledge Pyramid Construction for Multi-Level Retrieval-Augmented Generation
par: Chen, Rubing, et autres
Publié: (2024)
par: Chen, Rubing, et autres
Publié: (2024)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
par: Zhao, Enyu, et autres
Publié: (2025)
par: Zhao, Enyu, et autres
Publié: (2025)
In situ impregnation of cobalt‐doped tungstophosphoric acid on MOF‐801 toward enhanced catalytic activity for esterification
par: Qiuyun Zhang, et autres
Publié: (2024)
par: Qiuyun Zhang, et autres
Publié: (2024)
Search-on-Graph: Iterative Informed Navigation for Large Language Model Reasoning on Knowledge Graphs
par: Sun, Jia Ao, et autres
Publié: (2025)
par: Sun, Jia Ao, et autres
Publié: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
par: Gao, Zihan, et autres
Publié: (2025)
par: Gao, Zihan, et autres
Publié: (2025)
KOR-Bench: Benchmarking Language Models on Knowledge-Orthogonal Reasoning Tasks
par: Ma, Kaijing, et autres
Publié: (2024)
par: Ma, Kaijing, et autres
Publié: (2024)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
par: Wang, Yuxia, et autres
Publié: (2023)
par: Wang, Yuxia, et autres
Publié: (2023)
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
par: Xue, Siqiao, et autres
Publié: (2026)
par: Xue, Siqiao, et autres
Publié: (2026)
WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives
par: Yu, Yongan, et autres
Publié: (2025)
par: Yu, Yongan, et autres
Publié: (2025)
AL-Bench: A Benchmark for Automatic Logging
par: Tan, Boyin, et autres
Publié: (2025)
par: Tan, Boyin, et autres
Publié: (2025)
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work
par: Shen, Haiyang, et autres
Publié: (2026)
par: Shen, Haiyang, et autres
Publié: (2026)
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
par: Chou, Jason, et autres
Publié: (2025)
par: Chou, Jason, et autres
Publié: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
par: Kang, Li, et autres
Publié: (2026)
par: Kang, Li, et autres
Publié: (2026)
Assessment of Multimodal Large Language Models in Alignment with Human Values
par: Shi, Zhelun, et autres
Publié: (2024)
par: Shi, Zhelun, et autres
Publié: (2024)
TongSearch-QR: Reinforced Query Reasoning for Retrieval
par: Qin, Xubo, et autres
Publié: (2025)
par: Qin, Xubo, et autres
Publié: (2025)
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
par: Xi, Zhiheng, et autres
Publié: (2025)
par: Xi, Zhiheng, et autres
Publié: (2025)
Documents similaires
-
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
par: Zhou, Heng, et autres
Publié: (2026) -
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
par: Zhou, Heng, et autres
Publié: (2025) -
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
par: Tan, Zelin, et autres
Publié: (2025) -
LookBench: A Live and Holistic Open Benchmark for Fashion Image Retrieval
par: ai, Gensmo., et autres
Publié: (2026) -
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
par: Zhou, Heng, et autres
Publié: (2026)