LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yunfan, McKeown, Kathleen, Muresan, Smaranda |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
Forecasting Conversation Derailments Through Generation
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025)
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
von: Zhang, Weizhi, et al.
Veröffentlicht: (2025)
von: Zhang, Weizhi, et al.
Veröffentlicht: (2025)
FlashEvaluator: Expanding Search Space with Parallel Evaluation
von: Feng, Chao, et al.
Veröffentlicht: (2026)
von: Feng, Chao, et al.
Veröffentlicht: (2026)
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)
Search, Examine and Early-Termination: Fake News Detection with Annotation-Free Evidences
von: Yang, Yuzhou, et al.
Veröffentlicht: (2024)
von: Yang, Yuzhou, et al.
Veröffentlicht: (2024)
Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements
von: Terragni, Silvia, et al.
Veröffentlicht: (2024)
von: Terragni, Silvia, et al.
Veröffentlicht: (2024)
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2026)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2026)
RELIANCE: Reliable Ensemble Learning for Information and News Credibility Evaluation
von: Ramezani, Majid, et al.
Veröffentlicht: (2024)
von: Ramezani, Majid, et al.
Veröffentlicht: (2024)
Fake News Detection After LLM Laundering: Measurement and Explanation
von: Das, Rupak Kumar, et al.
Veröffentlicht: (2025)
von: Das, Rupak Kumar, et al.
Veröffentlicht: (2025)
Search Arena: Analyzing Search-Augmented LLMs
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
von: Miroyan, Mihran, et al.
Veröffentlicht: (2025)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
von: Verma, Neha, et al.
Veröffentlicht: (2026)
von: Verma, Neha, et al.
Veröffentlicht: (2026)
Evaluation of LLM-based Strategies for the Extraction of Food Product Information from Online Shops
von: Brosch, Christoph, et al.
Veröffentlicht: (2025)
von: Brosch, Christoph, et al.
Veröffentlicht: (2025)
Urdu News Article Recommendation Model using Natural Language Processing Techniques
von: Abbas, Syed Zain, et al.
Veröffentlicht: (2022)
von: Abbas, Syed Zain, et al.
Veröffentlicht: (2022)
LLM-based Relevance Assessment for Web-Scale Search Evaluation at Pinterest
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
Open Deep Search: Democratizing Search with Open-source Reasoning Agents
von: Alzubi, Salaheddin, et al.
Veröffentlicht: (2025)
von: Alzubi, Salaheddin, et al.
Veröffentlicht: (2025)
Lightweight Query Routing for Adaptive RAG: A Baseline Study on RAGRouter-Bench
von: Bansal, Prakhar, et al.
Veröffentlicht: (2026)
von: Bansal, Prakhar, et al.
Veröffentlicht: (2026)
What's happening in your neighborhood? A Weakly Supervised Approach to Detect Local News
von: Shah, Deven Santosh, et al.
Veröffentlicht: (2023)
von: Shah, Deven Santosh, et al.
Veröffentlicht: (2023)
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
von: Mühlbacher, Peter, et al.
Veröffentlicht: (2024)
von: Mühlbacher, Peter, et al.
Veröffentlicht: (2024)
AD-Bench: A Real-World, Trajectory-Aware Advertising Analytics Benchmark for LLM Agents
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2026)
NLP-Powered Repository and Search Engine for Academic Papers: A Case Study on Cyber Risk Literature with CyLit
von: Zhang, Linfeng, et al.
Veröffentlicht: (2024)
von: Zhang, Linfeng, et al.
Veröffentlicht: (2024)
Path-Constrained Retrieval: A Structural Approach to Reliable LLM Agent Reasoning Through Graph-Scoped Semantic Search
von: Oladokun, Joseph
Veröffentlicht: (2025)
von: Oladokun, Joseph
Veröffentlicht: (2025)
LawLLM: Law Large Language Model for the US Legal System
von: Shu, Dong, et al.
Veröffentlicht: (2024)
von: Shu, Dong, et al.
Veröffentlicht: (2024)
Beyond Sequential Reranking: Reranker-Guided Search Improves Reasoning Intensive Retrieval
von: Xu, Haike, et al.
Veröffentlicht: (2025)
von: Xu, Haike, et al.
Veröffentlicht: (2025)
Test-Time Compute for Frozen Embedding Models through Agentic Program Search
von: Xiao, Han
Veröffentlicht: (2026)
von: Xiao, Han
Veröffentlicht: (2026)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Training
von: Xia, Tianle, et al.
Veröffentlicht: (2026)
von: Xia, Tianle, et al.
Veröffentlicht: (2026)
Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools
von: Arat, Baris, et al.
Veröffentlicht: (2026)
von: Arat, Baris, et al.
Veröffentlicht: (2026)
Scaling Up LLM Reviews for Google Ads Content Moderation
von: Qiao, Wei, et al.
Veröffentlicht: (2024)
von: Qiao, Wei, et al.
Veröffentlicht: (2024)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
von: Gera, Ariel, et al.
Veröffentlicht: (2026)
von: Gera, Ariel, et al.
Veröffentlicht: (2026)
CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search
von: Wu, Kaixin, et al.
Veröffentlicht: (2024)
von: Wu, Kaixin, et al.
Veröffentlicht: (2024)
FGTR: Fine-Grained Multi-Table Retrieval via Hierarchical LLM Reasoning
von: Sun, Chaojie, et al.
Veröffentlicht: (2026)
von: Sun, Chaojie, et al.
Veröffentlicht: (2026)
Advancing Academic Knowledge Retrieval via LLM-enhanced Representation Similarity Fusion
von: Dai, Wei, et al.
Veröffentlicht: (2024)
von: Dai, Wei, et al.
Veröffentlicht: (2024)
Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
von: Su, Zhengyang, et al.
Veröffentlicht: (2026)
von: Su, Zhengyang, et al.
Veröffentlicht: (2026)
Transforming User Defined Criteria into Explainable Indicators with an Integrated LLM AHP System
von: Bang, Geonwoo, et al.
Veröffentlicht: (2025)
von: Bang, Geonwoo, et al.
Veröffentlicht: (2025)
VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference
von: Qi, Jasmine, et al.
Veröffentlicht: (2026)
von: Qi, Jasmine, et al.
Veröffentlicht: (2026)
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
ArXivBench: When You Should Avoid Using ChatGPT for Academic Writing
von: Li, Ning, et al.
Veröffentlicht: (2025)
von: Li, Ning, et al.
Veröffentlicht: (2025)
Layered Insights: Generalizable Analysis of Authorial Style by Leveraging All Transformer Layers
von: Alshomary, Milad, et al.
Veröffentlicht: (2025)
von: Alshomary, Milad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Chain-of-Thought Reasoning for Steerable Pluralistic Alignment
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025) -
Forecasting Conversation Derailments Through Generation
von: Zhang, Yunfan, et al.
Veröffentlicht: (2025) -
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
von: Zhang, Weizhi, et al.
Veröffentlicht: (2025) -
FlashEvaluator: Expanding Search Space with Parallel Evaluation
von: Feng, Chao, et al.
Veröffentlicht: (2026) -
Browsing Lost Unformed Recollections: A Benchmark for Tip-of-the-Tongue Search and Reasoning
von: CH-Wang, Sky, et al.
Veröffentlicht: (2025)