GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification
Fuente:
arXiv
Salvato in:
| Autori principali: | Fostiropoulos, Iordanis, Azhar, Muhammad Rafay, Sawwan, Abdalaziz, Fang, Boyu, Liu, Yuchen, Liu, Jiayi, Yu, Hanchao, Guo, Qi, Wang, Jianyu, Liu, Fei, Fan, Xiangjun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
2024 Google Scholar Research Interest Ranking for Top 3260 Computer Science Authors
di: Rasane, Atharva
Pubblicazione: (2024)
di: Rasane, Atharva
Pubblicazione: (2024)
Mind the Gap: Aligning Knowledge Bases with User Needs to Enhance Mental Health Retrieval
di: Chan, Amanda, et al.
Pubblicazione: (2025)
di: Chan, Amanda, et al.
Pubblicazione: (2025)
Session Context Embedding for Intent Understanding in Product Search
di: Mehrdad, Navid, et al.
Pubblicazione: (2024)
di: Mehrdad, Navid, et al.
Pubblicazione: (2024)
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
di: Zhang, Yiqing, et al.
Pubblicazione: (2026)
di: Zhang, Yiqing, et al.
Pubblicazione: (2026)
Criteria-Based LLM Relevance Judgments
di: Farzi, Naghmeh, et al.
Pubblicazione: (2025)
di: Farzi, Naghmeh, et al.
Pubblicazione: (2025)
NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance
di: Eyasir, Abrar, et al.
Pubblicazione: (2026)
di: Eyasir, Abrar, et al.
Pubblicazione: (2026)
When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
di: Shao, Jiaqi, et al.
Pubblicazione: (2026)
di: Shao, Jiaqi, et al.
Pubblicazione: (2026)
Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship
di: Pan, Yating, et al.
Pubblicazione: (2026)
di: Pan, Yating, et al.
Pubblicazione: (2026)
Scaling Multilingual Semantic Search in Uber Eats Delivery
di: Ling, Bo, et al.
Pubblicazione: (2026)
di: Ling, Bo, et al.
Pubblicazione: (2026)
NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data
di: Ming, Cong, et al.
Pubblicazione: (2026)
di: Ming, Cong, et al.
Pubblicazione: (2026)
Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
di: Liu, Jianan, et al.
Pubblicazione: (2026)
di: Liu, Jianan, et al.
Pubblicazione: (2026)
Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation
di: Ye, Hua, et al.
Pubblicazione: (2026)
di: Ye, Hua, et al.
Pubblicazione: (2026)
Automating Pharmacovigilance Evidence Generation: Using Large Language Models to Produce Context-Aware SQL
di: Painter, Jeffery L., et al.
Pubblicazione: (2024)
di: Painter, Jeffery L., et al.
Pubblicazione: (2024)
What Matters in LLM-Based Feature Extractor for Recommender? A Systematic Analysis of Prompts, Models, and Adaptation
di: Shi, Kainan, et al.
Pubblicazione: (2025)
di: Shi, Kainan, et al.
Pubblicazione: (2025)
Agentic Retrieval-Augmented Generation for Financial Document Question Answering
di: Shu, Yang, et al.
Pubblicazione: (2026)
di: Shu, Yang, et al.
Pubblicazione: (2026)
The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges
di: Liu, Jiqun, et al.
Pubblicazione: (2024)
di: Liu, Jiqun, et al.
Pubblicazione: (2024)
Pre-trained LLMs Meet Sequential Recommenders: Efficient User-Centric Knowledge Distillation
di: Severin, Nikita, et al.
Pubblicazione: (2026)
di: Severin, Nikita, et al.
Pubblicazione: (2026)
The Case for Intent-Based Query Rewriting
di: Nicolai, Gianna Lisa, et al.
Pubblicazione: (2025)
di: Nicolai, Gianna Lisa, et al.
Pubblicazione: (2025)
Motion Compensation for Real Time Ultrasound Scanning in Robotically Assisted Prostate Biopsy Procedures
di: Markulin, Matija, et al.
Pubblicazione: (2026)
di: Markulin, Matija, et al.
Pubblicazione: (2026)
SGMem: Sentence Graph Memory for Long-Term Conversational Agents
di: Wu, Yaxiong, et al.
Pubblicazione: (2025)
di: Wu, Yaxiong, et al.
Pubblicazione: (2025)
Less LLM, More Documents: Searching for Improved RAG
di: Ning, Jingjie, et al.
Pubblicazione: (2025)
di: Ning, Jingjie, et al.
Pubblicazione: (2025)
Query-Centric Graph Retrieval Augmented Generation
di: Wu, Yaxiong, et al.
Pubblicazione: (2025)
di: Wu, Yaxiong, et al.
Pubblicazione: (2025)
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
di: Zeng, Yangchen, et al.
Pubblicazione: (2026)
Understanding Multi-Agent LLM Frameworks: A Unified Benchmark and Experimental Analysis
di: Orogat, Abdelghny, et al.
Pubblicazione: (2026)
di: Orogat, Abdelghny, et al.
Pubblicazione: (2026)
ByteRover: Agent-Native Memory Through LLM-Curated Hierarchical Context
di: Nguyen, Andy, et al.
Pubblicazione: (2026)
di: Nguyen, Andy, et al.
Pubblicazione: (2026)
RLM-on-KG: Heuristics First, LLMs When Needed: Adaptive Retrieval Control over Mention Graphs for Scattered Evidence
di: Volpini, Andrea, et al.
Pubblicazione: (2026)
di: Volpini, Andrea, et al.
Pubblicazione: (2026)
A Systematic Framework for Enterprise Knowledge Retrieval: Leveraging LLM-Generated Metadata to Enhance RAG Systems
di: Mishra, Pranav Pushkar, et al.
Pubblicazione: (2025)
di: Mishra, Pranav Pushkar, et al.
Pubblicazione: (2025)
Deterministic Fuzzy Triage for Legal Compliance Classification and Evidence Retrieval
di: Atri, Rian
Pubblicazione: (2026)
di: Atri, Rian
Pubblicazione: (2026)
Routing End User Queries to Enterprise Databases
di: Sudarshan, Saikrishna, et al.
Pubblicazione: (2026)
di: Sudarshan, Saikrishna, et al.
Pubblicazione: (2026)
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
di: Yu, Boxi, et al.
Pubblicazione: (2026)
di: Yu, Boxi, et al.
Pubblicazione: (2026)
Flippi: End To End GenAI Assistant for E-Commerce
di: Rajasekar, Anand A., et al.
Pubblicazione: (2025)
di: Rajasekar, Anand A., et al.
Pubblicazione: (2025)
Fine-Grained Emotion Recognition via In-Context Learning
di: Ren, Zhaochun, et al.
Pubblicazione: (2025)
di: Ren, Zhaochun, et al.
Pubblicazione: (2025)
Adaptive ToR: Complexity-Aware Tree-Based Retrieval for Pareto-Optimal Multi-Intent NLU
di: Yoo, Hee-Kyong, et al.
Pubblicazione: (2026)
di: Yoo, Hee-Kyong, et al.
Pubblicazione: (2026)
MasterSet: A Large-Scale Benchmark for Must-Cite Citation Recommendation in the AI/ML Literature
di: Ratul, Md Toyaha Rahman, et al.
Pubblicazione: (2026)
di: Ratul, Md Toyaha Rahman, et al.
Pubblicazione: (2026)
NewsScope: Schema-Grounded Cross-Domain News Claim Extraction with Open Models
di: Pandya, Nidhi
Pubblicazione: (2025)
di: Pandya, Nidhi
Pubblicazione: (2025)
Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
di: Chhoun, Sovandara, et al.
Pubblicazione: (2026)
di: Chhoun, Sovandara, et al.
Pubblicazione: (2026)
Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables
di: Shen, Chen
Pubblicazione: (2026)
di: Shen, Chen
Pubblicazione: (2026)
BridgeRAG: Training-Free Bridge-Conditioned Retrieval for Multi-Hop Question Answering
di: Bacellar, Andre
Pubblicazione: (2026)
di: Bacellar, Andre
Pubblicazione: (2026)
Train Once, Use Flexibly: A Modular Framework for Multi-Aspect Neural News Recommendation
di: Iana, Andreea, et al.
Pubblicazione: (2023)
di: Iana, Andreea, et al.
Pubblicazione: (2023)
Towards Adaptive Context Management for Intelligent Conversational Question Answering
di: Perera, Manoj Madushanka, et al.
Pubblicazione: (2025)
di: Perera, Manoj Madushanka, et al.
Pubblicazione: (2025)
Documenti analoghi
-
2024 Google Scholar Research Interest Ranking for Top 3260 Computer Science Authors
di: Rasane, Atharva
Pubblicazione: (2024) -
Mind the Gap: Aligning Knowledge Bases with User Needs to Enhance Mental Health Retrieval
di: Chan, Amanda, et al.
Pubblicazione: (2025) -
Session Context Embedding for Intent Understanding in Product Search
di: Mehrdad, Navid, et al.
Pubblicazione: (2024) -
PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
di: Zhang, Yiqing, et al.
Pubblicazione: (2026) -
Criteria-Based LLM Relevance Judgments
di: Farzi, Naghmeh, et al.
Pubblicazione: (2025)