QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | V, Venktesh, Anand, Abhijit, Anand, Avishek, Setty, Vinay |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Surprising Effectiveness of Rankers Trained on Expanded Queries
par: Anand, Abhijit, et autres
Publié: (2024)
par: Anand, Abhijit, et autres
Publié: (2024)
LiveFC: A System for Live Fact-Checking of Audio Streams
par: V, Venktesh, et autres
Publié: (2024)
par: V, Venktesh, et autres
Publié: (2024)
Trust but Verify! A Survey on Verification Design for Test-time Scaling
par: Venktesh, V, et autres
Publié: (2025)
par: Venktesh, V, et autres
Publié: (2025)
Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
par: Chungkham, Primakov, et autres
Publié: (2025)
par: Chungkham, Primakov, et autres
Publié: (2025)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
par: Dumitru, Alexandru, et autres
Publié: (2025)
par: Dumitru, Alexandru, et autres
Publié: (2025)
DEXTER: A Benchmark for open-domain Complex Question Answering using LLMs
par: Prabhu, Venktesh V. Deepali, et autres
Publié: (2024)
par: Prabhu, Venktesh V. Deepali, et autres
Publié: (2024)
FactCheck Editor: Multilingual Text Editor with End-to-End fact-checking
par: Setty, Vinay
Publié: (2024)
par: Setty, Vinay
Publié: (2024)
DISCO: DISCovering Overfittings as Causal Rules for Text Classification Models
par: Zhang, Zijian, et autres
Publié: (2024)
par: Zhang, Zijian, et autres
Publié: (2024)
Understanding the User: An Intent-Based Ranking Dataset
par: Anand, Abhijit, et autres
Publié: (2024)
par: Anand, Abhijit, et autres
Publié: (2024)
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models
par: Setty, Vinay
Publié: (2024)
par: Setty, Vinay
Publié: (2024)
A Benchmark for Open-Domain Numerical Fact-Checking Enhanced by Claim Decomposition
par: Venktesh, V, et autres
Publié: (2025)
par: Venktesh, V, et autres
Publié: (2025)
AIC CTU@FEVER 8: On-premise fact checking through long context RAG
par: Ullrich, Herbert, et autres
Publié: (2025)
par: Ullrich, Herbert, et autres
Publié: (2025)
fact check AI at SemEval-2025 Task 7: Multilingual and Crosslingual Fact-checked Claim Retrieval
par: Rastogi, Pranshu
Publié: (2025)
par: Rastogi, Pranshu
Publié: (2025)
Truth Sleuth and Trend Bender: AI Agents to fact-check YouTube videos and influence opinions
par: Logé, Cécile, et autres
Publié: (2025)
par: Logé, Cécile, et autres
Publié: (2025)
Test-time Corpus Feedback: From Retrieval to RAG
par: Rathee, Mandeep, et autres
Publié: (2025)
par: Rathee, Mandeep, et autres
Publié: (2025)
TempRetriever: Fusion-based Temporal Dense Passage Retrieval for Time-Sensitive Questions
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
par: Abdallah, Abdelrahman, et autres
Publié: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
par: Barboule, Camille, et autres
Publié: (2024)
par: Barboule, Camille, et autres
Publié: (2024)
The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
par: Struß, Julia Maria, et autres
Publié: (2026)
par: Struß, Julia Maria, et autres
Publié: (2026)
What do people want to fact-check?
par: Ghafouri, Bijean, et autres
Publié: (2026)
par: Ghafouri, Bijean, et autres
Publié: (2026)
DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives
par: Moon, Sehwan, et autres
Publié: (2025)
par: Moon, Sehwan, et autres
Publié: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
par: Panagoulias, Dimitrios P., et autres
Publié: (2024)
par: Panagoulias, Dimitrios P., et autres
Publié: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
par: Pacchiardi, Lorenzo, et autres
Publié: (2024)
Exploration of Marker-Based Approaches in Argument Mining through Augmented Natural Language
par: Das, Nilmadhab, et autres
Publié: (2024)
par: Das, Nilmadhab, et autres
Publié: (2024)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
par: Anand, Dhruv, et autres
Publié: (2025)
par: Anand, Dhruv, et autres
Publié: (2025)
PKG API: A Tool for Personal Knowledge Graph Management
par: Bernard, Nolwenn, et autres
Publié: (2024)
par: Bernard, Nolwenn, et autres
Publié: (2024)
AKReF: An argumentative knowledge representation framework for structured argumentation
par: Bhattacharjee, Debarati, et autres
Publié: (2025)
par: Bhattacharjee, Debarati, et autres
Publié: (2025)
AWED-FiNER: Agents, Web applications, and Expert Detectors for Fine-grained Named Entity Recognition across 36 Languages for 6.6 Billion Speakers
par: Kaushik, Prachuryya, et autres
Publié: (2026)
par: Kaushik, Prachuryya, et autres
Publié: (2026)
A Reality check of the benefits of LLM in business
par: Cheung, Ming
Publié: (2024)
par: Cheung, Ming
Publié: (2024)
FlashCheck: Exploration of Efficient Evidence Retrieval for Fast Fact-Checking
par: Nanekhan, Kevin, et autres
Publié: (2025)
par: Nanekhan, Kevin, et autres
Publié: (2025)
Development and multi-center evaluation of domain-adapted speech recognition for human-AI teaming in real-world gastrointestinal endoscopy
par: Yang, Ruijie, et autres
Publié: (2026)
par: Yang, Ruijie, et autres
Publié: (2026)
TempPerturb-Eval: On the Joint Effects of Internal Temperature and External Perturbations in RAG Robustness
par: Zhou, Yongxin, et autres
Publié: (2025)
par: Zhou, Yongxin, et autres
Publié: (2025)
SysTemp: A Multi-Agent System for Template-Based Generation of SysML v2
par: Bouamra, Yasmine, et autres
Publié: (2025)
par: Bouamra, Yasmine, et autres
Publié: (2025)
Evaluating open-source Large Language Models for automated fact-checking
par: Fontana, Nicolo', et autres
Publié: (2025)
par: Fontana, Nicolo', et autres
Publié: (2025)
HalluciNot: Hallucination Detection Through Context and Common Knowledge Verification
par: Paudel, Bibek, et autres
Publié: (2025)
par: Paudel, Bibek, et autres
Publié: (2025)
FACT: Examining the Effectiveness of Iterative Context Rewriting for Multi-fact Retrieval
par: Wang, Jinlin, et autres
Publié: (2024)
par: Wang, Jinlin, et autres
Publié: (2024)
FactIR: A Real-World Zero-shot Open-Domain Retrieval Benchmark for Fact-Checking
par: V, Venktesh, et autres
Publié: (2025)
par: V, Venktesh, et autres
Publié: (2025)
End-to-End Argument Mining through Autoregressive Argumentative Structure Prediction
par: Das, Nilmadhab, et autres
Publié: (2025)
par: Das, Nilmadhab, et autres
Publié: (2025)
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models
par: Holtermann, Carolin, et autres
Publié: (2026)
par: Holtermann, Carolin, et autres
Publié: (2026)
QuestGen: Effectiveness of Question Generation Methods for Fact-Checking Applications
par: Setty, Ritvik, et autres
Publié: (2024)
par: Setty, Ritvik, et autres
Publié: (2024)
Attacks by Content: Automated Fact-checking is an AI Security Issue
par: Schlichtkrull, Michael
Publié: (2025)
par: Schlichtkrull, Michael
Publié: (2025)
Documents similaires
-
The Surprising Effectiveness of Rankers Trained on Expanded Queries
par: Anand, Abhijit, et autres
Publié: (2024) -
LiveFC: A System for Live Fact-Checking of Audio Streams
par: V, Venktesh, et autres
Publié: (2024) -
Trust but Verify! A Survey on Verification Design for Test-time Scaling
par: Venktesh, V, et autres
Publié: (2025) -
Think Right, Not More: Test-Time Scaling for Numerical Claim Verification
par: Chungkham, Primakov, et autres
Publié: (2025) -
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
par: Dumitru, Alexandru, et autres
Publié: (2025)