The FACTS Grounding Leaderboard: Benchmarking LLMs' Ability to Ground Responses to Long-Form Input
Fuente:
arXiv
Saved in:
| Main Authors: | Jacovi, Alon, Wang, Andrew, Alberti, Chris, Tao, Connie, Lipovetz, Jon, Olszewska, Kate, Haas, Lukas, Liu, Michelle, Keating, Nate, Bloniarz, Adam, Saroufim, Carl, Fry, Corey, Marcus, Dror, Kukliansky, Doron, Tomar, Gaurav Singh, Swirhun, James, Xing, Jinwei, Wang, Lily, Gurumurthy, Madhu, Aaron, Michael, Ambar, Moran, Fellinger, Rachana, Wang, Rui, Zhang, Zizhao, Goldshtein, Sasha, Das, Dipanjan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
by: Cheng, Aileen, et al.
Published: (2025)
by: Cheng, Aileen, et al.
Published: (2025)
CoverBench: A Challenging Benchmark for Complex Claim Verification
by: Jacovi, Alon, et al.
Published: (2024)
by: Jacovi, Alon, et al.
Published: (2024)
Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios
by: Ivry, Dror, et al.
Published: (2025)
by: Ivry, Dror, et al.
Published: (2025)
Quantum Circuit Tensors and Enumerators with Applications to Quantum Fault Tolerance
by: Kukliansky, Alon, et al.
Published: (2024)
by: Kukliansky, Alon, et al.
Published: (2024)
Pluralistic Leaderboards
by: Haghtalab, Nika, et al.
Published: (2026)
by: Haghtalab, Nika, et al.
Published: (2026)
Porous Carbon Nanospheres Derived From Caesalpinia Sappan Pods as Novel Antibacterial Agents
by: Suvadra Das, et al.
Published: (2025)
by: Suvadra Das, et al.
Published: (2025)
SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge
by: Haas, Lukas, et al.
Published: (2025)
by: Haas, Lukas, et al.
Published: (2025)
The Leaderboard Illusion
by: Singh, Shivalika, et al.
Published: (2025)
by: Singh, Shivalika, et al.
Published: (2025)
Open Universal Arabic ASR Leaderboard
by: Wang, Yingzhi, et al.
Published: (2024)
by: Wang, Yingzhi, et al.
Published: (2024)
Investigating Content Planning for Navigating Trade-offs in Knowledge-Grounded Dialogue
by: Chawla, Kushal, et al.
Published: (2024)
by: Chawla, Kushal, et al.
Published: (2024)
TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
by: Caciularu, Avi, et al.
Published: (2024)
by: Caciularu, Avi, et al.
Published: (2024)
Sustainability-driven Data Management Strategies
by: Fellinger, Engelbert, et al.
Published: (2025)
by: Fellinger, Engelbert, et al.
Published: (2025)
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
by: Li, Haonan, et al.
Published: (2024)
by: Li, Haonan, et al.
Published: (2024)
SafeGround: Know When to Trust GUI Grounding Models via Uncertainty Calibration
by: Wang, Qingni, et al.
Published: (2026)
by: Wang, Qingni, et al.
Published: (2026)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
by: Cattan, Arie, et al.
Published: (2025)
by: Cattan, Arie, et al.
Published: (2025)
FACTS: Fine-Grained Action Classification for Tactical Sports
by: Lai, Christopher, et al.
Published: (2024)
by: Lai, Christopher, et al.
Published: (2024)
COMMUNISTS: THE FACTS OF LIFE
Published: (1948)
Published: (1948)
FACTS TOWARDS LIFE
by: Ripu Ranjan Sinha
Published: (2019)
by: Ripu Ranjan Sinha
Published: (2019)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
Prompt-to-Leaderboard
by: Frick, Evan, et al.
Published: (2025)
by: Frick, Evan, et al.
Published: (2025)
FACTS: A Factored State-Space Framework For World Modelling
by: Nanbo, Li, et al.
Published: (2024)
by: Nanbo, Li, et al.
Published: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
by: Cao, Shengcao, et al.
Published: (2024)
by: Cao, Shengcao, et al.
Published: (2024)
Grounded Curriculum Learning
by: Wang, Linji, et al.
Published: (2024)
by: Wang, Linji, et al.
Published: (2024)
THE MIDDLE EAST: FACING FACTS
Published: (1958)
Published: (1958)
VGR: Visual Grounded Reasoning
by: Wang, Jiacong, et al.
Published: (2025)
by: Wang, Jiacong, et al.
Published: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
On the Workflows and Smells of Leaderboard Operations (LBOps): An Exploratory Study of Foundation Model Leaderboards
by: Zhao, Zhimin, et al.
Published: (2024)
by: Zhao, Zhimin, et al.
Published: (2024)
On the Evolution During Growth of Regular Boundaries of Bodies into Fractals
by: Goldshtein, Vladimir, et al.
Published: (2024)
by: Goldshtein, Vladimir, et al.
Published: (2024)
QFactor: A Domain-Specific Optimizer for Quantum Circuit Instantiation
by: Kukliansky, Alon, et al.
Published: (2023)
by: Kukliansky, Alon, et al.
Published: (2023)
Leveraging Quantum Machine Learning Generalization to Significantly Speed-up Quantum Compilation
by: Kukliansky, Alon, et al.
Published: (2024)
by: Kukliansky, Alon, et al.
Published: (2024)
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
by: Wang, Zhipin, et al.
Published: (2026)
by: Wang, Zhipin, et al.
Published: (2026)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
by: Zhao, Haoren, et al.
Published: (2026)
by: Zhao, Haoren, et al.
Published: (2026)
On the Naturalistic Grounds of Grounding
by: Raoni Arroyo, et al.
Published: (2026)
by: Raoni Arroyo, et al.
Published: (2026)
Calibration without Ground Truth
by: Kong, Yuqing, et al.
Published: (2026)
by: Kong, Yuqing, et al.
Published: (2026)
Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding
by: Xu, Zhengtong, et al.
Published: (2026)
by: Xu, Zhengtong, et al.
Published: (2026)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)
by: Zhang, Jiaxi, et al.
Published: (2026)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
A ranking of treatment efficacy in alopecia areata is not possible without head‐to‐head studies
by: Lidia Rudnicka, et al.
Published: (2024)
by: Lidia Rudnicka, et al.
Published: (2024)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
by: Hu, Rui, et al.
Published: (2025)
by: Hu, Rui, et al.
Published: (2025)
GroundSLAM: A Robust Visual SLAM System for Warehouse Robots Using Ground Textures
by: Xu, Kuan, et al.
Published: (2017)
by: Xu, Kuan, et al.
Published: (2017)
Similar Items
-
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
by: Cheng, Aileen, et al.
Published: (2025) -
CoverBench: A Challenging Benchmark for Complex Claim Verification
by: Jacovi, Alon, et al.
Published: (2024) -
Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios
by: Ivry, Dror, et al.
Published: (2025) -
Quantum Circuit Tensors and Enumerators with Applications to Quantum Fault Tolerance
by: Kukliansky, Alon, et al.
Published: (2024) -
Pluralistic Leaderboards
by: Haghtalab, Nika, et al.
Published: (2026)