An Analysis of Multilingual FActScore
Fuente:
arXiv
Saved in:
| Main Authors: | Vu, Kim Trong, Krumdick, Michael, Reddy, Varshini, Dernoncourt, Franck, Lai, Viet Dac |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024)
by: Shafayat, Sheikh, et al.
Published: (2024)
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024)
by: Lai, Viet Dac, et al.
Published: (2024)
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024)
by: Reddy, Varshini, et al.
Published: (2024)
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
by: Koncel-Kedziorski, Rik, et al.
Published: (2023)
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models
by: Man, Hieu, et al.
Published: (2025)
by: Man, Hieu, et al.
Published: (2025)
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
by: Krumdick, Michael, et al.
Published: (2025)
by: Krumdick, Michael, et al.
Published: (2025)
SlimLM: An Efficient Small Language Model for On-Device Document Assistance
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
Tokenization with Split Trees
by: Schmidt, Craig W., et al.
Published: (2026)
by: Schmidt, Craig W., et al.
Published: (2026)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
by: Lage, Lucas Fonseca, et al.
Published: (2025)
by: Lage, Lucas Fonseca, et al.
Published: (2025)
On Finding Inconsistencies in Documents
by: Lovering, Charles J., et al.
Published: (2025)
by: Lovering, Charles J., et al.
Published: (2025)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
DynaSaur: Large Language Agents Beyond Predefined Actions
by: Nguyen, Dang, et al.
Published: (2024)
by: Nguyen, Dang, et al.
Published: (2024)
Language Model Probabilities are Not Calibrated in Numeric Contexts
by: Lovering, Charles, et al.
Published: (2024)
by: Lovering, Charles, et al.
Published: (2024)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
OATS: Opinion Aspect Target Sentiment Quadruple Extraction Dataset for Aspect-Based Sentiment Analysis
by: Chebolu, Siva Uday Sampreeth, et al.
Published: (2023)
by: Chebolu, Siva Uday Sampreeth, et al.
Published: (2023)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
Lizard: An Efficient Linearization Framework for Large Language Models
by: Van Nguyen, Chien, et al.
Published: (2025)
by: Van Nguyen, Chien, et al.
Published: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
OWLViz: An Open-World Benchmark for Visual Question Answering
by: Nguyen, Thuy, et al.
Published: (2025)
by: Nguyen, Thuy, et al.
Published: (2025)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
by: Dongre, Vardhan, et al.
Published: (2026)
by: Dongre, Vardhan, et al.
Published: (2026)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
mSCoRe: a $M$ultilingual and Scalable Benchmark for $S$kill-based $Co$mmonsense $Re$asoning
by: Ngo, Nghia Trung, et al.
Published: (2025)
by: Ngo, Nghia Trung, et al.
Published: (2025)
CORG: Generating Answers from Complex, Interrelated Contexts
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
How Much is Enough? The Diminishing Returns of Tokenization Training Data
by: Reddy, Varshini, et al.
Published: (2025)
by: Reddy, Varshini, et al.
Published: (2025)
Cost-Efficient Estimation of General Abilities Across Benchmarks
by: Krumdick, Michael, et al.
Published: (2026)
by: Krumdick, Michael, et al.
Published: (2026)
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
Drift No More? Context Equilibria in Multi-Turn LLM Interactions
by: Dongre, Vardhan, et al.
Published: (2025)
by: Dongre, Vardhan, et al.
Published: (2025)
PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
ULLME: A Unified Framework for Large Language Model Embeddings with Generation-Augmented Learning
by: Man, Hieu, et al.
Published: (2024)
by: Man, Hieu, et al.
Published: (2024)
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
by: Le, Thanh-Long V., et al.
Published: (2025)
by: Le, Thanh-Long V., et al.
Published: (2025)
The Effect of Scripts and Formats on LLM Numeracy
by: Reddy, Varshini, et al.
Published: (2026)
by: Reddy, Varshini, et al.
Published: (2026)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
by: Zhou, Shijie, et al.
Published: (2025)
by: Zhou, Shijie, et al.
Published: (2025)
Explaining Pre-Trained Language Models with Attribution Scores: An Analysis in Low-Resource Settings
by: Zhou, Wei, et al.
Published: (2024)
by: Zhou, Wei, et al.
Published: (2024)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation
by: Suri, Manan, et al.
Published: (2024)
by: Suri, Manan, et al.
Published: (2024)
Similar Items
-
Multi-FAct: Assessing Factuality of Multilingual LLMs using FActScore
by: Shafayat, Sheikh, et al.
Published: (2024) -
SEC-QA: A Systematic Evaluation Corpus for Financial QA
by: Lai, Viet Dac, et al.
Published: (2024) -
DocFinQA: A Long-Context Financial Reasoning Dataset
by: Reddy, Varshini, et al.
Published: (2024) -
BizBench: A Quantitative Reasoning Benchmark for Business and Finance
by: Koncel-Kedziorski, Rik, et al.
Published: (2023) -
LUSIFER: Language Universal Space Integration for Enhanced Multilingual Embeddings with Large Language Models
by: Man, Hieu, et al.
Published: (2025)