BERTScoreVisualizer: A Web Tool for Understanding Simplified Text Evaluation with BERTScore
Fuente:
arXiv
Saved in:
| Main Authors: | Jaskowski, Sebastian, Chava, Sahasra, Shah, Agam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge
by: Shah, Agam, et al.
Published: (2025)
by: Shah, Agam, et al.
Published: (2025)
FinNuE: Exposing the Risks of Using BERTScore for Numerical Semantic Evaluation in Finance
by: Huang, Yu-Shiang, et al.
Published: (2025)
by: Huang, Yu-Shiang, et al.
Published: (2025)
Numerical Claim Detection in Finance: A New Financial Dataset, Weak-Supervision Model, and Market Analysis
by: Shah, Agam, et al.
Published: (2024)
by: Shah, Agam, et al.
Published: (2024)
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
by: Shor, Joel, et al.
Published: (2023)
by: Shor, Joel, et al.
Published: (2023)
Language Modeling for the Future of Finance: A Survey into Metrics, Tasks, and Data Opportunities
by: Tatarinov, Nikita, et al.
Published: (2025)
by: Tatarinov, Nikita, et al.
Published: (2025)
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
by: Ye, Liqin, et al.
Published: (2025)
by: Ye, Liqin, et al.
Published: (2025)
CoCoHD: Congress Committee Hearing Dataset
by: Hiray, Arnav, et al.
Published: (2024)
by: Hiray, Arnav, et al.
Published: (2024)
How Inclusively do LMs Perceive Social and Moral Norms?
by: Galarnyk, Michael, et al.
Published: (2025)
by: Galarnyk, Michael, et al.
Published: (2025)
FiNER-ORD: Financial Named Entity Recognition Open Research Dataset
by: Shah, Agam, et al.
Published: (2023)
by: Shah, Agam, et al.
Published: (2023)
Words That Unite The World: A Unified Framework for Deciphering Central Bank Communications Globally
by: Shah, Agam, et al.
Published: (2025)
by: Shah, Agam, et al.
Published: (2025)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
by: Saeki, Takaaki, et al.
Published: (2024)
by: Saeki, Takaaki, et al.
Published: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
by: Kishi, Minoru, et al.
Published: (2025)
by: Kishi, Minoru, et al.
Published: (2025)
ConfReady: A RAG based Assistant and Dataset for Conference Checklist Responses
by: Galarnyk, Michael, et al.
Published: (2024)
by: Galarnyk, Michael, et al.
Published: (2024)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
by: Galarnyk, Michael, et al.
Published: (2025)
by: Galarnyk, Michael, et al.
Published: (2025)
KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation
by: Tatarinov, Nikita, et al.
Published: (2025)
by: Tatarinov, Nikita, et al.
Published: (2025)
Arctic-Extract Technical Report
by: Chiliński, Mateusz, et al.
Published: (2025)
by: Chiliński, Mateusz, et al.
Published: (2025)
SubjECTive-QA: Measuring Subjectivity in Earnings Call Transcripts' QA Through Six-Dimensional Feature Analysis
by: Pardawala, Huzaifa, et al.
Published: (2024)
by: Pardawala, Huzaifa, et al.
Published: (2024)
Text Anomaly Detection with Simplified Isolation Kernel
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
Exploring Topic Trends in COVID-19 Research Literature using Non-Negative Matrix Factorization
by: Patel, Divya, et al.
Published: (2025)
by: Patel, Divya, et al.
Published: (2025)
From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities
by: Goyal, Agam, et al.
Published: (2026)
by: Goyal, Agam, et al.
Published: (2026)
ColorSwap: A Color and Word Order Dataset for Multimodal Evaluation
by: Burapacheep, Jirayu, et al.
Published: (2024)
by: Burapacheep, Jirayu, et al.
Published: (2024)
Palisade -- Prompt Injection Detection Framework
by: Kokkula, Sahasra, et al.
Published: (2024)
by: Kokkula, Sahasra, et al.
Published: (2024)
The Tool Illusion: Rethinking Tool Use in Web Agents
by: Lou, Renze, et al.
Published: (2026)
by: Lou, Renze, et al.
Published: (2026)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Finance Language Model Evaluation (FLaME)
by: Matlin, Glenn, et al.
Published: (2025)
by: Matlin, Glenn, et al.
Published: (2025)
ArgCMV: An Argument Summarization Benchmark for the LLM-era
by: Gurjar, Omkar, et al.
Published: (2025)
by: Gurjar, Omkar, et al.
Published: (2025)
Translate or Simplify First: An Analysis of Cross-lingual Text Simplification in English and French
by: Dahan, Ido, et al.
Published: (2026)
by: Dahan, Ido, et al.
Published: (2026)
Evaluating GenAI for Simplifying Texts for Education: Improving Accuracy and Consistency for Enhanced Readability
by: Day, Stephanie L., et al.
Published: (2025)
by: Day, Stephanie L., et al.
Published: (2025)
Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities
by: Säuberli, Andreas, et al.
Published: (2024)
by: Säuberli, Andreas, et al.
Published: (2024)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
Harnessing Webpage UIs for Text-Rich Visual Understanding
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
UI-Bench: A Benchmark for Evaluating Design Capabilities of AI Text-to-App Tools
by: Jung, Sam, et al.
Published: (2025)
by: Jung, Sam, et al.
Published: (2025)
Financial Instruction Following Evaluation (FIFE)
by: Matlin, Glenn, et al.
Published: (2025)
by: Matlin, Glenn, et al.
Published: (2025)
SimplifyMyText: An LLM-Based System for Inclusive Plain Language Text Simplification
by: Färber, Michael, et al.
Published: (2025)
by: Färber, Michael, et al.
Published: (2025)
Lost in the Mix: Evaluating LLM Understanding of Code-Switched Text
by: Mohamed, Amr, et al.
Published: (2025)
by: Mohamed, Amr, et al.
Published: (2025)
Learning to Simulate Human Dialogue
by: Gandhi, Kanishk, et al.
Published: (2026)
by: Gandhi, Kanishk, et al.
Published: (2026)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
From Text to Voice: A Reproducible and Verifiable Framework for Evaluating Tool Calling LLM Agents
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
by: Laskar, Md Tahmid Rahman, et al.
Published: (2026)
Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space
by: Wu, Si, et al.
Published: (2025)
by: Wu, Si, et al.
Published: (2025)
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
by: Penedo, Guilherme, et al.
Published: (2024)
by: Penedo, Guilherme, et al.
Published: (2024)
Similar Items
-
Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge
by: Shah, Agam, et al.
Published: (2025) -
FinNuE: Exposing the Risks of Using BERTScore for Numerical Semantic Evaluation in Finance
by: Huang, Yu-Shiang, et al.
Published: (2025) -
Numerical Claim Detection in Finance: A New Financial Dataset, Weak-Supervision Model, and Market Analysis
by: Shah, Agam, et al.
Published: (2024) -
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
by: Shor, Joel, et al.
Published: (2023) -
Language Modeling for the Future of Finance: A Survey into Metrics, Tasks, and Data Opportunities
by: Tatarinov, Nikita, et al.
Published: (2025)