WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
Fuente:
arXiv
Salvato in:
| Autori principali: | Ando, Kenichiro, Sekine, Satoshi, Komachi, Mamoru |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024)
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024)
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
di: Göpfert, Jan, et al.
Pubblicazione: (2025)
di: Göpfert, Jan, et al.
Pubblicazione: (2025)
Pruning Multilingual Large Language Models for Multilingual Inference
di: Kim, Hwichan, et al.
Pubblicazione: (2024)
di: Kim, Hwichan, et al.
Pubblicazione: (2024)
Revisiting Meta-evaluation for Grammatical Error Correction
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024)
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024)
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
di: Sato, Ayako, et al.
Pubblicazione: (2026)
di: Sato, Ayako, et al.
Pubblicazione: (2026)
Assessing the Capabilities of LLMs in Humor:A Multi-dimensional Analysis of Oogiri Generation and Evaluation
di: Sakabe, Ritsu, et al.
Pubblicazione: (2025)
di: Sakabe, Ritsu, et al.
Pubblicazione: (2025)
Aligning Large Language Model Behavior with Human Citation Preferences
di: Ando, Kenichiro, et al.
Pubblicazione: (2026)
di: Ando, Kenichiro, et al.
Pubblicazione: (2026)
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia
di: Semnani, Sina J., et al.
Pubblicazione: (2023)
di: Semnani, Sina J., et al.
Pubblicazione: (2023)
Tracking Temporal Dynamics of Vector Sets with Gaussian Process
di: Aida, Taichi, et al.
Pubblicazione: (2025)
di: Aida, Taichi, et al.
Pubblicazione: (2025)
Wiki-TabNER: Integrating Named Entity Recognition into Wikipedia Tables
di: Koleva, Aneta, et al.
Pubblicazione: (2024)
di: Koleva, Aneta, et al.
Pubblicazione: (2024)
Wiki Live Challenge: Challenging Deep Research Agents with Expert-Level Wikipedia Articles
di: Wang, Shaohan, et al.
Pubblicazione: (2026)
di: Wang, Shaohan, et al.
Pubblicazione: (2026)
ViWikiFC: Fact-Checking for Vietnamese Wikipedia-Based Textual Knowledge Source
di: Le, Hung Tuan, et al.
Pubblicazione: (2024)
di: Le, Hung Tuan, et al.
Pubblicazione: (2024)
Analyzing Continuous Semantic Shifts with Diachronic Word Similarity Matrices
di: Kiyama, Hajime, et al.
Pubblicazione: (2025)
di: Kiyama, Hajime, et al.
Pubblicazione: (2025)
Human-Grounded Multimodal Benchmark with 900K-Scale Aggregated Student Response Distributions from Japan's National Assessment of Academic Ability
di: Takami, Kyosuke, et al.
Pubblicazione: (2026)
di: Takami, Kyosuke, et al.
Pubblicazione: (2026)
Wikipedia is Not a Dictionary, Delete! Text Classification as a Proxy for Analysing Wiki Deletion Discussions
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2025)
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2025)
AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output
di: Suzuki, Hisami, et al.
Pubblicazione: (2025)
di: Suzuki, Hisami, et al.
Pubblicazione: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
di: Hou, Yufang, et al.
Pubblicazione: (2024)
di: Hou, Yufang, et al.
Pubblicazione: (2024)
WikiGap: Promoting Epistemic Equity by Surfacing Knowledge Gaps Between English Wikipedia and other Language Editions
di: Wang, Zining, et al.
Pubblicazione: (2025)
di: Wang, Zining, et al.
Pubblicazione: (2025)
A Large-Scale Benchmark for Vietnamese Sentence Paraphrases
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
di: Nguyen, Sang Quang, et al.
Pubblicazione: (2025)
Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks
di: Ohno, Keyaki, et al.
Pubblicazione: (2024)
di: Ohno, Keyaki, et al.
Pubblicazione: (2024)
Proper Noun Diacritization for Arabic Wikipedia: A Benchmark Dataset
di: Bondok, Rawan, et al.
Pubblicazione: (2025)
di: Bondok, Rawan, et al.
Pubblicazione: (2025)
Edisum: Summarizing and Explaining Wikipedia Edits at Scale
di: Šakota, Marija, et al.
Pubblicazione: (2024)
di: Šakota, Marija, et al.
Pubblicazione: (2024)
Datasets for Multilingual Answer Sentence Selection
di: Gabburo, Matteo, et al.
Pubblicazione: (2024)
di: Gabburo, Matteo, et al.
Pubblicazione: (2024)
WikiHint: A Human-Annotated Dataset for Hint Ranking and Generation
di: Mozafari, Jamshid, et al.
Pubblicazione: (2024)
di: Mozafari, Jamshid, et al.
Pubblicazione: (2024)
WikiFactDiff: A Large, Realistic, and Temporally Adaptable Dataset for Atomic Factual Knowledge Update in Causal Language Models
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2024)
di: Khodja, Hichem Ammar, et al.
Pubblicazione: (2024)
CHEW: A Dataset of CHanging Events in Wikipedia
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2024)
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2024)
PhantomWiki: On-Demand Datasets for Reasoning and Retrieval Evaluation
di: Gong, Albert, et al.
Pubblicazione: (2025)
di: Gong, Albert, et al.
Pubblicazione: (2025)
Should We Respect LLMs? A Cross-Lingual Study on the Influence of Prompt Politeness on LLM Performance
di: Yin, Ziqi, et al.
Pubblicazione: (2024)
di: Yin, Ziqi, et al.
Pubblicazione: (2024)
Hoaxpedia: A Unified Wikipedia Hoax Articles Dataset
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2024)
di: Borkakoty, Hsuvas, et al.
Pubblicazione: (2024)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
di: Hassan, Umair
Pubblicazione: (2025)
di: Hassan, Umair
Pubblicazione: (2025)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
di: Yin, Kayo, et al.
Pubblicazione: (2024)
di: Yin, Kayo, et al.
Pubblicazione: (2024)
Refining Sentence Embedding Model through Ranking Sentences Generation with Large Language Models
di: He, Liyang, et al.
Pubblicazione: (2025)
di: He, Liyang, et al.
Pubblicazione: (2025)
COIG-P: A High-Quality and Large-Scale Chinese Preference Dataset for Alignment with Human Values
di: P Team, et al.
Pubblicazione: (2025)
di: P Team, et al.
Pubblicazione: (2025)
Subspace Representations for Soft Set Operations and Sentence Similarities
di: Ishibashi, Yoichi, et al.
Pubblicazione: (2022)
di: Ishibashi, Yoichi, et al.
Pubblicazione: (2022)
Japanese-English Sentence Translation Exercises Dataset for Automatic Grading
di: Miura, Naoki, et al.
Pubblicazione: (2024)
di: Miura, Naoki, et al.
Pubblicazione: (2024)
Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset
di: Samir, Farhan, et al.
Pubblicazione: (2024)
di: Samir, Farhan, et al.
Pubblicazione: (2024)
How Good is Your Wikipedia? Auditing Data Quality for Low-resource and Multilingual NLP
di: Tatariya, Kushal, et al.
Pubblicazione: (2024)
di: Tatariya, Kushal, et al.
Pubblicazione: (2024)
Is a Document Educational or Just Wikipedia-Style? -- Pitfalls of Classifier-Based Quality Filtering
di: Klimaszewski, Mateusz, et al.
Pubblicazione: (2026)
di: Klimaszewski, Mateusz, et al.
Pubblicazione: (2026)
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
di: Kovalev, Grigory, et al.
Pubblicazione: (2025)
CoheMark: A Novel Sentence-Level Watermark for Enhanced Text Quality
di: Zhang, Junyan, et al.
Pubblicazione: (2025)
di: Zhang, Junyan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024) -
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
di: Göpfert, Jan, et al.
Pubblicazione: (2025) -
Pruning Multilingual Large Language Models for Multilingual Inference
di: Kim, Hwichan, et al.
Pubblicazione: (2024) -
Revisiting Meta-evaluation for Grammatical Error Correction
di: Kobayashi, Masamune, et al.
Pubblicazione: (2024) -
Exploring the Effects of Alignment on Numerical Bias in Large Language Models
di: Sato, Ayako, et al.
Pubblicazione: (2026)