Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks
Fuente:
arXiv
Saved in:
| Main Authors: | Ohno, Keyaki, Kameko, Hirotaka, Shirai, Keisuke, Nishimura, Taichi, Mori, Shinsuke |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022)
by: Nishimura, Taichi, et al.
Published: (2022)
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
by: Haneji, Yuto, et al.
Published: (2024)
by: Haneji, Yuto, et al.
Published: (2024)
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
by: Nishimoto, Tomohiro, et al.
Published: (2024)
by: Nishimoto, Tomohiro, et al.
Published: (2024)
Coarse-Tuning for Ad-hoc Document Retrieval Using Pre-trained Language Models
by: Keyaki, Atsushi, et al.
Published: (2024)
by: Keyaki, Atsushi, et al.
Published: (2024)
DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
by: Ju, Rui-Yang, et al.
Published: (2025)
by: Ju, Rui-Yang, et al.
Published: (2025)
Restoration-Guided Kuzushiji Character Recognition Framework under Seal Interference
by: Ju, Rui-Yang, et al.
Published: (2026)
by: Ju, Rui-Yang, et al.
Published: (2026)
Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models
by: Semnani, Sina J., et al.
Published: (2025)
by: Semnani, Sina J., et al.
Published: (2025)
Biomedical Entity Linking for Dutch: Fine-tuning a Self-alignment BERT Model on an Automatically Generated Wikipedia Corpus
by: Hartendorp, Fons, et al.
Published: (2024)
by: Hartendorp, Fons, et al.
Published: (2024)
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2025)
by: Yoshida, Tomoya, et al.
Published: (2025)
Text-driven Affordance Learning from Egocentric Vision
by: Yoshida, Tomoya, et al.
Published: (2024)
by: Yoshida, Tomoya, et al.
Published: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Vision-Language Interpreter for Robot Task Planning
by: Shirai, Keisuke, et al.
Published: (2023)
by: Shirai, Keisuke, et al.
Published: (2023)
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
by: Ando, Kenichiro, et al.
Published: (2023)
by: Ando, Kenichiro, et al.
Published: (2023)
Edisum: Summarizing and Explaining Wikipedia Edits at Scale
by: Šakota, Marija, et al.
Published: (2024)
by: Šakota, Marija, et al.
Published: (2024)
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
by: Simonotti, Martina, et al.
Published: (2026)
by: Simonotti, Martina, et al.
Published: (2026)
Leveraging Corpus Metadata to Detect Template-based Translation: An Exploratory Case Study of the Egyptian Arabic Wikipedia Edition
by: Alshahrani, Saied, et al.
Published: (2024)
by: Alshahrani, Saied, et al.
Published: (2024)
Vaporetto: Efficient Japanese Tokenization Based on Improved Pointwise Linear Classification
by: Akabe, Koichi, et al.
Published: (2024)
by: Akabe, Koichi, et al.
Published: (2024)
LegalRikai: Open Benchmark -- Benchmark for Complex Japanese Corporate Legal Tasks
by: Fujita, Shogo, et al.
Published: (2025)
by: Fujita, Shogo, et al.
Published: (2025)
EduDial: Constructing a Large-scale Multi-turn Teacher-Student Dialogue Corpus
by: Wei, Shouang, et al.
Published: (2025)
by: Wei, Shouang, et al.
Published: (2025)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
AE-LLM: Adaptive Efficiency Optimization for Large Language Models
by: Tanaka, Kaito, et al.
Published: (2026)
by: Tanaka, Kaito, et al.
Published: (2026)
Detecting Sockpuppetry on Wikipedia Using Meta-Learning
by: Raszewski, Luc, et al.
Published: (2025)
by: Raszewski, Luc, et al.
Published: (2025)
A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
by: Achkar, Pierre, et al.
Published: (2026)
by: Achkar, Pierre, et al.
Published: (2026)
RAG System for Supporting Japanese Litigation Procedures: Faithful Response Generation Complying with Legal Norms
by: Ishihara, Yuya, et al.
Published: (2025)
by: Ishihara, Yuya, et al.
Published: (2025)
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024)
by: Munakata, Hokuto, et al.
Published: (2024)
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
by: Munakata, Hokuto, et al.
Published: (2025)
by: Munakata, Hokuto, et al.
Published: (2025)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
by: Ohkawa, Takehiko, et al.
Published: (2023)
by: Ohkawa, Takehiko, et al.
Published: (2023)
DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain
by: Cruz, Walter Hernandez, et al.
Published: (2026)
by: Cruz, Walter Hernandez, et al.
Published: (2026)
Developing Vision-Language-Action Model from Egocentric Videos
by: Yoshida, Tomoya, et al.
Published: (2025)
by: Yoshida, Tomoya, et al.
Published: (2025)
C-VARC: A Large-Scale Chinese Value Rule Corpus for Value Alignment of Large Language Models
by: Wu, Ping, et al.
Published: (2025)
by: Wu, Ping, et al.
Published: (2025)
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
by: Gao, Fan, et al.
Published: (2023)
by: Gao, Fan, et al.
Published: (2023)
Rakuten Data Release: A Large-Scale and Long-Term Reviews Corpus for Hotel Domain
by: Nakayama, Yuki, et al.
Published: (2025)
by: Nakayama, Yuki, et al.
Published: (2025)
JaParaPat: A Large-Scale Japanese-English Parallel Patent Application Corpus
by: Nagata, Masaaki, et al.
Published: (2025)
by: Nagata, Masaaki, et al.
Published: (2025)
KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus
by: Shi, Xiaoming, et al.
Published: (2025)
by: Shi, Xiaoming, et al.
Published: (2025)
Multilingual Pretraining Using a Large Corpus Machine-Translated from a Single Source Language
by: Wang, Jiayi, et al.
Published: (2024)
by: Wang, Jiayi, et al.
Published: (2024)
Matina: A Large-Scale 73B Token Persian Text Corpus
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
by: Hosseinbeigi, Sara Bourbour, et al.
Published: (2025)
The Rise of AI-Generated Content in Wikipedia
by: Brooks, Creston, et al.
Published: (2024)
by: Brooks, Creston, et al.
Published: (2024)
Verifying Claims About Metaphors with Large-Scale Automatic Metaphor Identification
by: Aono, Kotaro, et al.
Published: (2024)
by: Aono, Kotaro, et al.
Published: (2024)
LAMUS: A Large-Scale Corpus for Legal Argument Mining from U.S. Caselaw using LLMs
by: Wang, Serene, et al.
Published: (2026)
by: Wang, Serene, et al.
Published: (2026)
Similar Items
-
Recipe Generation from Unsegmented Cooking Videos
by: Nishimura, Taichi, et al.
Published: (2022) -
EgoOops: A Dataset for Mistake Action Detection from Egocentric Videos referring to Procedural Texts
by: Haneji, Yuto, et al.
Published: (2024) -
BioVL-QR: Egocentric Biochemical Vision-and-Language Dataset Using Micro QR Codes
by: Nishimoto, Tomohiro, et al.
Published: (2024) -
Coarse-Tuning for Ad-hoc Document Retrieval Using Pre-trained Language Models
by: Keyaki, Atsushi, et al.
Published: (2024) -
DKDS: A Benchmark Dataset of Degraded Kuzushiji Documents with Seals for Detection and Binarization
by: Ju, Rui-Yang, et al.
Published: (2025)