CC-GPX: Extracting High-Quality Annotated Geospatial Data from Common Crawl
Fuente:
arXiv
Saved in:
| Main Authors: | Ilyankou, Ilya, Wang, Meihui, Cavazzi, Stefano, Haworth, James |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
Much of Geospatial Web Search Is Beyond Traditional GIS
by: Ilyankou, Ilya, et al.
Published: (2026)
by: Ilyankou, Ilya, et al.
Published: (2026)
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?
by: Ilyankou, Ilya, et al.
Published: (2024)
by: Ilyankou, Ilya, et al.
Published: (2024)
The Scenic Route to Deception: Dark Patterns and Explainability Pitfalls in Conversational Navigation
by: Ilyankou, Ilya, et al.
Published: (2026)
by: Ilyankou, Ilya, et al.
Published: (2026)
CycleTrajectory: An End-to-End Pipeline for Enriching and Analyzing GPS Trajectories to Understand Cycling Behavior and Environment
by: Wang, Meihui, et al.
Published: (2024)
by: Wang, Meihui, et al.
Published: (2024)
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
by: Su, Dan, et al.
Published: (2024)
by: Su, Dan, et al.
Published: (2024)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Leveraging Web-Crawled Data for High-Quality Fine-Tuning
by: Zhou, Jing, et al.
Published: (2024)
by: Zhou, Jing, et al.
Published: (2024)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images
by: Ilyankou, Ilya, et al.
Published: (2025)
by: Ilyankou, Ilya, et al.
Published: (2025)
Multiple Object Detection and Tracking in Panoramic Videos for Cycling Safety Analysis
by: Guo, Jingwei, et al.
Published: (2024)
by: Guo, Jingwei, et al.
Published: (2024)
The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project
by: Aquino, Angelina A., et al.
Published: (2025)
by: Aquino, Angelina A., et al.
Published: (2025)
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages
by: Tessema, Bethel Melesse, et al.
Published: (2024)
by: Tessema, Bethel Melesse, et al.
Published: (2024)
GeoLLM: Extracting Geospatial Knowledge from Large Language Models
by: Manvi, Rohin, et al.
Published: (2023)
by: Manvi, Rohin, et al.
Published: (2023)
Introducing the NewsPaLM MBR and QE Dataset: LLM-Generated High-Quality Parallel Data Outperforms Traditional Web-Crawled Data
by: Finkelstein, Mara, et al.
Published: (2024)
by: Finkelstein, Mara, et al.
Published: (2024)
Infini-News: Efficiently Queryable Access to 1.3 Billion Processed Common Crawl News Articles
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
by: Lazzaroni, Ruggero Marino, et al.
Published: (2026)
V-RoAst: Visual Road Assessment. Can VLM be a Road Safety Assessor Using the iRAP Standard?
by: Jongwiriyanurak, Natchapon, et al.
Published: (2024)
by: Jongwiriyanurak, Natchapon, et al.
Published: (2024)
WanJuan-CC: A Safe and High-Quality Open-sourced English Webtext Dataset
by: Qiu, Jiantao, et al.
Published: (2024)
by: Qiu, Jiantao, et al.
Published: (2024)
An LLM Agent for Automatic Geospatial Data Analysis
by: Chen, Yuxing, et al.
Published: (2024)
by: Chen, Yuxing, et al.
Published: (2024)
DialogCC: An Automated Pipeline for Creating High-Quality Multi-Modal Dialogue Dataset
by: Lee, Young-Jun, et al.
Published: (2022)
by: Lee, Young-Jun, et al.
Published: (2022)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
by: van Noord, Rik, et al.
Published: (2024)
by: van Noord, Rik, et al.
Published: (2024)
Craw4LLM: Efficient Web Crawling for LLM Pretraining
by: Yu, Shi, et al.
Published: (2025)
by: Yu, Shi, et al.
Published: (2025)
Smart Bilingual Focused Crawling of Parallel Documents
by: García-Romero, Cristian, et al.
Published: (2024)
by: García-Romero, Cristian, et al.
Published: (2024)
Web Page Classification using LLMs for Crawling Support
by: Sasazawa, Yuichi, et al.
Published: (2025)
by: Sasazawa, Yuichi, et al.
Published: (2025)
On Efficient and Statistical Quality Estimation for Data Annotation
by: Klie, Jan-Christoph, et al.
Published: (2024)
by: Klie, Jan-Christoph, et al.
Published: (2024)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
Large Language Models as Annotators for Machine Translation Quality Estimation
by: Wang, Sidi, et al.
Published: (2026)
by: Wang, Sidi, et al.
Published: (2026)
Analyzing Dataset Annotation Quality Management in the Wild
by: Klie, Jan-Christoph, et al.
Published: (2023)
by: Klie, Jan-Christoph, et al.
Published: (2023)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
CC-LEARN: Cohort-based Consistency Learning
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
by: Lu, Qingyu, et al.
Published: (2024)
by: Lu, Qingyu, et al.
Published: (2024)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
by: Pungeršek, Taja Kuzman, et al.
Published: (2026)
Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key
by: Chen, Yingda, et al.
Published: (2024)
by: Chen, Yingda, et al.
Published: (2024)
Transcending the Attention Paradigm: Representation Learning from Geospatial Social Media Data
by: DiSanto, Nick, et al.
Published: (2023)
by: DiSanto, Nick, et al.
Published: (2023)
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
by: Huang, Chau-Jian
Published: (2024)
by: Huang, Chau-Jian
Published: (2024)
ThreatCrawl: A BERT-based Focused Crawler for the Cybersecurity Domain
by: Kuehn, Philipp, et al.
Published: (2023)
by: Kuehn, Philipp, et al.
Published: (2023)
Selective Annotation via Data Allocation: These Data Should Be Triaged to Experts for Annotation Rather Than the Model
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
Language Model as an Annotator: Unsupervised Context-aware Quality Phrase Generation
by: Zhang, Zhihao, et al.
Published: (2023)
by: Zhang, Zhihao, et al.
Published: (2023)
EduCoder: An Open-Source Annotation System for Education Transcript Data
by: Ashraf, Saad, et al.
Published: (2025)
by: Ashraf, Saad, et al.
Published: (2025)
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data
by: Wang, Yudong, et al.
Published: (2025)
by: Wang, Yudong, et al.
Published: (2025)
Similar Items
-
Quantifying Geospatial in the Common Crawl Corpus
by: Ilyankou, Ilya, et al.
Published: (2024) -
Much of Geospatial Web Search Is Beyond Traditional GIS
by: Ilyankou, Ilya, et al.
Published: (2026) -
Do Sentence Transformers Learn Quasi-Geospatial Concepts from General Text?
by: Ilyankou, Ilya, et al.
Published: (2024) -
The Scenic Route to Deception: Dark Patterns and Explainability Pitfalls in Conversational Navigation
by: Ilyankou, Ilya, et al.
Published: (2026) -
CycleTrajectory: An End-to-End Pipeline for Enriching and Analyzing GPS Trajectories to Understand Cycling Behavior and Environment
by: Wang, Meihui, et al.
Published: (2024)