The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
Fuente:
arXiv
Saved in:
| Main Authors: | Lequeu, Pierre-Antoine, Labat, Léo, Cave, Laurène, Lejeune, Gaël, Yvon, François, Piwowarski, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026)
by: Labat, Léo, et al.
Published: (2026)
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025)
by: Bouthors, Maxime, et al.
Published: (2025)
ECLAIR: Enhanced Clarification for Interactive Responses in an Enterprise AI Assistant
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
by: Yuan, Weikang, et al.
Published: (2025)
by: Yuan, Weikang, et al.
Published: (2025)
ECLAIR: Enhanced Clarification for Interactive Responses
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
AI Appeals Processor: A Deep Learning Approach to Automated Classification of Citizen Appeals in Government Services
by: Beskorovainyi, Vladimir
Published: (2026)
by: Beskorovainyi, Vladimir
Published: (2026)
GUMBridge: a Corpus for Varieties of Bridging Anaphora
by: Levine, Lauren, et al.
Published: (2025)
by: Levine, Lauren, et al.
Published: (2025)
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
by: Grofsky, Matthew
Published: (2025)
by: Grofsky, Matthew
Published: (2025)
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions
by: Pan, Huitong, et al.
Published: (2024)
by: Pan, Huitong, et al.
Published: (2024)
Corpus Considerations for Annotator Modeling and Scaling
by: Sarumi, Olufunke O., et al.
Published: (2024)
by: Sarumi, Olufunke O., et al.
Published: (2024)
ELCC: the Emergent Language Corpus Collection
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus
by: Boratyn, Daria, et al.
Published: (2026)
by: Boratyn, Daria, et al.
Published: (2026)
IWLV-Ramayana: A Sarga-Aligned Parallel Corpus of Valmiki's Ramayana Across Indian Languages
by: VP, Sumesh
Published: (2026)
by: VP, Sumesh
Published: (2026)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
CorpusStudio: Surfacing Emergent Patterns in a Corpus of Prior Work while Writing
by: Dang, Hai, et al.
Published: (2025)
by: Dang, Hai, et al.
Published: (2025)
AI-assisted German Employment Contract Review: A Benchmark Dataset
by: Wardas, Oliver, et al.
Published: (2025)
by: Wardas, Oliver, et al.
Published: (2025)
Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus
by: Alzetta, Chiara, et al.
Published: (2025)
by: Alzetta, Chiara, et al.
Published: (2025)
PathBench: Speech Intelligibility Benchmark for Automatic Pathological Speech Assessment
by: Halpern, Bence Mark, et al.
Published: (2026)
by: Halpern, Bence Mark, et al.
Published: (2026)
The cohomological Kudla conjecture for unitary Shimura varieties
by: Greer, François, et al.
Published: (2025)
by: Greer, François, et al.
Published: (2025)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Low-resource neural machine translation with morphological modeling
by: Nzeyimana, Antoine
Published: (2024)
by: Nzeyimana, Antoine
Published: (2024)
Identifying Fairness Issues in Automatically Generated Testing Content
by: Stowe, Kevin, et al.
Published: (2024)
by: Stowe, Kevin, et al.
Published: (2024)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
by: Ge, Danying, et al.
Published: (2025)
by: Ge, Danying, et al.
Published: (2025)
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
by: Signoroni, Edoardo, et al.
Published: (2026)
by: Signoroni, Edoardo, et al.
Published: (2026)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
by: Costa-jussà, Marta R., et al.
Published: (2024)
by: Costa-jussà, Marta R., et al.
Published: (2024)
Physical oceanography during Walther Herwig I cruise WH27
by: Anonymous
Published: (2011)
by: Anonymous
Published: (2011)
Primary production within the euphotic layer of the Norwegian Sea in July 1997: measurements by different methods
by: Sapozhnikov, Victor V, et al.
Published: (2000)
by: Sapozhnikov, Victor V, et al.
Published: (2000)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
by: DeRenzi, Brian, et al.
Published: (2025)
by: DeRenzi, Brian, et al.
Published: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
Hydrochemistry measured on water bottle samples during Professor Kolesnikov cruise PK27
by: Piontkovski, Sergey, et al.
Published: (2011)
by: Piontkovski, Sergey, et al.
Published: (2011)
Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
Automatic Prediction of the Performance of Every Parser
by: Biçici, Ergun
Published: (2024)
by: Biçici, Ergun
Published: (2024)
Historical Ink: 19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction
by: Manrique-Gómez, Laura, et al.
Published: (2024)
by: Manrique-Gómez, Laura, et al.
Published: (2024)
Clinical Document Corpora -- Real Ones, Translated and Synthetic Substitutes, and Assorted Domain Proxies: A Survey of Diversity in Corpus Design, with Focus on German Text Data
by: Hahn, Udo
Published: (2024)
by: Hahn, Udo
Published: (2024)
Sharp threshold for the ballisticity of the random walk on the exclusion process
by: Conchon--Kerjan, Guillaume, et al.
Published: (2024)
by: Conchon--Kerjan, Guillaume, et al.
Published: (2024)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
by: Tran, Hung, et al.
Published: (2026)
by: Tran, Hung, et al.
Published: (2026)
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Similar Items
-
Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encoders
by: Lequeu, Pierre-Antoine, et al.
Published: (2026) -
Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions
by: Labat, Léo, et al.
Published: (2026) -
Improving Retrieval-Augmented Neural Machine Translation with Monolingual Data
by: Bouthors, Maxime, et al.
Published: (2025) -
ECLAIR: Enhanced Clarification for Interactive Responses in an Enterprise AI Assistant
by: Murzaku, John, et al.
Published: (2025) -
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)