Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Meisenbacher, Stephen, Nestorov, Svetlozar, Norlander, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
por: Meisenbacher, Stephen, et al.
Publicado: (2026)
por: Meisenbacher, Stephen, et al.
Publicado: (2026)
Towards A Structured Overview of Use Cases for Natural Language Processing in the Legal Domain: A German Perspective
por: Vladika, Juraj, et al.
Publicado: (2024)
por: Vladika, Juraj, et al.
Publicado: (2024)
Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
por: Ahnert, Georg, et al.
Publicado: (2024)
por: Ahnert, Georg, et al.
Publicado: (2024)
Applying BioBERT to Extract Germline Gene-Disease Associations for Building a Knowledge Graph from the Biomedical Literature
por: Gonzalez, Armando D. Diaz, et al.
Publicado: (2023)
por: Gonzalez, Armando D. Diaz, et al.
Publicado: (2023)
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
por: Olufemi, Victor Tolulope, et al.
Publicado: (2026)
por: Olufemi, Victor Tolulope, et al.
Publicado: (2026)
With Privacy, Size Matters: On the Importance of Dataset Size in Differentially Private Text Rewriting
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten Text
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
por: Litterer, Benjamin, et al.
Publicado: (2024)
por: Litterer, Benjamin, et al.
Publicado: (2024)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
por: Ousidhoum, Nedjma, et al.
Publicado: (2024)
por: Ousidhoum, Nedjma, et al.
Publicado: (2024)
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
por: Pham, Nhi, et al.
Publicado: (2024)
por: Pham, Nhi, et al.
Publicado: (2024)
The Moral Foundations Reddit Corpus
por: Trager, Jackson, et al.
Publicado: (2022)
por: Trager, Jackson, et al.
Publicado: (2022)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
por: Berriche, Manon, et al.
Publicado: (2025)
por: Berriche, Manon, et al.
Publicado: (2025)
Industry Risk Assessment via Hierarchical Financial Data Using Stock Market Sentiment Indicators
por: Zhu, Hongyin
Publicado: (2023)
por: Zhu, Hongyin
Publicado: (2023)
The Cambridge Law Corpus: A Dataset for Legal AI Research
por: Östling, Andreas, et al.
Publicado: (2023)
por: Östling, Andreas, et al.
Publicado: (2023)
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
por: Fillies, Jan, et al.
Publicado: (2025)
por: Fillies, Jan, et al.
Publicado: (2025)
Corpus-Based Approaches to Igbo Diacritic Restoration
por: Ezeani, Ignatius
Publicado: (2026)
por: Ezeani, Ignatius
Publicado: (2026)
ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging
por: Nacar, Omer, et al.
Publicado: (2026)
por: Nacar, Omer, et al.
Publicado: (2026)
Using Twitter Data to Understand Public Perceptions of Approved versus Off-label Use for COVID-19-related Medications
por: Hua, Yining, et al.
Publicado: (2022)
por: Hua, Yining, et al.
Publicado: (2022)
Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
Lexical Substitution is not Synonym Substitution: On the Importance of Producing Contextually Relevant Word Substitutes
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
On the Impact of Noise in Differentially Private Text Rewriting
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
por: Meisenbacher, Stephen, et al.
Publicado: (2025)
A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation
por: Meisenbacher, Stephen, et al.
Publicado: (2026)
por: Meisenbacher, Stephen, et al.
Publicado: (2026)
A Collocation-based Method for Addressing Challenges in Word-level Metric Differential Privacy
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
por: Meisenbacher, Stephen, et al.
Publicado: (2024)
A Thematic Framework for Analyzing Large-scale Self-reported Social Media Data on Opioid Use Disorder Treatment Using Buprenorphine Product
por: Basak, Madhusudan, et al.
Publicado: (2024)
por: Basak, Madhusudan, et al.
Publicado: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
por: Karamolegkou, Antonia, et al.
Publicado: (2024)
por: Karamolegkou, Antonia, et al.
Publicado: (2024)
Computational Studies in Influencer Marketing: A Systematic Literature Review
por: Gui, Haoyang, et al.
Publicado: (2025)
por: Gui, Haoyang, et al.
Publicado: (2025)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
por: Tasnim, Nazia, et al.
Publicado: (2024)
por: Tasnim, Nazia, et al.
Publicado: (2024)
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
por: Frazzetto, Paolo, et al.
Publicado: (2025)
por: Frazzetto, Paolo, et al.
Publicado: (2025)
Extracting memorized pieces of (copyrighted) books from open-weight language models
por: Cooper, A. Feder, et al.
Publicado: (2025)
por: Cooper, A. Feder, et al.
Publicado: (2025)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
por: Yilmaz, Berk, et al.
Publicado: (2025)
por: Yilmaz, Berk, et al.
Publicado: (2025)
REInstruct: Building Instruction Data from Unlabeled Corpus
por: Chen, Shu, et al.
Publicado: (2024)
por: Chen, Shu, et al.
Publicado: (2024)
Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
por: Filimonovic, Dragan, et al.
Publicado: (2025)
por: Filimonovic, Dragan, et al.
Publicado: (2025)
Sovereign AI-based Public Services are Viable and Affordable
por: Branco, António, et al.
Publicado: (2026)
por: Branco, António, et al.
Publicado: (2026)
Evaluating LLM-Generated Legal Explanations for Regulatory Compliance in Social Media Influencer Marketing
por: Gui, Haoyang, et al.
Publicado: (2025)
por: Gui, Haoyang, et al.
Publicado: (2025)
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
por: Thomas, Danielle R., et al.
Publicado: (2025)
por: Thomas, Danielle R., et al.
Publicado: (2025)
Dual Use Concerns of Generative AI and Large Language Models
por: Grinbaum, Alexei, et al.
Publicado: (2023)
por: Grinbaum, Alexei, et al.
Publicado: (2023)
Measuring the Gap Between Media Coverage and Public Information Demand: Evidence from the 2026 Lebanon Conflict
por: Soufan, Mohamed
Publicado: (2026)
por: Soufan, Mohamed
Publicado: (2026)
Ejemplares similares
-
Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
por: Meisenbacher, Stephen, et al.
Publicado: (2026) -
Towards A Structured Overview of Use Cases for Natural Language Processing in the Legal Domain: A German Perspective
por: Vladika, Juraj, et al.
Publicado: (2024) -
Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
por: Ahnert, Georg, et al.
Publicado: (2024) -
Applying BioBERT to Extract Germline Gene-Disease Associations for Building a Knowledge Graph from the Biomedical Literature
por: Gonzalez, Armando D. Diaz, et al.
Publicado: (2023) -
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
por: Olufemi, Victor Tolulope, et al.
Publicado: (2026)