Extracting O*NET Features from the NLx Corpus to Build Public Use Aggregate Labor Market Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meisenbacher, Stephen, Nestorov, Svetlozar, Norlander, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2026)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2026)
Towards A Structured Overview of Use Cases for Natural Language Processing in the Legal Domain: A German Perspective
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
von: Ahnert, Georg, et al.
Veröffentlicht: (2024)
von: Ahnert, Georg, et al.
Veröffentlicht: (2024)
Applying BioBERT to Extract Germline Gene-Disease Associations for Building a Knowledge Graph from the Biomedical Literature
von: Gonzalez, Armando D. Diaz, et al.
Veröffentlicht: (2023)
von: Gonzalez, Armando D. Diaz, et al.
Veröffentlicht: (2023)
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
von: Olufemi, Victor Tolulope, et al.
Veröffentlicht: (2026)
von: Olufemi, Victor Tolulope, et al.
Veröffentlicht: (2026)
With Privacy, Size Matters: On the Importance of Dataset Size in Differentially Private Text Rewriting
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten Text
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus
von: Litterer, Benjamin, et al.
Veröffentlicht: (2024)
von: Litterer, Benjamin, et al.
Veröffentlicht: (2024)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Building Better: Avoiding Pitfalls in Developing Language Resources when Data is Scarce
von: Ousidhoum, Nedjma, et al.
Veröffentlicht: (2024)
von: Ousidhoum, Nedjma, et al.
Veröffentlicht: (2024)
Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties
von: Pham, Nhi, et al.
Veröffentlicht: (2024)
von: Pham, Nhi, et al.
Veröffentlicht: (2024)
The Moral Foundations Reddit Corpus
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
von: Trager, Jackson, et al.
Veröffentlicht: (2022)
SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversations
von: Berriche, Manon, et al.
Veröffentlicht: (2025)
von: Berriche, Manon, et al.
Veröffentlicht: (2025)
Industry Risk Assessment via Hierarchical Financial Data Using Stock Market Sentiment Indicators
von: Zhu, Hongyin
Veröffentlicht: (2023)
von: Zhu, Hongyin
Veröffentlicht: (2023)
The Cambridge Law Corpus: A Dataset for Legal AI Research
von: Östling, Andreas, et al.
Veröffentlicht: (2023)
von: Östling, Andreas, et al.
Veröffentlicht: (2023)
Mapping Toxic Comments Across Demographics: A Dataset from German Public Broadcasting
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
von: Fillies, Jan, et al.
Veröffentlicht: (2025)
Corpus-Based Approaches to Igbo Diacritic Restoration
von: Ezeani, Ignatius
Veröffentlicht: (2026)
von: Ezeani, Ignatius
Veröffentlicht: (2026)
ARCADE: A City-Scale Corpus for Fine-Grained Arabic Dialect Tagging
von: Nacar, Omer, et al.
Veröffentlicht: (2026)
von: Nacar, Omer, et al.
Veröffentlicht: (2026)
Using Twitter Data to Understand Public Perceptions of Approved versus Off-label Use for COVID-19-related Medications
von: Hua, Yining, et al.
Veröffentlicht: (2022)
von: Hua, Yining, et al.
Veröffentlicht: (2022)
Leveraging Semantic Triples for Private Document Generation with Local Differential Privacy Guarantees
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
Lexical Substitution is not Synonym Substitution: On the Importance of Producing Contextually Relevant Word Substitutes
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
von: Vladika, Juraj, et al.
Veröffentlicht: (2025)
On the Impact of Noise in Differentially Private Text Rewriting
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2025)
A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2026)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2026)
A Collocation-based Method for Addressing Challenges in Word-level Metric Differential Privacy
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2024)
A Thematic Framework for Analyzing Large-scale Self-reported Social Media Data on Opioid Use Disorder Treatment Using Buprenorphine Product
von: Basak, Madhusudan, et al.
Veröffentlicht: (2024)
von: Basak, Madhusudan, et al.
Veröffentlicht: (2024)
Ethical Concern Identification in NLP: A Corpus of ACL Anthology Ethics Statements
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2024)
von: Karamolegkou, Antonia, et al.
Veröffentlicht: (2024)
Computational Studies in Influencer Marketing: A Systematic Literature Review
von: Gui, Haoyang, et al.
Veröffentlicht: (2025)
von: Gui, Haoyang, et al.
Veröffentlicht: (2025)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
von: Tasnim, Nazia, et al.
Veröffentlicht: (2024)
von: Tasnim, Nazia, et al.
Veröffentlicht: (2024)
From Text to Talent: A Pipeline for Extracting Insights from Candidate Profiles
von: Frazzetto, Paolo, et al.
Veröffentlicht: (2025)
von: Frazzetto, Paolo, et al.
Veröffentlicht: (2025)
Extracting memorized pieces of (copyrighted) books from open-weight language models
von: Cooper, A. Feder, et al.
Veröffentlicht: (2025)
von: Cooper, A. Feder, et al.
Veröffentlicht: (2025)
Towards Equitable AI: Detecting Bias in Using Large Language Models for Marketing
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
von: Yilmaz, Berk, et al.
Veröffentlicht: (2025)
REInstruct: Building Instruction Data from Unlabeled Corpus
von: Chen, Shu, et al.
Veröffentlicht: (2024)
von: Chen, Shu, et al.
Veröffentlicht: (2024)
Does Scientific Writing Converge to U.S. English? Evidence from Generative AI-Assisted Publications
von: Filimonovic, Dragan, et al.
Veröffentlicht: (2025)
von: Filimonovic, Dragan, et al.
Veröffentlicht: (2025)
Sovereign AI-based Public Services are Viable and Affordable
von: Branco, António, et al.
Veröffentlicht: (2026)
von: Branco, António, et al.
Veröffentlicht: (2026)
Evaluating LLM-Generated Legal Explanations for Regulatory Compliance in Social Media Influencer Marketing
von: Gui, Haoyang, et al.
Veröffentlicht: (2025)
von: Gui, Haoyang, et al.
Veröffentlicht: (2025)
LLM-Generated Feedback Supports Learning If Learners Choose to Use It
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
von: Thomas, Danielle R., et al.
Veröffentlicht: (2025)
Dual Use Concerns of Generative AI and Large Language Models
von: Grinbaum, Alexei, et al.
Veröffentlicht: (2023)
von: Grinbaum, Alexei, et al.
Veröffentlicht: (2023)
Measuring the Gap Between Media Coverage and Public Information Demand: Evidence from the 2026 Lebanon Conflict
von: Soufan, Mohamed
Veröffentlicht: (2026)
von: Soufan, Mohamed
Veröffentlicht: (2026)
Ähnliche Einträge
-
Building a Custom Taxonomy of AI Skills and Tasks from the Ground Up with Job Postings
von: Meisenbacher, Stephen, et al.
Veröffentlicht: (2026) -
Towards A Structured Overview of Use Cases for Natural Language Processing in the Legal Domain: A German Perspective
von: Vladika, Juraj, et al.
Veröffentlicht: (2024) -
Extracting Affect Aggregates from Longitudinal Social Media Data with Temporal Adapters for Large Language Models
von: Ahnert, Georg, et al.
Veröffentlicht: (2024) -
Applying BioBERT to Extract Germline Gene-Disease Associations for Building a Knowledge Graph from the Biomedical Literature
von: Gonzalez, Armando D. Diaz, et al.
Veröffentlicht: (2023) -
WAXAL-NET: Finetuned Edge ASR Across 19 African Languages
von: Olufemi, Victor Tolulope, et al.
Veröffentlicht: (2026)