Current State in Privacy-Preserving Text Preprocessing for Domain-Agnostic NLP
Fuente:
arXiv
Saved in:
| Main Authors: | Sinha, Abhirup, Saha, Pritilata, Saha, Tithi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Study of Privacy-preserving Language Modeling Approaches
by: Saha, Pritilata, et al.
Published: (2025)
by: Saha, Pritilata, et al.
Published: (2025)
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
by: Kumar, Raghvendra, et al.
Published: (2026)
by: Kumar, Raghvendra, et al.
Published: (2026)
Detecting Statements in Text: A Domain-Agnostic Few-Shot Solution
by: Chausson, Sandrine, et al.
Published: (2024)
by: Chausson, Sandrine, et al.
Published: (2024)
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling
by: Akash, Pritom Saha, et al.
Published: (2025)
by: Akash, Pritom Saha, et al.
Published: (2025)
Text Categorization Can Enhance Domain-Agnostic Stopword Extraction
by: Turki, Houcemeddine, et al.
Published: (2024)
by: Turki, Houcemeddine, et al.
Published: (2024)
SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP
by: Duan, Decheng, et al.
Published: (2025)
by: Duan, Decheng, et al.
Published: (2025)
NLP Workbench: Efficient and Extensible Integration of State-of-the-art Text Mining Tools
by: Yao, Peiran, et al.
Published: (2023)
by: Yao, Peiran, et al.
Published: (2023)
Privacy Evaluation Benchmarks for NLP Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
TACIT: A Target-Agnostic Feature Disentanglement Framework for Cross-Domain Text Classification
by: Song, Rui, et al.
Published: (2023)
by: Song, Rui, et al.
Published: (2023)
Two eyes, Two views, and finally, One summary! Towards Multi-modal Multi-tasking Knowledge-Infused Medical Dialogue Summarization
by: Saha, Anisha, et al.
Published: (2024)
by: Saha, Anisha, et al.
Published: (2024)
MASE: Interpretable NLP Models via Model-Agnostic Saliency Estimation
by: Yang, Zhou, et al.
Published: (2025)
by: Yang, Zhou, et al.
Published: (2025)
Empirical Analysis of the Effect of Context in the Task of Automated Essay Scoring in Transformer-Based Models
by: Chakravarty, Abhirup
Published: (2025)
by: Chakravarty, Abhirup
Published: (2025)
Comparison of Large Language Models for Deployment Requirements
by: Yaman, Alper, et al.
Published: (2025)
by: Yaman, Alper, et al.
Published: (2025)
Knowledge-Aware Self-Correction in Language Models via Structured Memory Graphs
by: Saha, Swayamjit
Published: (2025)
by: Saha, Swayamjit
Published: (2025)
Measuring the Robustness of NLP Models to Domain Shifts
by: Calderon, Nitay, et al.
Published: (2023)
by: Calderon, Nitay, et al.
Published: (2023)
Explainability of Text Processing and Retrieval Methods: A Survey
by: Saha, Sourav, et al.
Published: (2022)
by: Saha, Sourav, et al.
Published: (2022)
Investigating Large Language Models' Linguistic Abilities for Text Preprocessing
by: Braga, Marco, et al.
Published: (2025)
by: Braga, Marco, et al.
Published: (2025)
Evolutionary Feature-wise Thresholding for Binary Representation of NLP Embeddings
by: Sinha, Soumen, et al.
Published: (2025)
by: Sinha, Soumen, et al.
Published: (2025)
Enhancing Short-Text Topic Modeling with LLM-Driven Context Expansion and Prefix-Tuned VAEs
by: Akash, Pritom Saha, et al.
Published: (2024)
by: Akash, Pritom Saha, et al.
Published: (2024)
When does MAML Work the Best? An Empirical Study on Model-Agnostic Meta-Learning in NLP Applications
by: Liu, Zequn, et al.
Published: (2020)
by: Liu, Zequn, et al.
Published: (2020)
Transparent NLP: Using RAG and LLM Alignment for Privacy Q&A
by: Leschanowsky, Anna, et al.
Published: (2025)
by: Leschanowsky, Anna, et al.
Published: (2025)
NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey
by: Goswami, Dhiman, et al.
Published: (2026)
by: Goswami, Dhiman, et al.
Published: (2026)
How Does A Text Preprocessing Pipeline Affect Ontology Matching?
by: Qiang, Zhangcheng, et al.
Published: (2024)
by: Qiang, Zhangcheng, et al.
Published: (2024)
WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
Prompting Away Stereotypes? Evaluating Bias in Text-to-Image Models for Occupations
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Synthesizing Privacy-Preserving Text Data via Finetuning without Finetuning Billion-Scale LLMs
by: Tan, Bowen, et al.
Published: (2025)
by: Tan, Bowen, et al.
Published: (2025)
NAP^2: A Benchmark for Naturalness and Privacy-Preserving Text Rewriting by Learning from Human
by: Huang, Shuo, et al.
Published: (2024)
by: Huang, Shuo, et al.
Published: (2024)
Anonymous-by-Construction: An LLM-Driven Framework for Privacy-Preserving Text
by: Albanese, Federico, et al.
Published: (2026)
by: Albanese, Federico, et al.
Published: (2026)
Transforming Sensitive Documents into Quantitative Data: An AI-Based Preprocessing Toolchain for Structured and Privacy-Conscious Analysis
by: Ledberg, Anders, et al.
Published: (2025)
by: Ledberg, Anders, et al.
Published: (2025)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
by: Saha, Binita, et al.
Published: (2025)
by: Saha, Binita, et al.
Published: (2025)
ParsiPy: NLP Toolkit for Historical Persian Texts in Python
by: Farsi, Farhan, et al.
Published: (2025)
by: Farsi, Farhan, et al.
Published: (2025)
State-of-the-art generalisation research in NLP: A taxonomy and review
by: Hupkes, Dieuwke, et al.
Published: (2022)
by: Hupkes, Dieuwke, et al.
Published: (2022)
Evaluation Metrics for Text Data Augmentation in NLP
by: Amadeus, Marcellus, et al.
Published: (2024)
by: Amadeus, Marcellus, et al.
Published: (2024)
1-Diffractor: Efficient and Utility-Preserving Text Obfuscation Leveraging Word-Level Metric Differential Privacy
by: Meisenbacher, Stephen, et al.
Published: (2024)
by: Meisenbacher, Stephen, et al.
Published: (2024)
An EcoSage Assistant: Towards Building A Multimodal Plant Care Dialogue Assistant
by: Tomar, Mohit, et al.
Published: (2024)
by: Tomar, Mohit, et al.
Published: (2024)
One Word is Enough: Minimal Adversarial Perturbations for Neural Text Ranking
by: Karmakar, Tanmay, et al.
Published: (2026)
by: Karmakar, Tanmay, et al.
Published: (2026)
ATEB: Evaluating and Improving Advanced NLP Tasks for Text Embedding Models
by: Han, Simeng, et al.
Published: (2025)
by: Han, Simeng, et al.
Published: (2025)
State of NLP in Kenya: A Survey
by: Amol, Cynthia Jayne, et al.
Published: (2024)
by: Amol, Cynthia Jayne, et al.
Published: (2024)
User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs
by: Saha, Sougata, et al.
Published: (2025)
by: Saha, Sougata, et al.
Published: (2025)
Similar Items
-
A Study of Privacy-preserving Language Modeling Approaches
by: Saha, Pritilata, et al.
Published: (2025) -
BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources
by: Kumar, Raghvendra, et al.
Published: (2026) -
Detecting Statements in Text: A Domain-Agnostic Few-Shot Solution
by: Chausson, Sandrine, et al.
Published: (2024) -
Understanding Cross-Domain Adaptation in Low-Resource Topic Modeling
by: Akash, Pritom Saha, et al.
Published: (2025) -
Text Categorization Can Enhance Domain-Agnostic Stopword Extraction
by: Turki, Houcemeddine, et al.
Published: (2024)