The study of short texts in digital politics: Document aggregation for topic modeling
Fuente:
arXiv
Salvato in:
| Autori principali: | Nakka, Nitheesha, Yalcin, Omer F., Desmarais, Bruce A., Rajtmajer, Sarah, Monroe, Burt |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Human-interpretable clustering of short-text using large language models
di: Miller, Justin K., et al.
Pubblicazione: (2024)
di: Miller, Justin K., et al.
Pubblicazione: (2024)
A Bayesian approach to modeling topic-metadata relationships
di: Schulze, P., et al.
Pubblicazione: (2021)
di: Schulze, P., et al.
Pubblicazione: (2021)
Stylometry recognizes human and LLM-generated texts in short samples
di: Przystalski, Karol, et al.
Pubblicazione: (2025)
di: Przystalski, Karol, et al.
Pubblicazione: (2025)
CFTM: Continuous time fractional topic model
di: Nakagawa, Kei, et al.
Pubblicazione: (2024)
di: Nakagawa, Kei, et al.
Pubblicazione: (2024)
Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
di: Medvecki, Darija, et al.
Pubblicazione: (2024)
di: Medvecki, Darija, et al.
Pubblicazione: (2024)
Machine-generated text detection prevents language model collapse
di: Drayson, George, et al.
Pubblicazione: (2025)
di: Drayson, George, et al.
Pubblicazione: (2025)
PRISM: PRIor from corpus Statistics for topic Modeling
di: Ishon, Tal, et al.
Pubblicazione: (2026)
di: Ishon, Tal, et al.
Pubblicazione: (2026)
Critical biblical studies via word frequency analysis: unveiling text authorship
di: Faigenbaum-Golovin, Shira, et al.
Pubblicazione: (2024)
di: Faigenbaum-Golovin, Shira, et al.
Pubblicazione: (2024)
PrivacyScalpel: Enhancing LLM Privacy via Interpretable Feature Intervention with Sparse Autoencoders
di: Frikha, Ahmed, et al.
Pubblicazione: (2025)
di: Frikha, Ahmed, et al.
Pubblicazione: (2025)
TAGLAS: An atlas of text-attributed graph datasets in the era of large graph and language models
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
di: Feng, Jiarui, et al.
Pubblicazione: (2024)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
Forma mentis networks predict creativity ratings of short texts via interpretable artificial intelligence in human and GPT-simulated raters
di: Haim, Edith, et al.
Pubblicazione: (2024)
di: Haim, Edith, et al.
Pubblicazione: (2024)
Seeded Poisson Factorization: leveraging domain knowledge to fit topic models
di: Prostmaier, Bernd, et al.
Pubblicazione: (2025)
di: Prostmaier, Bernd, et al.
Pubblicazione: (2025)
Silence and Noise: Self-censorship and Opinion Expression on Social Media
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
di: Wang, Xinyu, et al.
Pubblicazione: (2026)
Modeling citation worthiness by using attention-based bidirectional long short-term memory networks and interpretable models
di: Zeng, Tong, et al.
Pubblicazione: (2024)
di: Zeng, Tong, et al.
Pubblicazione: (2024)
Failure Modes of Maximum Entropy RLHF
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
di: Çağatan, Ömer Veysel, et al.
Pubblicazione: (2025)
Document Summarization with Conformal Importance Guarantees
di: Kuwahara, Bruce, et al.
Pubblicazione: (2025)
di: Kuwahara, Bruce, et al.
Pubblicazione: (2025)
LLM-based feature generation from text for interpretable machine learning
di: Balek, Vojtěch, et al.
Pubblicazione: (2024)
di: Balek, Vojtěch, et al.
Pubblicazione: (2024)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
di: Nakka, Krishna Kanth, et al.
Pubblicazione: (2024)
Representative Language Generation
di: Peale, Charlotte, et al.
Pubblicazione: (2025)
di: Peale, Charlotte, et al.
Pubblicazione: (2025)
Detecting out-of-distribution text using topological features of transformer-based language models
di: Pollano, Andres, et al.
Pubblicazione: (2023)
di: Pollano, Andres, et al.
Pubblicazione: (2023)
Extractive text summarisation of Privacy Policy documents using machine learning approaches
di: Choi, Chanwoo
Pubblicazione: (2024)
di: Choi, Chanwoo
Pubblicazione: (2024)
AIDetx: a compression-based method for identification of machine-learning generated text
di: Almeida, Leonardo, et al.
Pubblicazione: (2024)
di: Almeida, Leonardo, et al.
Pubblicazione: (2024)
Cropping outperforms dropout as an augmentation strategy for self-supervised training of text embeddings
di: González-Márquez, Rita, et al.
Pubblicazione: (2025)
di: González-Márquez, Rita, et al.
Pubblicazione: (2025)
Large Language Model Augmented Exercise Retrieval for Personalized Language Learning
di: Xu, Austin, et al.
Pubblicazione: (2024)
di: Xu, Austin, et al.
Pubblicazione: (2024)
Discovering influential text using convolutional neural networks
di: Ayers, Megan, et al.
Pubblicazione: (2024)
di: Ayers, Megan, et al.
Pubblicazione: (2024)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
di: Gao, Peiheng, et al.
Pubblicazione: (2025)
di: Gao, Peiheng, et al.
Pubblicazione: (2025)
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
di: Gatto, Joseph, et al.
Pubblicazione: (2024)
di: Gatto, Joseph, et al.
Pubblicazione: (2024)
Combining topic modelling and citation network analysis to study case law from the European Court on Human Rights on the right to respect for private and family life
di: Mohammadi, M., et al.
Pubblicazione: (2024)
di: Mohammadi, M., et al.
Pubblicazione: (2024)
ChatGPT for automated grading of short answer questions in mechanical ventilation
di: Jade, Tejas, et al.
Pubblicazione: (2025)
di: Jade, Tejas, et al.
Pubblicazione: (2025)
AlleNoise: large-scale text classification benchmark dataset with real-world label noise
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
di: Rączkowska, Alicja, et al.
Pubblicazione: (2024)
A quantitative analysis of semantic information in deep representations of text and images
di: Acevedo, Santiago, et al.
Pubblicazione: (2025)
di: Acevedo, Santiago, et al.
Pubblicazione: (2025)
English offensive text detection using CNN based Bi-GRU model
di: Roy, Tonmoy, et al.
Pubblicazione: (2024)
di: Roy, Tonmoy, et al.
Pubblicazione: (2024)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
di: Shingi, Geet, et al.
Pubblicazione: (2021)
di: Shingi, Geet, et al.
Pubblicazione: (2021)
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification
di: Dubey, Kush
Pubblicazione: (2024)
di: Dubey, Kush
Pubblicazione: (2024)
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
IncogniText: Privacy-enhancing Conditional Text Anonymization via LLM-based Private Attribute Randomization
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
di: Frikha, Ahmed, et al.
Pubblicazione: (2024)
Phase transition on a context-sensitive random language model with short range interactions
di: Toji, Yuma, et al.
Pubblicazione: (2026)
di: Toji, Yuma, et al.
Pubblicazione: (2026)
Clinical ModernBERT: An efficient and long context encoder for biomedical text
di: Lee, Simon A., et al.
Pubblicazione: (2025)
di: Lee, Simon A., et al.
Pubblicazione: (2025)
BP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
di: Li, Fengyi, et al.
Pubblicazione: (2025)
di: Li, Fengyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Human-interpretable clustering of short-text using large language models
di: Miller, Justin K., et al.
Pubblicazione: (2024) -
A Bayesian approach to modeling topic-metadata relationships
di: Schulze, P., et al.
Pubblicazione: (2021) -
Stylometry recognizes human and LLM-generated texts in short samples
di: Przystalski, Karol, et al.
Pubblicazione: (2025) -
CFTM: Continuous time fractional topic model
di: Nakagawa, Kei, et al.
Pubblicazione: (2024) -
Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
di: Medvecki, Darija, et al.
Pubblicazione: (2024)