AI for Monitoring and Classifying Data Used in Research Literature
Fuente:
arXiv
Saved in:
| Main Authors: | Macalaba, Rafael, Solatorio, Aivin V. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models and Synthetic Data for Monitoring Dataset Mentions in Research Papers
by: Solatorio, Aivin V., et al.
Published: (2025)
by: Solatorio, Aivin V., et al.
Published: (2025)
GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning
by: Solatorio, Aivin V.
Published: (2024)
by: Solatorio, Aivin V.
Published: (2024)
Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification
by: Solatorio, Aivin V.
Published: (2025)
by: Solatorio, Aivin V.
Published: (2025)
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers
by: Solatorio, Aivin V., et al.
Published: (2024)
by: Solatorio, Aivin V., et al.
Published: (2024)
Modelling and Classifying the Components of a Literature Review
by: Bolaños, Francisco, et al.
Published: (2025)
by: Bolaños, Francisco, et al.
Published: (2025)
Leveraging LLMs for Translating and Classifying Mental Health Data
by: Skianis, Konstantinos, et al.
Published: (2024)
by: Skianis, Konstantinos, et al.
Published: (2024)
Could the Road to Grounded, Neuro-symbolic AI be Paved with Words-as-Classifiers?
by: Kennington, Casey, et al.
Published: (2025)
by: Kennington, Casey, et al.
Published: (2025)
LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection
by: Wu, Jian, et al.
Published: (2025)
by: Wu, Jian, et al.
Published: (2025)
Train a Unified Multimodal Data Quality Classifier with Synthetic Data
by: Wang, Weizhi, et al.
Published: (2025)
by: Wang, Weizhi, et al.
Published: (2025)
Re-defining Humor Data Objects for AI Humor Research
by: Arnett, Anna, et al.
Published: (2026)
by: Arnett, Anna, et al.
Published: (2026)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
by: Birkenmaier, Lukas, et al.
Published: (2024)
by: Birkenmaier, Lukas, et al.
Published: (2024)
*-PLUIE: Personalisable metric with Llm Used for Improved Evaluation
by: Lemesle, Quentin, et al.
Published: (2026)
by: Lemesle, Quentin, et al.
Published: (2026)
Detecting Turkish Synonyms Used in Different Time Periods
by: Yazar, Umur Togay, et al.
Published: (2024)
by: Yazar, Umur Togay, et al.
Published: (2024)
On-Device Emoji Classifier Trained with GPT-based Data Augmentation for a Mobile Keyboard
by: Amer, Hossam, et al.
Published: (2024)
by: Amer, Hossam, et al.
Published: (2024)
Classifying Human-Generated and AI-Generated Election Claims in Social Media
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
Your Extreme Multi-label Classifier is Secretly a Hierarchical Text Classifier for Free
by: Bertalis, Nerijus, et al.
Published: (2024)
by: Bertalis, Nerijus, et al.
Published: (2024)
Exploring Complex Mental Health Symptoms via Classifying Social Media Data with Explainable LLMs
by: Chen, Kexin, et al.
Published: (2024)
by: Chen, Kexin, et al.
Published: (2024)
CHAIR -- Classifier of Hallucination as Improver
by: Sun, Ao
Published: (2025)
by: Sun, Ao
Published: (2025)
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
by: Pomerenke, David, et al.
Published: (2025)
by: Pomerenke, David, et al.
Published: (2025)
Toward Cross-Lingual Quality Classifiers for Multilingual Pretraining Data Selection
by: Turki, Yassine, et al.
Published: (2026)
by: Turki, Yassine, et al.
Published: (2026)
Chapter 7 Review of Data-Driven Generative AI Models for Knowledge Extraction from Scientific Literature in Healthcare
by: Kopitar, Leon, et al.
Published: (2024)
by: Kopitar, Leon, et al.
Published: (2024)
Classifying several dialectal Nawatl varieties
by: Guzmán-Landa, Juan-José, et al.
Published: (2026)
by: Guzmán-Landa, Juan-José, et al.
Published: (2026)
Automated Adversarial Discovery for Safety Classifiers
by: Lal, Yash Kumar, et al.
Published: (2024)
by: Lal, Yash Kumar, et al.
Published: (2024)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
by: Saada, Thiziri Nait, et al.
Published: (2025)
by: Saada, Thiziri Nait, et al.
Published: (2025)
A Systematic Review of Open Datasets Used in Text-to-Image (T2I) Gen AI Model Safety
by: Rouf, Rakeen, et al.
Published: (2025)
by: Rouf, Rakeen, et al.
Published: (2025)
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026)
by: Miok, Kristian, et al.
Published: (2026)
An Annotation Scheme and Classifier for Personal Facts in Dialogue
by: Zaitsev, Konstantin
Published: (2026)
by: Zaitsev, Konstantin
Published: (2026)
Classifying Unreliable Narrators with Large Language Models
by: Brei, Anneliese, et al.
Published: (2025)
by: Brei, Anneliese, et al.
Published: (2025)
How to Make LMs Strong Node Classifiers?
by: Xu, Zhe, et al.
Published: (2024)
by: Xu, Zhe, et al.
Published: (2024)
Are LLMs Good Zero-Shot Fallacy Classifiers?
by: Pan, Fengjun, et al.
Published: (2024)
by: Pan, Fengjun, et al.
Published: (2024)
Isotropy, Clusters, and Classifiers
by: Mickus, Timothee, et al.
Published: (2024)
by: Mickus, Timothee, et al.
Published: (2024)
Automatic Classifiers Underdetect Emotions Expressed by Men
by: Smirnov, Ivan, et al.
Published: (2026)
by: Smirnov, Ivan, et al.
Published: (2026)
A Network Analysis Approach to Conlang Research Literature
by: Gonzalez, Simon
Published: (2024)
by: Gonzalez, Simon
Published: (2024)
Can Unconfident LLM Annotations Be Used for Confident Conclusions?
by: Gligorić, Kristina, et al.
Published: (2024)
by: Gligorić, Kristina, et al.
Published: (2024)
Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI
by: Sandoval, Aaron, et al.
Published: (2025)
by: Sandoval, Aaron, et al.
Published: (2025)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Cloaked Classifiers: Pseudonymization Strategies on Sensitive Classification Tasks
by: Riabi, Arij, et al.
Published: (2024)
by: Riabi, Arij, et al.
Published: (2024)
LLM with Relation Classifier for Document-Level Relation Extraction
by: Li, Xingzuo, et al.
Published: (2024)
by: Li, Xingzuo, et al.
Published: (2024)
Embedded Named Entity Recognition using Probing Classifiers
by: Popovič, Nicholas, et al.
Published: (2024)
by: Popovič, Nicholas, et al.
Published: (2024)
Similar Items
-
Large Language Models and Synthetic Data for Monitoring Dataset Mentions in Research Papers
by: Solatorio, Aivin V., et al.
Published: (2025) -
GISTEmbed: Guided In-sample Selection of Training Negatives for Text Embedding Fine-tuning
by: Solatorio, Aivin V.
Published: (2024) -
Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification
by: Solatorio, Aivin V.
Published: (2025) -
Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers
by: Solatorio, Aivin V., et al.
Published: (2024) -
Modelling and Classifying the Components of a Literature Review
by: Bolaños, Francisco, et al.
Published: (2025)