Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Karch, Tristan, Engel, Luca, Schwaller, Philippe, Kaplan, Frédéric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
von: Milička, Jiří, et al.
Veröffentlicht: (2025)
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026)
MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction
von: Nasser, Mahdi, et al.
Veröffentlicht: (2025)
von: Nasser, Mahdi, et al.
Veröffentlicht: (2025)
Bias in News Summarization: Measures, Pitfalls and Corpora
von: Steen, Julius, et al.
Veröffentlicht: (2023)
von: Steen, Julius, et al.
Veröffentlicht: (2023)
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
von: Tang, Raphael, et al.
Veröffentlicht: (2024)
von: Tang, Raphael, et al.
Veröffentlicht: (2024)
Multilingual and Explainable Text Detoxification with Parallel Corpora
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
CAMEO: Collection of Multilingual Emotional Speech Corpora
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
von: Christop, Iwona, et al.
Veröffentlicht: (2025)
LLM Agents for Interactive Exploration of Historical Cadastre Data: Framework and Application to Venice
von: Karch, Tristan, et al.
Veröffentlicht: (2025)
von: Karch, Tristan, et al.
Veröffentlicht: (2025)
SimplifyMyText: An LLM-Based System for Inclusive Plain Language Text Simplification
von: Färber, Michael, et al.
Veröffentlicht: (2025)
von: Färber, Michael, et al.
Veröffentlicht: (2025)
Discovering Multi-Scale Semantic Structure in Text Corpora Using Density-Based Trees and LLM Embeddings
von: Haschka, Thomas, et al.
Veröffentlicht: (2025)
von: Haschka, Thomas, et al.
Veröffentlicht: (2025)
From Words to Worth: Newborn Article Impact Prediction with LLM
von: Zhao, Penghai, et al.
Veröffentlicht: (2024)
von: Zhao, Penghai, et al.
Veröffentlicht: (2024)
Readability Measures and Automatic Text Simplification: In the Search of a Construct
von: Cardon, Rémi, et al.
Veröffentlicht: (2025)
von: Cardon, Rémi, et al.
Veröffentlicht: (2025)
Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora
von: Urchs, Stefanie, et al.
Veröffentlicht: (2025)
von: Urchs, Stefanie, et al.
Veröffentlicht: (2025)
Deep Active Learning for Data Mining from Conflict Text Corpora
von: Croicu, Mihai
Veröffentlicht: (2024)
von: Croicu, Mihai
Veröffentlicht: (2024)
From Raw Corpora to Domain Benchmarks: Automated Evaluation of LLM Domain Expertise
von: Sharma, Nitin, et al.
Veröffentlicht: (2025)
von: Sharma, Nitin, et al.
Veröffentlicht: (2025)
Policy-based Sentence Simplification: Replacing Parallel Corpora with LLM-as-a-Judge
von: Wu, Xuanxin, et al.
Veröffentlicht: (2025)
von: Wu, Xuanxin, et al.
Veröffentlicht: (2025)
Individual Text Corpora Predict Openness, Interests, Knowledge and Level of Education
von: Hofmann, Markus J., et al.
Veröffentlicht: (2024)
von: Hofmann, Markus J., et al.
Veröffentlicht: (2024)
Leveraging Large Language Models to Measure Gender Representation Bias in Gendered Language Corpora
von: Derner, Erik, et al.
Veröffentlicht: (2024)
von: Derner, Erik, et al.
Veröffentlicht: (2024)
AlignAR: Generative Sentence Alignment for Arabic-English Parallel Corpora of Legal and Literary Texts
von: Huang, Baorong, et al.
Veröffentlicht: (2025)
von: Huang, Baorong, et al.
Veröffentlicht: (2025)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
von: Kaplan, Guy, et al.
Veröffentlicht: (2025)
von: Kaplan, Guy, et al.
Veröffentlicht: (2025)
LLM-Augmented Chemical Synthesis and Design Decision Programs
von: Wang, Haorui, et al.
Veröffentlicht: (2025)
von: Wang, Haorui, et al.
Veröffentlicht: (2025)
My LLM might Mimic AAE -- But When Should it?
von: Sandoval, Sandra C., et al.
Veröffentlicht: (2025)
von: Sandoval, Sandra C., et al.
Veröffentlicht: (2025)
Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval
von: Liu, Ze, et al.
Veröffentlicht: (2025)
von: Liu, Ze, et al.
Veröffentlicht: (2025)
NLP for Knowledge Discovery and Information Extraction from Energetics Corpora
von: VanGessel, Francis G., et al.
Veröffentlicht: (2024)
von: VanGessel, Francis G., et al.
Veröffentlicht: (2024)
Validating and Exploring Large Geographic Corpora
von: Dunn, Jonathan
Veröffentlicht: (2024)
von: Dunn, Jonathan
Veröffentlicht: (2024)
A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key Tokens
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
von: Nie, Zhijie, et al.
Veröffentlicht: (2024)
Identifying Emerging Concepts in Large Corpora
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
von: Ma, Sibo, et al.
Veröffentlicht: (2025)
Measuring Contextual Informativeness in Child-Directed Text
von: Valentini, Maria, et al.
Veröffentlicht: (2024)
von: Valentini, Maria, et al.
Veröffentlicht: (2024)
Do Language Models Care About Text Quality? Evaluating Web-Crawled Corpora Across 11 Languages
von: van Noord, Rik, et al.
Veröffentlicht: (2024)
von: van Noord, Rik, et al.
Veröffentlicht: (2024)
Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascades
von: Bouchard, Dylan
Veröffentlicht: (2026)
von: Bouchard, Dylan
Veröffentlicht: (2026)
Probing Ethical Framework Representations in Large Language Models: Structure, Entanglement, and Methodological Challenges
von: Xu, Weilun, et al.
Veröffentlicht: (2026)
von: Xu, Weilun, et al.
Veröffentlicht: (2026)
"Be My Cheese?": Assessing Cultural Nuance in Multilingual LLM Translations
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
von: Van Doren, Madison, et al.
Veröffentlicht: (2025)
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
von: Yang, Yi, et al.
Veröffentlicht: (2024)
von: Yang, Yi, et al.
Veröffentlicht: (2024)
Hey AI Can You Grade My Essay?: Automatic Essay Grading
von: Maliha, Maisha, et al.
Veröffentlicht: (2024)
von: Maliha, Maisha, et al.
Veröffentlicht: (2024)
Comparable Corpora: Opportunities for New Research Directions
von: Church, Kenneth
Veröffentlicht: (2025)
von: Church, Kenneth
Veröffentlicht: (2025)
New Textual Corpora for Serbian Language Modeling
von: Škorić, Mihailo, et al.
Veröffentlicht: (2024)
von: Škorić, Mihailo, et al.
Veröffentlicht: (2024)
TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora
von: Kargupta, Priyanka, et al.
Veröffentlicht: (2025)
von: Kargupta, Priyanka, et al.
Veröffentlicht: (2025)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
von: Pungeršek, Taja Kuzman, et al.
Veröffentlicht: (2026)
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
von: Arvin, Chuck
Veröffentlicht: (2025)
von: Arvin, Chuck
Veröffentlicht: (2025)
Mathematical Entities: Corpora and Benchmarks
von: Collard, Jacob, et al.
Veröffentlicht: (2024)
von: Collard, Jacob, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
von: Milička, Jiří, et al.
Veröffentlicht: (2025) -
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
von: Kersting, Nicholas S., et al.
Veröffentlicht: (2026) -
MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction
von: Nasser, Mahdi, et al.
Veröffentlicht: (2025) -
Bias in News Summarization: Measures, Pitfalls and Corpora
von: Steen, Julius, et al.
Veröffentlicht: (2023) -
Words Worth a Thousand Pictures: Measuring and Understanding Perceptual Variability in Text-to-Image Generation
von: Tang, Raphael, et al.
Veröffentlicht: (2024)