Identifying Fairness Issues in Automatically Generated Testing Content
Fuente:
arXiv
Saved in:
| Main Authors: | Stowe, Kevin, Longwill, Benny, Francis, Alyssa, Aoyama, Tatsuya, Ghosh, Debanjan, Somasundaran, Swapna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026)
by: Stowe, Kevin, et al.
Published: (2026)
Identifying Bias in Machine-generated Text Detection
by: Stowe, Kevin, et al.
Published: (2025)
by: Stowe, Kevin, et al.
Published: (2025)
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025)
by: Dejl, Adam, et al.
Published: (2025)
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
by: Gupta, Abhay, et al.
Published: (2025)
by: Gupta, Abhay, et al.
Published: (2025)
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025)
by: Yun, Janghyeon, et al.
Published: (2025)
Generator-Guided Crowd Reaction Assessment
by: Ghosh, Sohom, et al.
Published: (2024)
by: Ghosh, Sohom, et al.
Published: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
by: Dugan, Liam, et al.
Published: (2024)
by: Dugan, Liam, et al.
Published: (2024)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
by: Ge, Danying, et al.
Published: (2025)
by: Ge, Danying, et al.
Published: (2025)
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
by: Signoroni, Edoardo, et al.
Published: (2026)
by: Signoroni, Edoardo, et al.
Published: (2026)
Augmenting Dialog with Think-Aloud Utterances for Modeling Individual Personality Traits by LLM
by: Ishikura, Seiya, et al.
Published: (2025)
by: Ishikura, Seiya, et al.
Published: (2025)
Synthetic Voice Data for Automatic Speech Recognition in African Languages
by: DeRenzi, Brian, et al.
Published: (2025)
by: DeRenzi, Brian, et al.
Published: (2025)
Automatic Generation of Conversational Interfaces for Tabular Data Analysis
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
by: Gomez-Vazquez, Marcos, et al.
Published: (2023)
Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
by: Barkhordar, Ehsan, et al.
Published: (2026)
by: Barkhordar, Ehsan, et al.
Published: (2026)
Quality Estimation with $k$-nearest Neighbors and Automatic Evaluation for Model-specific Quality Estimation
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Sensitive Content Classification in Social Media: A Holistic Resource and Evaluation
by: Antypas, Dimosthenis, et al.
Published: (2024)
by: Antypas, Dimosthenis, et al.
Published: (2024)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Breaking the HISCO Barrier: Automatic Occupational Standardization with OccCANINE
by: Dahl, Christian Møller, et al.
Published: (2024)
by: Dahl, Christian Møller, et al.
Published: (2024)
Identifying Intensity of the Structure and Content in Tweets and the Discriminative Power of Attributes in Context with Referential Translation Machines
by: Biçici, Ergun
Published: (2024)
by: Biçici, Ergun
Published: (2024)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
by: Riemenschneider, Frederick, et al.
Published: (2024)
by: Riemenschneider, Frederick, et al.
Published: (2024)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
by: Kugler, Kai
Published: (2025)
by: Kugler, Kai
Published: (2025)
Test-Time Scaling of Reasoning Models for Machine Translation
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
RedHerring Attack: Testing the Reliability of Attack Detection
by: Rusert, Jonathan
Published: (2025)
by: Rusert, Jonathan
Published: (2025)
A Likelihood Ratio Test of Genetic Relationship among Languages
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2024)
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
by: Galvan-Sosa, Diana, et al.
Published: (2025)
by: Galvan-Sosa, Diana, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
COSTAR-A: A prompting framework for enhancing Large Language Model performance on Point-of-View questions
by: Ohalete, Nzubechukwu C., et al.
Published: (2025)
by: Ohalete, Nzubechukwu C., et al.
Published: (2025)
StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis
by: Opara, Chidimma
Published: (2024)
by: Opara, Chidimma
Published: (2024)
What Language Models Know But Don't Say: Non-Generative Prior Extraction for Generalization
by: Rezaeimanesh, Sara, et al.
Published: (2026)
by: Rezaeimanesh, Sara, et al.
Published: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
After Retrieval, Before Generation: Enhancing the Trustworthiness of Large Language Models in Retrieval-Augmented Generation
by: Dai, Xinbang, et al.
Published: (2025)
by: Dai, Xinbang, et al.
Published: (2025)
KinyaColBERT: A Lexically Grounded Retrieval Model for Low-Resource Retrieval-Augmented Generation
by: Nzeyimana, Antoine, et al.
Published: (2025)
by: Nzeyimana, Antoine, et al.
Published: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
by: Collado-Montañez, Jaime, et al.
Published: (2025)
by: Collado-Montañez, Jaime, et al.
Published: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
by: Smădu, Răzvan-Alexandru, et al.
Published: (2025)
Atomic Inference for NLI with Generated Facts as Atoms
by: Stacey, Joe, et al.
Published: (2023)
by: Stacey, Joe, et al.
Published: (2023)
Learning to Generate Structured Output with Schema Reinforcement Learning
by: Lu, Yaxi, et al.
Published: (2025)
by: Lu, Yaxi, et al.
Published: (2025)
Similar Items
-
Spotlights and Blindspots: Evaluating Machine-Generated Text Detection
by: Stowe, Kevin, et al.
Published: (2026) -
Identifying Bias in Machine-generated Text Detection
by: Stowe, Kevin, et al.
Published: (2025) -
Comprehensiveness Metrics for Automatic Evaluation of Factual Recall in Text Generation
by: Dejl, Adam, et al.
Published: (2025) -
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
by: Gupta, Abhay, et al.
Published: (2025) -
SEED: Enhancing Text-to-SQL Performance and Practical Usability Through Automatic Evidence Generation
by: Yun, Janghyeon, et al.
Published: (2025)