Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Bedemariam, Rewina, Perez, Natalie, Bhaduri, Sreyoshi, Kapoor, Satya, Gil, Alex, Conjar, Elizabeth, Itoku, Ikkei, Theil, David, Chadha, Aman, Nayyar, Naumaan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transforming Expert Knowledge into Scalable Ontology via Large Language Models
di: Itoku, Ikkei, et al.
Pubblicazione: (2025)
di: Itoku, Ikkei, et al.
Pubblicazione: (2025)
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
di: Perez, Natalie, et al.
Pubblicazione: (2026)
di: Perez, Natalie, et al.
Pubblicazione: (2026)
Reconciling Methodological Paradigms: Employing Large Language Models as Novice Qualitative Research Assistants in Talent Management Research
di: Bhaduri, Sreyoshi, et al.
Pubblicazione: (2024)
di: Bhaduri, Sreyoshi, et al.
Pubblicazione: (2024)
Qualitative Insights Tool (QualIT): LLM Enhanced Topic Modeling
di: Kapoor, Satya, et al.
Pubblicazione: (2024)
di: Kapoor, Satya, et al.
Pubblicazione: (2024)
Decoding the Diversity: A Review of the Indic AI Research Landscape
di: KJ, Sankalp, et al.
Pubblicazione: (2024)
di: KJ, Sankalp, et al.
Pubblicazione: (2024)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
di: KJ, Sankalp, et al.
Pubblicazione: (2025)
di: KJ, Sankalp, et al.
Pubblicazione: (2025)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
di: Balne, Charith Chandra Sai, et al.
Pubblicazione: (2024)
di: Balne, Charith Chandra Sai, et al.
Pubblicazione: (2024)
Levers of Power in the Field of AI
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
What We Do Not Know: GPT Use in Business and Management
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
THELMA: Task Based Holistic Evaluation of Large Language Model Applications-RAG Question Answering
di: Patel, Udita, et al.
Pubblicazione: (2025)
di: Patel, Udita, et al.
Pubblicazione: (2025)
Canary in the Mine: An LLM Augmented Survey of Disciplinary Complaints to the Ordre des ingénieurs du Québec (OIQ)
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
di: Mackenzie, Tammy, et al.
Pubblicazione: (2025)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
di: Singh, Smriti, et al.
Pubblicazione: (2024)
di: Singh, Smriti, et al.
Pubblicazione: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
From Fog to Failure: The Unintended Consequences of Dehazing on Object Detection in Clear Images
di: Kumar, Ashutosh, et al.
Pubblicazione: (2025)
di: Kumar, Ashutosh, et al.
Pubblicazione: (2025)
The Perils & Promises of Fact-checking with Large Language Models
di: Quelle, Dorian, et al.
Pubblicazione: (2023)
di: Quelle, Dorian, et al.
Pubblicazione: (2023)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
di: Chowdhury, Arijit Ghosh, et al.
Pubblicazione: (2023)
di: Chowdhury, Arijit Ghosh, et al.
Pubblicazione: (2023)
Breaking Language Barriers: A Question Answering Dataset for Hindi and Marathi
di: Sabane, Maithili, et al.
Pubblicazione: (2023)
di: Sabane, Maithili, et al.
Pubblicazione: (2023)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
di: Gorti, Atmika, et al.
Pubblicazione: (2024)
di: Gorti, Atmika, et al.
Pubblicazione: (2024)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
di: Sinha, Neelabh, et al.
Pubblicazione: (2024)
LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data
di: Zhang, Hanyu, et al.
Pubblicazione: (2024)
di: Zhang, Hanyu, et al.
Pubblicazione: (2024)
The Complexity of Pure Strategy Relevant Equilibria in Concurrent Games
di: Bhaduri, Purandar
Pubblicazione: (2025)
di: Bhaduri, Purandar
Pubblicazione: (2025)
Can Large Language Models Infer Causal Relationships from Real-World Text?
di: Saklad, Ryan, et al.
Pubblicazione: (2025)
di: Saklad, Ryan, et al.
Pubblicazione: (2025)
MedVisionLlama: Leveraging Pre-Trained Large Language Model Layers to Enhance Medical Image Segmentation
di: Kumar, Gurucharan Marthi Krishna, et al.
Pubblicazione: (2024)
di: Kumar, Gurucharan Marthi Krishna, et al.
Pubblicazione: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
di: Haider, Batool, et al.
Pubblicazione: (2025)
di: Haider, Batool, et al.
Pubblicazione: (2025)
From Prejudice to Parity: A New Approach to Debiasing Large Language Model Word Embeddings
di: Rakshit, Aishik, et al.
Pubblicazione: (2024)
di: Rakshit, Aishik, et al.
Pubblicazione: (2024)
HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent
di: Xu, Weijie, et al.
Pubblicazione: (2024)
di: Xu, Weijie, et al.
Pubblicazione: (2024)
Density Adaptive Attention is All You Need: Robust Parameter-Efficient Fine-Tuning Across Multiple Modalities
di: Ioannides, Georgios, et al.
Pubblicazione: (2024)
di: Ioannides, Georgios, et al.
Pubblicazione: (2024)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
di: Khoshnoodi, Mahsa, et al.
Pubblicazione: (2024)
di: Khoshnoodi, Mahsa, et al.
Pubblicazione: (2024)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
di: Sahoo, Pranab, et al.
Pubblicazione: (2024)
di: Sahoo, Pranab, et al.
Pubblicazione: (2024)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
di: Thakur, Aman Singh, et al.
Pubblicazione: (2024)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
di: Kharchenko, Julia, et al.
Pubblicazione: (2024)
di: Kharchenko, Julia, et al.
Pubblicazione: (2024)
I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
di: Kharchenko, Julia, et al.
Pubblicazione: (2025)
di: Kharchenko, Julia, et al.
Pubblicazione: (2025)
The Potential and Perils of Generative Artificial Intelligence for Quality Improvement and Patient Safety
di: Jalilian, Laleh, et al.
Pubblicazione: (2024)
di: Jalilian, Laleh, et al.
Pubblicazione: (2024)
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model Security
di: Fan, Yihe, et al.
Pubblicazione: (2024)
di: Fan, Yihe, et al.
Pubblicazione: (2024)
Perils of current DAO governance
di: Kharman, Aida Manzano, et al.
Pubblicazione: (2024)
di: Kharman, Aida Manzano, et al.
Pubblicazione: (2024)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
How Culturally Aware are Vision-Language Models?
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
di: Burda-Lassen, Olena, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Transforming Expert Knowledge into Scalable Ontology via Large Language Models
di: Itoku, Ikkei, et al.
Pubblicazione: (2025) -
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
di: Perez, Natalie, et al.
Pubblicazione: (2026) -
Reconciling Methodological Paradigms: Employing Large Language Models as Novice Qualitative Research Assistants in Talent Management Research
di: Bhaduri, Sreyoshi, et al.
Pubblicazione: (2024) -
Qualitative Insights Tool (QualIT): LLM Enhanced Topic Modeling
di: Kapoor, Satya, et al.
Pubblicazione: (2024) -
Decoding the Diversity: A Review of the Indic AI Research Landscape
di: KJ, Sankalp, et al.
Pubblicazione: (2024)