Depth and Autonomy: A Framework for Evaluating LLM Applications in Social Science Research
Fuente:
arXiv
Saved in:
| Main Authors: | Sanaei, Ali, Rajabzadeh, Ali |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications
by: Baes, Naomi, et al.
Published: (2024)
by: Baes, Naomi, et al.
Published: (2024)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
by: Ye, Xiao, et al.
Published: (2025)
by: Ye, Xiao, et al.
Published: (2025)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
by: Veuthey, Jaime Raldua, et al.
Published: (2025)
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026)
by: Rajabzadeh, Hossein, et al.
Published: (2026)
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
by: Yang, Yi, et al.
Published: (2024)
by: Yang, Yi, et al.
Published: (2024)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
by: Kennedy, Molly, et al.
Published: (2026)
by: Kennedy, Molly, et al.
Published: (2026)
Supporting the Digital Autonomy of Elders Through LLM Assistance
by: Roberts, Jesse, et al.
Published: (2024)
by: Roberts, Jesse, et al.
Published: (2024)
LATA: A Tool for LLM-Assisted Translation Annotation
by: Huang, Baorong, et al.
Published: (2026)
by: Huang, Baorong, et al.
Published: (2026)
Evaluating the Retrieval Component in LLM-Based Question Answering Systems
by: Alinejad, Ashkan, et al.
Published: (2024)
by: Alinejad, Ashkan, et al.
Published: (2024)
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026)
by: Chatelain, Arnault, et al.
Published: (2026)
Memory Dial: A Training Framework for Controllable Memorization in Language Models
by: Zhang, Xiangbo, et al.
Published: (2026)
by: Zhang, Xiangbo, et al.
Published: (2026)
KGHaluBench: A Knowledge Graph-Based Hallucination Benchmark for Evaluating the Breadth and Depth of LLM Knowledge
by: Robertson, Alex, et al.
Published: (2026)
by: Robertson, Alex, et al.
Published: (2026)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
by: Mekky, Ali, et al.
Published: (2025)
by: Mekky, Ali, et al.
Published: (2025)
LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
by: Khamis, Ahmed Khaled, et al.
Published: (2026)
by: Khamis, Ahmed Khaled, et al.
Published: (2026)
QDyLoRA: Quantized Dynamic Low-Rank Adaptation for Efficient Large Language Model Tuning
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
EchoAtt: Attend, Copy, then Adjust for More Efficient Large Language Models
by: Rajabzadeh, Hossein, et al.
Published: (2024)
by: Rajabzadeh, Hossein, et al.
Published: (2024)
LLM Sensitivity Evaluation Framework for Clinical Diagnosis
by: Yan, Chenwei, et al.
Published: (2025)
by: Yan, Chenwei, et al.
Published: (2025)
DKG-LLM : A Framework for Medical Diagnosis and Personalized Treatment Recommendations via Dynamic Knowledge Graph and Large Language Model Integration
by: Sarabadani, Ali, et al.
Published: (2025)
by: Sarabadani, Ali, et al.
Published: (2025)
Application of LLM Agents in Recruitment: A Novel Framework for Resume Screening
by: Gan, Chengguang, et al.
Published: (2024)
by: Gan, Chengguang, et al.
Published: (2024)
Evaluating Cultural and Social Awareness of LLM Web Agents
by: Qiu, Haoyi, et al.
Published: (2024)
by: Qiu, Haoyi, et al.
Published: (2024)
Systematic Framework of Application Methods for Large Language Models in Language Sciences
by: Sun, Kun, et al.
Published: (2025)
by: Sun, Kun, et al.
Published: (2025)
LLM Prompt Evaluation for Educational Applications
by: Holmes, Langdon, et al.
Published: (2026)
by: Holmes, Langdon, et al.
Published: (2026)
Trustworthy LLM-Mediated Communication: Evaluating Information Fidelity in LLM as a Communicator (LAAC) Framework in Multiple Application Domains
by: Rafi, Mohammed Musthafa, et al.
Published: (2025)
by: Rafi, Mohammed Musthafa, et al.
Published: (2025)
Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
EvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
by: Dejl, Adam, et al.
Published: (2026)
by: Dejl, Adam, et al.
Published: (2026)
TALE: A Tool-Augmented Framework for Reference-Free Evaluation of Large Language Models
by: Badshah, Sher, et al.
Published: (2025)
by: Badshah, Sher, et al.
Published: (2025)
PeeriScope: A Multi-Faceted Framework for Evaluating Peer Review Quality
by: Ebrahimi, Sajad, et al.
Published: (2026)
by: Ebrahimi, Sajad, et al.
Published: (2026)
MedRCube: A Multidimensional Framework for Fine-Grained and In-Depth Evaluation of MLLMs in Medical Imaging
by: Bao, Zhijie, et al.
Published: (2026)
by: Bao, Zhijie, et al.
Published: (2026)
Rethinking LLM Bias Probing Using Lessons from the Social Sciences
by: Morehouse, Kirsten N., et al.
Published: (2025)
by: Morehouse, Kirsten N., et al.
Published: (2025)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
by: Mohammadkhani, Ali Ghiasvand
Published: (2024)
FedMental: Evaluating Federated Learning for Mental Health Detection from Social Media Data
by: Abdelkadir, Nuredin Ali, et al.
Published: (2026)
by: Abdelkadir, Nuredin Ali, et al.
Published: (2026)
Mind the Blind Spots: A Focus-Level Evaluation Framework for LLM Reviews
by: Shin, Hyungyu, et al.
Published: (2025)
by: Shin, Hyungyu, et al.
Published: (2025)
Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation
by: Wang, Siyuan, et al.
Published: (2024)
by: Wang, Siyuan, et al.
Published: (2024)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
by: Nguyen, Bang, et al.
Published: (2026)
by: Nguyen, Bang, et al.
Published: (2026)
Navigating the Prompt Space: Improving LLM Classification of Social Science Texts Through Prompt Engineering
by: Gunes, Erkan, et al.
Published: (2026)
by: Gunes, Erkan, et al.
Published: (2026)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
by: Jain, Shomik, et al.
Published: (2025)
by: Jain, Shomik, et al.
Published: (2025)
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows
by: Borovkov, Nikita, et al.
Published: (2026)
by: Borovkov, Nikita, et al.
Published: (2026)
AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents
by: Liang, Yuanzhi, et al.
Published: (2024)
by: Liang, Yuanzhi, et al.
Published: (2024)
Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
by: Al-Athba, Aisha Ali, et al.
Published: (2026)
Exploring Boundaries and Intensities in Offensive and Hate Speech: Unveiling the Complex Spectrum of Social Media Discourse
by: Ayele, Abinew Ali, et al.
Published: (2024)
by: Ayele, Abinew Ali, et al.
Published: (2024)
Similar Items
-
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications
by: Baes, Naomi, et al.
Published: (2024) -
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
by: Ye, Xiao, et al.
Published: (2025) -
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
by: Veuthey, Jaime Raldua, et al.
Published: (2025) -
LoRA-Drop: Temporal LoRA Decoding for Efficient LLM Inference
by: Rajabzadeh, Hossein, et al.
Published: (2026) -
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
by: Yang, Yi, et al.
Published: (2024)