Reasoning Language Models for complex assessments tasks: Evaluating parental cooperation from child protection case reports
Fuente:
arXiv
Saved in:
| Main Authors: | Stoll, Dragan, Perron, Brian E., Qi, Zia, Steinmann, Selina, Eicher, Nicole F., Jud, Andreas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Validation of a Small Language Model for DSM-5 Substance Category Classification in Child Welfare Records
by: Perron, Brian E., et al.
Published: (2026)
by: Perron, Brian E., et al.
Published: (2026)
Small Models Achieve Large Language Model Performance: Evaluating Reasoning-Enabled AI for Secure Child Welfare Research
by: Qi, Zia, et al.
Published: (2025)
by: Qi, Zia, et al.
Published: (2025)
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
by: Perron, Brian E., et al.
Published: (2025)
by: Perron, Brian E., et al.
Published: (2025)
AI and Cultural Context: An Empirical Investigation of Large Language Models' Performance on Chinese Social Work Professional Standards
by: Qi, Zia, et al.
Published: (2024)
by: Qi, Zia, et al.
Published: (2024)
A Primer on Word Embeddings: AI Techniques for Text Analysis in Social Work
by: Perron, Brian E., et al.
Published: (2024)
by: Perron, Brian E., et al.
Published: (2024)
AI-Assisted Curation of Conference Scholarship: Compiling, Structuring, and Analyzing Two Decades of Presentations at the Society for Social Work and Research
by: Perron, Brian, et al.
Published: (2026)
by: Perron, Brian, et al.
Published: (2026)
Reducing Selection Bias in Large Language Models
by: Eicher, J. E., et al.
Published: (2024)
by: Eicher, J. E., et al.
Published: (2024)
The influence of persona and conversational task on social interactions with a LLM-controlled embodied conversational agent
by: Kroczek, Leon O. H., et al.
Published: (2024)
by: Kroczek, Leon O. H., et al.
Published: (2024)
MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
by: Wang, Weiqi, et al.
Published: (2024)
by: Wang, Weiqi, et al.
Published: (2024)
Facilitating learners' self‐assessment during formative writing tasks using writing analytics toolkit
by: Luzhen Tang, et al.
Published: (2024)
by: Luzhen Tang, et al.
Published: (2024)
Silent Hill: The Terror Engine
by: Perron, Bernard
Published: (2019)
by: Perron, Bernard
Published: (2019)
MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
by: Wu, Kevin, et al.
Published: (2025)
by: Wu, Kevin, et al.
Published: (2025)
Temporal Causal Reasoning with (Non-Recursive) Structural Equation Models
by: Gladyshev, Maksim, et al.
Published: (2025)
by: Gladyshev, Maksim, et al.
Published: (2025)
"You tell me": A Dataset of GPT-4-Based Behaviour Change Support Conversations
by: Meyer, Selina, et al.
Published: (2024)
by: Meyer, Selina, et al.
Published: (2024)
Evaluating the Performance of LLMs on Technical Language Processing tasks
by: Kernycky, Andrew, et al.
Published: (2024)
by: Kernycky, Andrew, et al.
Published: (2024)
POPCat: Propagation of particles for complex annotation tasks
by: Yang, Adam Srebrnjak, et al.
Published: (2024)
by: Yang, Adam Srebrnjak, et al.
Published: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Large Language Models Meet Symbolic Provers for Logical Reasoning Evaluation
by: Qi, Chengwen, et al.
Published: (2025)
by: Qi, Chengwen, et al.
Published: (2025)
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
by: Lan, Tian, et al.
Published: (2025)
by: Lan, Tian, et al.
Published: (2025)
Demystifying Application Programming Interfaces (APIs): Unlocking the Power of Large Language Models and Other Web-based AI Services in Social Work Research
by: Perron, Brian E., et al.
Published: (2024)
by: Perron, Brian E., et al.
Published: (2024)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
by: Szymanski, Annalisa, et al.
Published: (2024)
by: Szymanski, Annalisa, et al.
Published: (2024)
News Classification in Low‐Resource Languages: Insights From Transformer and Baseline Models
by: Wubetu Barud Demilie, et al.
Published: (2026)
by: Wubetu Barud Demilie, et al.
Published: (2026)
Graph Elicitation for Guiding Multi-Step Reasoning in Large Language Models
by: Park, Jinyoung, et al.
Published: (2023)
by: Park, Jinyoung, et al.
Published: (2023)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
by: Ioannou, Antreas, et al.
Published: (2025)
by: Ioannou, Antreas, et al.
Published: (2025)
SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science
by: Ying, Jie, et al.
Published: (2025)
by: Ying, Jie, et al.
Published: (2025)
How Reliable are LLMs for Reasoning on the Re-ranking task?
by: Islam, Nafis Tanveer, et al.
Published: (2025)
by: Islam, Nafis Tanveer, et al.
Published: (2025)
Towards Robust and Generalizable Lensless Imaging with Modular Learned Reconstruction
by: Bezzam, Eric, et al.
Published: (2025)
by: Bezzam, Eric, et al.
Published: (2025)
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing
by: Anderson, Madeline, et al.
Published: (2025)
by: Anderson, Madeline, et al.
Published: (2025)
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
by: Lau, Gregory Kang Ruey, et al.
Published: (2024)
Configurational-force-driven adaptive refinement and coarsening in topology optimization
by: Stankiewicz, Gabriel, et al.
Published: (2025)
by: Stankiewicz, Gabriel, et al.
Published: (2025)
A novel multi-thickness topology optimization method for balancing structural performance and manufacturability
by: Stankiewicz, Gabriel, et al.
Published: (2025)
by: Stankiewicz, Gabriel, et al.
Published: (2025)
Deblurring structural edges in variable thickness topology optimization via density-gradient-informed projection
by: Stankiewicz, Gabriel, et al.
Published: (2026)
by: Stankiewicz, Gabriel, et al.
Published: (2026)
Differentiable, Bit-shifting, and Scalable Quantization without training neural network from scratch
by: Badar, Zia
Published: (2025)
by: Badar, Zia
Published: (2025)
Adapting Vision Transformers to Ultra-High Resolution Semantic Segmentation with Relay Tokens
by: Perron, Yohann, et al.
Published: (2026)
by: Perron, Yohann, et al.
Published: (2026)
"Jutters"
by: Driessen, Meike, et al.
Published: (2025)
by: Driessen, Meike, et al.
Published: (2025)
Multi-task convolutional neural network for image aesthetic assessment
by: Soydaner, Derya, et al.
Published: (2023)
by: Soydaner, Derya, et al.
Published: (2023)
Fueling Volunteer Growth: the case of Wikipedia Administrators
by: Asikin-Garmager, Eli, et al.
Published: (2026)
by: Asikin-Garmager, Eli, et al.
Published: (2026)
The cost of coordination can exceed the benefit of collaboration in performing complex tasks
by: Straub, Vince J., et al.
Published: (2020)
by: Straub, Vince J., et al.
Published: (2020)
Headlines You Won't Forget: Can Pronoun Insertion Increase Memorability?
by: Meyer, Selina, et al.
Published: (2026)
by: Meyer, Selina, et al.
Published: (2026)
Similar Items
-
Validation of a Small Language Model for DSM-5 Substance Category Classification in Child Welfare Records
by: Perron, Brian E., et al.
Published: (2026) -
Small Models Achieve Large Language Model Performance: Evaluating Reasoning-Enabled AI for Secure Child Welfare Research
by: Qi, Zia, et al.
Published: (2025) -
Human services organizations and the responsible integration of AI: Considering ethics and contextualizing risk(s)
by: Perron, Brian E., et al.
Published: (2025) -
AI and Cultural Context: An Empirical Investigation of Large Language Models' Performance on Chinese Social Work Professional Standards
by: Qi, Zia, et al.
Published: (2024) -
A Primer on Word Embeddings: AI Techniques for Text Analysis in Social Work
by: Perron, Brian E., et al.
Published: (2024)