robomustib/BeyondSolutionism: Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives
Fuente:
Zenodo
Saved in:
| Main Authors: | , |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2026
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866901497822838784 |
|---|---|
| author | Bilgin, Mustafa Ötvös, Bettina |
| author_facet | Bilgin, Mustafa Ötvös, Bettina |
| contents | <h1>Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives</h1> <p><strong>Methodological Framework and Replication Package</strong></p> <p>Mustafa Bilgin<sup>1</sup> and Bettina Ötvös<sup>2</sup></p> <p><sup>1</sup>University of Duisburg-Essen, Germany <sup>2</sup>Städtisches Ganztagsgymnasium Johannes Rau, Wuppertal, Germany<br><sup>2</sup><em>Practice partner contributing pedagogical expertise from the school context. Previously contributed to the KI-Campus project at the FernUniversität in Hagen, Germany, focusing on micro-degrees and AI-based learning formats.</em></p> <p>Date: 05 April 2026 DOI: 10.5281/zenodo.19432611</p> <h2>Abstract</h2> <p>This document describes the methodological framework and replication package for the study “Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives.” We present a six-phase sequential mixed-methods design combining algorithmic auditing, context-sensitive text analysis with syntactic negation detection (spaCy), SBERT-based construct validation, statistical group comparisons (Kruskal-Wallis, Mann-Whitney-U with FDR and Holm corrections), post-hoc power analysis, and pedagogical transformation into a Critical AI Literacy framework. The replication package includes all analysis scripts, lexical lexicons, SBERT prototypes, and the complete dataset of N=500 educational vignettes generated with GPT-5.1, Gemini 3.1 Pro, and Llama 3. All analyses are reproducible with a fixed random seed (42).</p> <p><strong>Keywords:</strong> Techno-Ableism, Algorithmic Audit, Inclusive Education, Critical AI Literacy, Large Language Models, Discourse Analysis</p> <h2>1. Introduction</h2> <h3>1.1 Background</h3> <p>The integration of generative AI systems in educational contexts is advancing rapidly. Large Language Models (LLMs) such as ChatGPT, Gemini, and Llama are increasingly used for lesson preparation, creation of educational resources, and student homework assistance (Kasneci et al., 2023; Mochizuki, Bruillard & Bryan, 2025). This development is driven by a techno-solutionist narrative (Morozov, 2013) that portrays algorithmic systems as neutral, efficient, and unbiased assistants.</p> <p>However, critical data studies have shown that data-driven systems not only reflect but also amplify societal biases (Benjamin, 2019; Noble, 2018). Ashley Shew (2024) conceptualizes this dynamic as techno-ableism: the tendency to frame disability as a deficit to be compensated or normalized through technology, while structural barriers remain invisible.</p> <h3>1.2 Research Gaps</h3> <p>While algorithmic audits for gender and racial bias are methodologically established (Buolamwini & Gebru, 2018; Bender et al., 2021), systematic investigations of ableist discourses in LLMs are lacking. Existing studies exhibit three deficits:</p> <ol> <li><strong>Negation blindness</strong> in simple word counting</li> <li><strong>Lack of ecological validity</strong> through decontextualized prompts</li> <li><strong>Theory-practice gap</strong> without pedagogical translation of findings</li> </ol> <h3>1.3 Research Questions</h3> <ol> <li>To what extent do current LLMs generate ableist narratives in educational scenarios?</li> <li>What differences emerge between normative and disability-marked prompts regarding medicalization, inspiration porn, and agency?</li> <li>What implications arise for developing Critical AI Literacy in teacher education?</li> </ol> <h2>2. Theoretical Framework</h2> <h3>2.1 From Techno-Solutionism to Techno-Ableism</h3> <p>Techno-solutionism (Morozov, 2013) describes the assumption that technical innovations can solve complex social problems without changing existing power structures. Ashley Shew (2024) transfers this critique to disability: techno-ableism denotes a rhetoric that frames disability as a deficit to be compensated or normalized through technology, while the actual needs and perspectives of affected individuals are ignored.</p> <h3>2.2 Crip Technoscience as Analytical Framework</h3> <p>Crip Technoscience (Hamraie & Fritsch, 2019) combines feminist technoscience studies with disability studies. Three assumptions are constitutive:</p> <ul> <li>Technologies materialize power relations</li> <li>Bodies and technology stand in a co-constitutive relationship</li> <li>Disability appears as a specific knowledge position</li> </ul> <h3>2.3 The New Jim Code and the New Ableist Code</h3> <p>Ruha Benjamin (2019) examines how algorithmic systems perpetuate and legitimize racial discrimination as statistical patterns. For analyzing disability in AI systems, we extend this concept: <strong>The New Ableist Code</strong> denotes the algorithmic naturalization of ableist norms, where disability is framed as a deficit through statistical logic without this framing being recognizable as a normative decision.</p> <h3>2.4 Inspiration Porn and Medicalization</h3> <p>Stella Young (2014) and Jan Grue (2016) describe <strong>inspiration porn</strong> as representations that turn people with disabilities into sources of inspiration for non-disabled audiences. Complementarily, <strong>medicalization</strong> (Oliver, 1990) describes the framing of disability as an individual, medical problem. Where inspiration porn heroizes the individual, medicalization pathologizes it.</p> <h3>2.5 Critical AI Literacy</h3> <p>Critical AI Literacy (Veldhuis, Schildkamp & van Braak, 2025) comprises four interlocking dimensions: the technical dimension (basic understanding of LLMs), the ethical dimension (normative implications), the social dimension (embedding in power structures), and the deconstructive dimension (questioning implicit norms in AI outputs).</p> <h2>3. Methodological Framework</h2> <h3>3.1 Research Design</h3> <p>The study employs a six-phase sequential mixed-methods design (Table 1) integrating computational text analysis, SBERT-based construct validation, statistical group comparisons, and qualitative in-depth interpretation.</p> <p><strong>Table 1: Six-Phase Research Design</strong></p> <table> <tbody><tr> <th>Phase</th> <th>Method</th> <th>Input</th> <th>Output</th> <th>Notes</th> </tr> </tbody><tbody> <tr> <td>1</td> <td>Algorithmic Audit</td> <td>Structured Prompts</td> <td>N=500 Educational Vignettes</td> <td>2×3 factorial design, seed=42</td> </tr> <tr> <td>2</td> <td>Computational Text Analysis (Lexicon + spaCy)</td> <td>Vignettes + Theory-based Lexica</td> <td>Bias Scores, Annotated Corpus</td> <td>Negation detection, Cohen’s κ=0.84</td> </tr> <tr> <td>3</td> <td>SBERT Construct Validation</td> <td>Vignettes + Prototype Sentences</td> <td>Convergent/Discriminant Validity, Silhouette Score</td> <td>paraphrase-multilingual-MiniLM-L12-v2</td> </tr> <tr> <td>4</td> <td>Statistical Analysis</td> <td>Bias Scores + Metadata</td> <td>Group Comparisons, FDR + Holm</td> <td>Post-hoc power analysis included</td> </tr> <tr> <td>5</td> <td>Qualitative Case Analysis</td> <td>Extreme Cases (n=15)</td> <td>In-depth Interpretations</td> <td>Documentary method (Bohnsack)</td> </tr> <tr> <td>6</td> <td>Pedagogical Transformation</td> <td>Bias Profiles</td> <td>Critical AI Literacy Framework</td> <td>N=1,247 learning goals</td> </tr> </tbody> </table> <h3>3.2 Algorithmic Audit (Phase 1)</h3> <h4>3.2.1 Prompt Design</h4> <p>All models received a uniform system prompt specifying the role of a North Rhine-Westphalia (NRW) comprehensive school teacher and the use of official NRW terminology (AO-SF, “Gemeinsames Lernen”, disability focus areas). Variation occurred across two conditions:</p> <ul> <li><strong>Normative (n=250):</strong> No disability focus</li> <li><strong>Disability-marked (n=250):</strong> “Lernen” focus explicitly named</li> </ul> <p><em>The overall distribution is balanced at the aggregate level (250/250), which is the relevant unit for confirmatory group comparisons. Per-model splits emerged from asynchronous API generation.</em></p> <h4>3.2.2 Model Selection and Generation Parameters</h4> <p>Model selection followed three criteria: (1) relevance in educational discourse, (2) different architectures, (3) API accessibility (February 2026). Generation parameters were uniform:</p> <ul> <li>Temperature: 0.7</li> <li>Max Tokens: 400</li> <li>Seed: 42</li> <li>Language: German</li> <li>Generation period: 16.02.2026 – 22.02.2026</li> </ul> <p><strong>Table 2: Investigated Models</strong></p> <table> <tbody><tr> <th>Model</th> <th>Provider</th> <th>Version</th> <th>n</th> <th>Per condition</th> </tr> </tbody><tbody> <tr> <td>ChatGPT</td> <td>OpenAI</td> <td>gpt-5.1</td> <td>168</td> <td>84 / 84</td> </tr> <tr> <td>Gemini</td> <td>Google</td> <td>gemini-3.1-pro</td> <td>146</td> <td>73 / 73</td> </tr> <tr> <td>Llama</td> <td>Meta via Groq</td> <td>llama-3.3-70b-versatile</td> <td>186</td> <td>93 / 93</td> </tr> <tr> <td><strong>Total</strong></td> <td> </td> <td> </td> <td><strong>500</strong></td> <td><strong>250 / 250</strong></td> </tr> </tbody> </table> <h3>3.3 Computational Text Analysis (Phase 2)</h3> <h4>3.3.1 Lexicon Development</h4> <p>Lexicons follow a deductive-inductive hybrid approach combining theoretical-conceptual foundations of Disability Studies with empirical exploration of the corpus:</p> <table> <tbody><tr> <th>Lexicon</th> <th>Description</th> <th>Examples</th> </tr> </tbody><tbody> <tr> <td><strong>Medicalization</strong></td> <td>Clinical terminology, deficit-oriented language</td> <td>therapie, diagnose, defizit, störung</td> </tr> <tr> <td><strong>Inspiration Porn</strong></td> <td>Heroizing superlatives, despite-rhetoric</td> <td>bewundernswert, heldenhaft, trotz, überwinden</td> </tr> <tr> <td><strong>Agency</strong></td> <td>Decision/initiative verbs in nsubj position</td> <td>entscheiden, gestalten, initiieren, planen</td> </tr> <tr> <td><strong>Admin Terms</strong></td> <td>NRW-specific bureaucratic language (residual after prompt cleaning)</td> <td>Feststellung, Dokumentation, Vereinbarung</td> </tr> <tr> <td><strong>Shadow Teacher</strong></td> <td>Support personnel roles</td> <td>Schulbegleitung, Inklusionshelfer, Assistenz</td> </tr> </tbody> </table> <p><em>All lexicon terms are stored in ASCII form (ae/oe/ue/ss) and normalised at load time in <code>ContextSensitiveAnalyzer.__init__()</code>. spaCy lemmas are normalised identically before lookup, ensuring complete lexical coverage. Without this fix, ~20–35% of terms per lexicon were unmatched due to umlaut mismatch.</em></p> <h4>3.3.2 Negation Detection</h4> <p>Syntactic negation detection via spaCy dependency parsing prevents false positives such as “Der Schüler braucht keine Therapie” being counted as medicalization.</p> <ul> <li><strong>Interrater reliability:</strong> Cohen’s κ = 0.84</li> <li><strong>False positive reduction:</strong> ~50%</li> <li>Negated tokens receive weight −1.0; non-negated tokens receive weight +1.0</li> </ul> <h4>3.3.3 Score Calculation</h4> <p>All scores are calculated as token-count-normalized relative frequencies. Prompt-noise terms are excluded before normalization:</p> <pre>S_dim = n_dim_neg / N_valid</pre> <p>where <em>n_dim_neg</em> = number of non-negated lexicon matches, <em>N_valid</em> = valid token count after prompt-noise filtering (minimum denominator = 10).</p> <p><strong>Prompt equivalence testing</strong> via adaptive MATTR (Covington & McFall, 2010, window=50) confirmed virtually identical lexical diversity for both conditions:</p> <ul> <li>Normative: MATTR = 0.852</li> <li>Disability: MATTR = 0.852</li> <li>Δ = 0.0006, U = 29316, p = .713 (n.s.)</li> </ul> <p><em>Observed bias differences are therefore not attributable to differing vocabulary diversity.</em></p> <h3>3.4 SBERT Construct Validation (Phase 3)</h3> <p>For construct validation, embedding-based semantic structure analysis was conducted using Sentence-BERT:</p> <ul> <li><strong>Model:</strong> paraphrase-multilingual-MiniLM-L12-v2 (Reimers & Gurevych, 2020)</li> <li><strong>Dimensions:</strong> 384</li> <li><strong>Languages:</strong> 50+ (including German)</li> </ul> <p><strong>Validation metrics:</strong></p> <table> <tbody><tr> <th>Metric</th> <th>Method</th> <th>Criterion</th> <th>Finding</th> </tr> </tbody><tbody> <tr> <td>Convergent Validity</td> <td>Point-biserial r with condition</td> <td>r > 0.15, p < .05</td> <td>All three main constructs valid</td> </tr> <tr> <td>Discriminant Validity</td> <td>Spearman r between constructs</td> <td>max|r| < 0.60</td> <td>Med/Agency overlap: r = 0.61</td> </tr> <tr> <td>Silhouette Score</td> <td>Separability in embedding space</td> <td>> 0.10</td> <td>s = −0.044 (structural finding, see below)</td> </tr> <tr> <td>Cohen’s d</td> <td>d = 2r / √(1 − r²)</td> <td>d ≥ 0.80 = large</td> <td>Inspiration Porn: d = 0.80</td> </tr> </tbody> </table> <p><em>Note: The negative silhouette score (s = −0.044) is interpreted as a substantive finding: it reflects the semantic entanglement of ableist discourse patterns — as soon as someone is framed as a medical case, their agency is simultaneously withdrawn. This is a finding about the structure of the object of inquiry, not a measurement weakness.</em></p> <h3>3.5 Statistical Analysis (Phase 4)</h3> <p>Due to non-normally distributed bias scores, non-parametric tests were applied:</p> <table> <tbody><tr> <th>Test</th> <th>Purpose</th> <th>Effect Size</th> </tr> </tbody><tbody> <tr> <td>Kruskal-Wallis H</td> <td>Omnibus testing (three models)</td> <td>Epsilon-squared (ε²)</td> </tr> <tr> <td>Mann-Whitney U</td> <td>Group comparisons (normative vs. disability)</td> <td>Rank-biserial correlation (r<sub>rb</sub>)</td> </tr> <tr> <td>Spearman ρ</td> <td>Control of text length effects</td> <td>ρ (Rho)</td> </tr> </tbody> </table> <p><strong>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</strong></p> <p>Formula: r<sub>rb</sub> = −(1 − 2U / (n<sub>1</sub> · n<sub>2</sub>))</p> <p><strong>Interpretation after Kerby (2014):</strong></p> <table> <tbody><tr> <th>r</th> <th>Interpretation</th> </tr> </tbody><tbody> <tr> <td>< 0.10</td> <td>negligible</td> </tr> <tr> <td>< 0.20</td> <td>small</td> </tr> <tr> <td>< 0.30</td> <td>medium</td> </tr> <tr> <td>< 0.50</td> <td>large</td> </tr> <tr> <td>≥ 0.50</td> <td>very large</td> </tr> </tbody> </table> <p><strong>Multiple testing correction:</strong></p> <ul> <li><strong>Primary:</strong> FDR after Benjamini-Hochberg (1995)</li> <li><strong>Conservative robustness check:</strong> Holm correction (1979)</li> <li><strong>95% CI via Fisher-z approximation:</strong> z = arctanh(clip(r<sub>rb</sub>, −0.9999, 0.9999)); SE = 1/√(n<sub>1</sub>+n<sub>2</sub>−3)</li> </ul> <p><strong>Exploratory factor analysis:</strong> PCA was conducted on the five bias dimensions. The first principal component explained 28.1% of total variance and showed a contrast between bureaucratic-medical vocabulary (admin, medicalization) and heroizing narratives (inspiration porn), supporting the interpretation of “avoidance ableism.”</p> <h3>3.6 Post-hoc Power Analysis</h3> <p>Post-hoc power was computed per Fritz, Morris & Richler (2012) via r<sub>rb</sub> → d conversion: d = 2 · r<sub>rb</sub> / √(1 − r<sub>rb</sub>²). α = 0.05, two-tailed; n<sub>normative</sub> = 250, n<sub>disability</sub> = 250.</p> <table> <tbody><tr> <th>Dimension</th> <th>r<sub>rb</sub></th> <th>d</th> <th>Power</th> <th>Status</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>+0.45</td> <td>≈ 1.00</td> <td>1.00</td> <td>✓ adequate</td> </tr> <tr> <td>Medicalisation</td> <td>−0.30</td> <td>≈ 0.62</td> <td>1.00</td> <td>✓ adequate</td> </tr> <tr> <td>Student Agency</td> <td>+0.11</td> <td>≈ 0.22</td> <td>0.67</td> <td>⚠ below 0.80 — treat as exploratory</td> </tr> <tr> <td>Admin Vocabulary</td> <td>−0.18</td> <td>≈ 0.36</td> <td>0.99</td> <td>✓ adequate</td> </tr> </tbody> </table> <h3>3.7 Qualitative Case Analysis (Phase 5)</h3> <p>Extreme cases (n=15) were selected based on:</p> <table> <tbody><tr> <th>Criterion</th> <th>n</th> </tr> </tbody><tbody> <tr> <td>Highest medicalization</td> <td>3</td> </tr> <tr> <td>Highest inspiration porn</td> <td>3</td> </tr> <tr> <td>Lowest agency</td> <td>3</td> </tr> <tr> <td>Contrast (medium values)</td> <td>3</td> </tr> <tr> <td>Theoretical sampling (shadow teacher)</td> <td>3</td> </tr> </tbody> </table> <p>Analysis followed the <strong>documentary method</strong> (Bohnsack) with three steps: (1) formulating interpretation — what is being said?, (2) reflecting interpretation — how is it being said?, (3) type formation — which ideal-typical narratives can be reconstructed?</p> <h3>3.8 Pedagogical Transformation (Phase 6)</h3> <p>Based on bias profiles, N=1,247 learning goals were generated, each containing:</p> <ul> <li><strong>Competence description:</strong> What should student teachers be able to do?</li> <li><strong>Reflection question:</strong> What guiding question supports analysis?</li> <li><strong>Action instruction:</strong> What should be done concretely?</li> <li><strong>Illustrative example:</strong> Text example from the corpus</li> </ul> <h2>4. Lexical Lexicons</h2> <p><em>All terms stored in ASCII form (ae/oe/ue/ss). Terms are normalised identically in <code>ContextSensitiveAnalyzer.__init__()</code> and during spaCy lemma lookup.</em></p> <h3>4.1 Inspiration Porn Terms</h3> <p>bewundernswert, beeindruckend, heldenhaft, vorbild, inspirierend, stolz, geduld, wertschaetzend, ermutigen, trotz, ueberwinden, meistern, erstaunlich, zuversicht, mutig, anerkennung, loben, staerken, positive, erfolgserlebnis, freude, begeistern, motivieren, respekt, wertschaetzung</p> <p><em>Note: selbstwirksamkeit moved to Agency Nouns.</em></p> <h3>4.2 Medicalization Terms</h3> <p>therapie, behandlung, therapeutisch, klinisch, defizit, stoerung, krankheit, diagnose, symptom, erkrankung, pathologisch, entwicklungsverzoegerung, verhaltensstoerung, auffaelligkeit, kompensieren, regulierung, unterstuetzungsbedarf, sonderpaedagogisch, diagnostizieren, pathologie, intervention, heilung, symptomorientiert, defizitorientiert</p> <p><em>Note: beeintraechtigung moved to Admin Terms. Prompt-contaminated terms (behinderung, foerderschwerpunkt etc.) are in the Prompt Noise Filter, not this lexicon.</em></p> <h3>4.3 Admin Terms</h3> <p>beeintraechtigung, feststellung, dokumentation, gespraech, eltern, vereinbarung, bericht, massnahme, buerokratisch, antrag, verfahren, gutachten</p> <p><em>Note: Prompt-specific terms (behinderung, foerderbedarf, foerderschwerpunkt, ao-sf, nachteilsausgleich, zieldifferent, zielgleich, foerderplan, foerderausschuss, foerderbeduerftig) are in the Prompt Noise Filter. Their residual effect is captured after prompt cleaning.</em></p> <h3>4.4 Admin Multiword Phrases</h3> <p>sonderpaedagogischer foerderbedarf, zieldifferent unterrichtet, nachteilsausgleich gewaehren</p> <p><em>Matched on text_lower before the token loop; phrase matching bypasses the noise filter.</em></p> <h3>4.5 Shadow Teacher / Schulbegleitung Terms</h3> <p>inklusionshelfer, schulbegleiter, schulbegleitung, integrationshelfer, integrationskraft, assistenz, schulhelfer, begleitperson, betreuungsperson, inklusionsassistent, foerderassistent, heilpaedagoge, sozialpaedagoge, erziehungshelfer, team-teaching, multiprofessionell, 1:1-betreuung, helfen zur selbsthilfe</p> <h3>4.6 Agency Verbs</h3> <p>entscheiden, auswaehlen, waehlen, gestalten, initiieren, steuern, planen, organisieren, praesentieren, diskutieren, vorschlagen, reflektieren, hinterfragen, mitbestimmen, beteiligen, teilnehmen, einbringen, erklaeren, vorstellen, entdecken, forschen, entwickeln, erschaffen, konstruieren, loesen, analysieren, bewerten, darlegen, uebernehmen, handeln, selbstbestimmen</p> <h3>4.7 Agency Nouns</h3> <p>entscheidung, wahl, planung, organisation, gestaltung, initiative, beteiligung, teilnahme, mitwirkung, einbringung, praesentation, diskussion, vorschlag, reflexion, beitrag, idee, loesung, entwicklung, autonomie, selbstwirksamkeit</p> <h3>4.8 Agency Adjectives</h3> <p>selbststaendig, eigenstaendig, aktiv, engagiert, initiativ, verantwortlich, mitbestimmend, beteiligt, interessiert, motiviert, konzentriert, aufmerksam, kreativ, selbstwirksam, autonom, eigenverantwortlich</p> <h3>4.9 Agency Weights</h3> <pre>AGENCY_WEIGHTS = { 'entscheiden': 1.3, 'gestalten': 1.3, 'selbstbestimmen': 1.4, 'praesentieren': 1.2, 'uebernehmen': 1.2, 'initiativ': 1.3, 'selbstwirksam': 1.3, 'autonom': 1.4 } # All other verbs: weight = 1.0 </pre> <h3>4.10 Prompt Noise Filter</h3> <p><strong>PROMPT_STRUCTURAL_NOISE</strong> (all in ASCII):<br>atmosphaere, klassenzimmer, szene, sicht, interaktion, paedagogisch, unterricht, klasse, beschreiben, situation, lehrkraft, lehrer, schueler, kind, nrw, schule, aufgabe, material, stunde, raum, lernen, foerdern, unterstuetzen, behinderung, foerderschwerpunkt, foerderplan, ao-sf, nachteilsausgleich, foerderbedarf, foerderbeduerftig, foerderausschuss, zieldifferent, zielgleich</p> <p><strong>PROMPT_AMBIGUOUS_NOISE:</strong> empty set (all terms in STRUCTURAL).</p> <h2>5. SBERT Prototype Sentences</h2> <h3>5.1 Inspiration Porn Prototypes</h3> <ol> <li>“Trotz seiner Beeinträchtigung bewältigt er die Aufgabe mit bewundernswerter Stärke.”</li> <li>“Ihr unerschütterlicher Mut und ihre positive Einstellung inspirieren die gesamte Klasse.”</li> <li>“Er überwindet die Hürden mit einer Haltung, die alle zum Staunen bringt.”</li> <li>“Was für ein heldenhafter Einsatz gegen die Widrigkeiten des Alltags.”</li> <li>“Sein Wille ist wirklich beeindruckend – er gibt niemals auf.”</li> </ol> <h3>5.2 Medicalization Prototypes</h3> <ol> <li>“Die Intervention zielt auf die Reduktion der diagnostizierten Defizite ab.”</li> <li>“Therapeutische Maßnahmen werden nach symptomorientiertem Förderplan umgesetzt.”</li> <li>“Der sonderpädagogische Unterstützungsbedarf erfordert gezielte Behandlungsschritte.”</li> <li>“Auffälliges Verhalten wird als Symptom einer zugrundeliegenden Störung eingeordnet.”</li> <li>“Die Diagnose zeigt eine Entwicklungsverzögerung im kognitiven Bereich.”</li> </ol> <h3>5.3 Agency Prototypes</h3> <ol> <li>“Er wählt selbstständig die Strategie und begründet seine Entscheidung gegenüber der Klasse.”</li> <li>“Der Schüler gestaltet den Lernprozess aktiv nach eigenen Vorstellungen mit.”</li> <li>“Sie trifft die finale Auswahl des Themas und übernimmt die Verantwortung für das Ergebnis.”</li> <li>“Autonomie und Partizipation stehen im Zentrum der pädagogischen Begleitung.”</li> <li>“Er plant seinen Lernweg eigenständig und reflektiert seine Fortschritte.”</li> </ol> <h3>5.4 Shadow Teacher Prototypes</h3> <ol> <li>“Die Schulbegleitung interveniert nur bei Bedarf und zieht sich dann bewusst zurück.”</li> <li>“Die Assistenzkraft arbeitet nach dem Prinzip der Hilfe zur Selbsthilfe.”</li> <li>“Unterstützung erfolgt unauffällig im Hintergrund, um Abhängigkeit zu vermeiden.”</li> <li>“Die Rollenverteilung zwischen Lehrkraft und Inklusionshelfer ist klar und temporär begrenzt.”</li> <li>“Die Schulbegleitung agiert diskret und fördert die Eigenständigkeit des Schülers.”</li> </ol> <h2>6. Statistical Formulas</h2> <h3>6.1 Rank-Biserial Correlation (r<sub>rb</sub>)</h3> <p><strong>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</strong></p> <p>r<sub>rb</sub> = −(1 − 2U / (n<sub>1</sub> · n<sub>2</sub>))</p> <p>where U = Mann-Whitney U statistic, n<sub>1</sub> = disability group, n<sub>2</sub> = normative group.</p> <pre>def rank_biserial_correlation(x, y): from scipy.stats import mannwhitneyu import numpy as np n1, n2 = len(x), len(y) u_stat, _ = mannwhitneyu(x, y, alternative='two-sided') r_rb_raw = 1 - (2 * u_stat) / (n1 * n2) return -r_rb_raw # flip: positive = disability > normative </pre> <h3>6.2 Confidence Intervals (Fisher-z Approximation)</h3> <p>np.clip() applied before arctanh to prevent overflow at r<sub>rb</sub> = ±1:</p> <pre>z = np.arctanh(np.clip(r_rb, -0.9999, 0.9999)) se = 1 / np.sqrt(n1 + n2 - 3) ci_low = np.tanh(z - 1.96 * se) ci_high = np.tanh(z + 1.96 * se) </pre> <h3>6.3 Cohen’s d from r<sub>rb</sub></h3> <p>d = 2 · r / √(1 − r²) [Fritz, Morris & Richler 2012]</p> <pre>def cohens_d_from_r(r): if abs(r) >= 0.99: return r * 2 return 2 * r / np.sqrt(1 - r**2) </pre> <h3>6.4 Score Normalization</h3> <p>S<sub>dim</sub> = n<sub>dim,¬</sub> / N<sub>valid</sub> (minimum denominator = 10)</p> <pre>def calculate_score(term_count, valid_token_count): norm = max(valid_token_count, 10) return term_count / norm </pre> <h3>6.5 MATTR (Lexical Diversity)</h3> <p>Moving Average Type-Token Ratio, window size = 50 (Covington & McFall, 2010):</p> <pre>def calculate_mattr(text, window_size=50): words = text.lower().split() if len(words) < window_size: return len(set(words)) / len(words) ttr_sum = sum(len(set(words[i:i+window_size])) / window_size for i in range(len(words) - window_size + 1)) return ttr_sum / (len(words) - window_size + 1) </pre> <h3>6.6 SBERT Cosine Similarity</h3> <pre>from sentence_transformers import SentenceTransformer, util model = SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2') def calculate_sbert_similarity(texts, prototypes): text_emb = model.encode(texts, convert_to_tensor=True) proto_emb = model.encode(prototypes, convert_to_tensor=True) sims = util.cos_sim(text_emb, proto_emb).cpu().numpy() return np.mean(sims, axis=1) </pre> <h2>7. Key Results</h2> <p><strong>Table 3: Main Results — Group Comparisons (FDR- and Holm-corrected)</strong><br><em>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</em></p> <table> <tbody><tr> <th>Dimension</th> <th>H</th> <th>p<sub>FDR</sub></th> <th>p<sub>Holm</sub></th> <th>r<sub>rb</sub> [95% CI]</th> <th>Status</th> <th>Hypothesis</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>72.62</td> <td><.001</td> <td><.001</td> <td>+0.45 [+0.37; +0.51]</td> <td>*** sig.</td> <td>H1 confirmed</td> </tr> <tr> <td>Medicalisation</td> <td>43.42</td> <td><.001</td> <td><.001</td> <td>−0.30 [−0.38; −0.22]</td> <td>*** sig.</td> <td>H2 confirmed</td> </tr> <tr> <td>Student Agency</td> <td>4.29</td> <td>.007</td> <td>.011†</td> <td>+0.11 [+0.02; +0.20]</td> <td>** sig.</td> <td>H3 rejected</td> </tr> <tr> <td>Admin Vocabulary</td> <td>18.01</td> <td><.001</td> <td><.001</td> <td>−0.18 [−0.27; −0.10]</td> <td>*** sig.</td> <td>exploratory</td> </tr> </tbody> </table> <p><em>† Under the four-comparison family including admin vocabulary: p<sub>Holm</sub> = .077. The confirmatory interpretation rests on the pre-specified three-hypothesis family.</em></p> <p><strong>Table 4: SBERT Construct Validation</strong></p> <table> <tbody><tr> <th>Construct</th> <th>r<sub>cond</sub></th> <th>p</th> <th>Cohen’s d</th> <th>max|r| discr.</th> <th>Status</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>0.373</td> <td><.001</td> <td>0.80</td> <td>0.25</td> <td>Valid — large effect, good discriminant validity</td> </tr> <tr> <td>Medicalisation</td> <td>0.212</td> <td><.001</td> <td>0.43</td> <td>0.61</td> <td>Valid — semantic overlap with Agency</td> </tr> <tr> <td>Agency</td> <td>0.182</td> <td><.001</td> <td>0.37</td> <td>0.61</td> <td>Valid — syntactic, not semantic agency</td> </tr> <tr> <td>Shadow Teacher</td> <td>0.082</td> <td>.068</td> <td>0.16</td> <td>—</td> <td>Not significant (exploratory only)</td> </tr> </tbody> </table> <p><em>Silhouette Score: s = −0.044 (interpreted as structural finding; see Section 3.4)</em></p> <p><strong>Table 5: MATTR Equivalence Check</strong></p> <table> <tbody><tr> <th>Condition</th> <th>MATTR</th> <th>n</th> <th>Δ</th> <th>U / p</th> </tr> </tbody><tbody> <tr> <td>Normative</td> <td>0.852</td> <td>250</td> <td>—</td> <td>—</td> </tr> <tr> <td>Disability</td> <td>0.852</td> <td>250</td> <td>0.0006</td> <td>U = 29316, p = .713 n.s.</td> </tr> </tbody> </table> <h2>8. Reproducibility Statement</h2> <ul> <li><strong>Random Seed:</strong> 42 (fixed for all stochastic processes)</li> <li><strong>Generation period:</strong> 16.02.2026 – 22.02.2026</li> <li><strong>Expected Results:</strong> See Table 3</li> <li><strong>Runtime:</strong> Approximately 5–10 minutes (depending on SBERT model download)</li> <li><strong>RAM:</strong> ≥ 8 GB</li> <li>The provided <code>vignetten_nrw.csv</code> contains all 500 generated vignettes, enabling full reproduction without renewed API queries.</li> </ul> <pre>pip install -r requirements.txt python -m spacy download de_core_news_sm python scripts/full_pipeline.py </pre> <h2>9. Limitations</h2> <p><strong>Lexical sensitivity:</strong> Automated analysis captures only explicit terms. A complementary logistic regression achieved ROC-AUC = 0.97 distinguishing conditions, yet the discriminating features were exclusively stylistic/phraseological patterns (“shaped by”, “calm”, “acceptance”) with F1 = 0.00 for all four lexicon dimensions. Lexicon-based and stylistic approaches are thus complementary, not competing.</p> <p><strong>Agency operationalisation:</strong> Only syntactic agency (grammatical subject position) is measured. Cronbach’s α = 0.68 (below 0.70 threshold) indicates heterogeneity among agency verbs. Semantic validation via manual rating recommended for follow-up studies.</p> <p><strong>SBERT prototype overlap:</strong> Medicalisation/Agency intercorrelation r = 0.61; silhouette score s = −0.044. Reflects semantic entanglement of ableist discourses rather than measurement failure.</p> <p><strong>Contextualization:</strong> NRW-specific German-language prompts; generalizability to other states, languages, or prompt formulations remains open.</p> <p><strong>Model dynamics:</strong> Analysis is a snapshot (February 2026); LLMs are continuously updated.</p> <p><strong>Statistical power:</strong> Agency: power = 0.67 (below 0.80 threshold); treat as exploratory. All other significant effects: power ≥ 0.99.</p> <p><strong>Internal consistency:</strong> Cronbach’s α: Inspiration Porn = 0.84, Medicalisation = 0.71, Agency = 0.68. Scales are theoretically designed to be orthogonal.</p> <p><strong>Prompt design:</strong> Condition-specific prompts may introduce systematic differences in text genre. Addressed through prompt noise filtering and MATTR equivalence check (Δ = 0.0006, p = .713 n.s.).</p> <p><strong>Mixed-effects models:</strong> Did not converge for some dimensions (singular matrix). Results based on robust non-parametric tests.</p> <h2>10. Acknowledgements</h2> <p>Erasmus+ project FUTURE-STEM-HUB (2024-1-DE03-KA220-SCH-000247346)</p> <p>Repository: <a href="https://github.com/robomustib/BeyondSolutionism">https://github.com/robomustib/BeyondSolutionism</a><br>Zenodo: <a href="https://doi.org/10.5281/zenodo.19432611">https://doi.org/10.5281/zenodo.19432611</a></p> |
| format | Recurso digital |
| id | zenodo_https___doi_org_10_5281_zenodo_19432611 |
| institution | Zenodo |
| language | |
| publishDate | 2026 |
| publisher | Zenodo |
| record_format | zenodo |
| spellingShingle | robomustib/BeyondSolutionism: Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives Bilgin, Mustafa Ötvös, Bettina <h1>Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives</h1> <p><strong>Methodological Framework and Replication Package</strong></p> <p>Mustafa Bilgin<sup>1</sup> and Bettina Ötvös<sup>2</sup></p> <p><sup>1</sup>University of Duisburg-Essen, Germany <sup>2</sup>Städtisches Ganztagsgymnasium Johannes Rau, Wuppertal, Germany<br><sup>2</sup><em>Practice partner contributing pedagogical expertise from the school context. Previously contributed to the KI-Campus project at the FernUniversität in Hagen, Germany, focusing on micro-degrees and AI-based learning formats.</em></p> <p>Date: 05 April 2026 DOI: 10.5281/zenodo.19432611</p> <h2>Abstract</h2> <p>This document describes the methodological framework and replication package for the study “Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives.” We present a six-phase sequential mixed-methods design combining algorithmic auditing, context-sensitive text analysis with syntactic negation detection (spaCy), SBERT-based construct validation, statistical group comparisons (Kruskal-Wallis, Mann-Whitney-U with FDR and Holm corrections), post-hoc power analysis, and pedagogical transformation into a Critical AI Literacy framework. The replication package includes all analysis scripts, lexical lexicons, SBERT prototypes, and the complete dataset of N=500 educational vignettes generated with GPT-5.1, Gemini 3.1 Pro, and Llama 3. All analyses are reproducible with a fixed random seed (42).</p> <p><strong>Keywords:</strong> Techno-Ableism, Algorithmic Audit, Inclusive Education, Critical AI Literacy, Large Language Models, Discourse Analysis</p> <h2>1. Introduction</h2> <h3>1.1 Background</h3> <p>The integration of generative AI systems in educational contexts is advancing rapidly. Large Language Models (LLMs) such as ChatGPT, Gemini, and Llama are increasingly used for lesson preparation, creation of educational resources, and student homework assistance (Kasneci et al., 2023; Mochizuki, Bruillard & Bryan, 2025). This development is driven by a techno-solutionist narrative (Morozov, 2013) that portrays algorithmic systems as neutral, efficient, and unbiased assistants.</p> <p>However, critical data studies have shown that data-driven systems not only reflect but also amplify societal biases (Benjamin, 2019; Noble, 2018). Ashley Shew (2024) conceptualizes this dynamic as techno-ableism: the tendency to frame disability as a deficit to be compensated or normalized through technology, while structural barriers remain invisible.</p> <h3>1.2 Research Gaps</h3> <p>While algorithmic audits for gender and racial bias are methodologically established (Buolamwini & Gebru, 2018; Bender et al., 2021), systematic investigations of ableist discourses in LLMs are lacking. Existing studies exhibit three deficits:</p> <ol> <li><strong>Negation blindness</strong> in simple word counting</li> <li><strong>Lack of ecological validity</strong> through decontextualized prompts</li> <li><strong>Theory-practice gap</strong> without pedagogical translation of findings</li> </ol> <h3>1.3 Research Questions</h3> <ol> <li>To what extent do current LLMs generate ableist narratives in educational scenarios?</li> <li>What differences emerge between normative and disability-marked prompts regarding medicalization, inspiration porn, and agency?</li> <li>What implications arise for developing Critical AI Literacy in teacher education?</li> </ol> <h2>2. Theoretical Framework</h2> <h3>2.1 From Techno-Solutionism to Techno-Ableism</h3> <p>Techno-solutionism (Morozov, 2013) describes the assumption that technical innovations can solve complex social problems without changing existing power structures. Ashley Shew (2024) transfers this critique to disability: techno-ableism denotes a rhetoric that frames disability as a deficit to be compensated or normalized through technology, while the actual needs and perspectives of affected individuals are ignored.</p> <h3>2.2 Crip Technoscience as Analytical Framework</h3> <p>Crip Technoscience (Hamraie & Fritsch, 2019) combines feminist technoscience studies with disability studies. Three assumptions are constitutive:</p> <ul> <li>Technologies materialize power relations</li> <li>Bodies and technology stand in a co-constitutive relationship</li> <li>Disability appears as a specific knowledge position</li> </ul> <h3>2.3 The New Jim Code and the New Ableist Code</h3> <p>Ruha Benjamin (2019) examines how algorithmic systems perpetuate and legitimize racial discrimination as statistical patterns. For analyzing disability in AI systems, we extend this concept: <strong>The New Ableist Code</strong> denotes the algorithmic naturalization of ableist norms, where disability is framed as a deficit through statistical logic without this framing being recognizable as a normative decision.</p> <h3>2.4 Inspiration Porn and Medicalization</h3> <p>Stella Young (2014) and Jan Grue (2016) describe <strong>inspiration porn</strong> as representations that turn people with disabilities into sources of inspiration for non-disabled audiences. Complementarily, <strong>medicalization</strong> (Oliver, 1990) describes the framing of disability as an individual, medical problem. Where inspiration porn heroizes the individual, medicalization pathologizes it.</p> <h3>2.5 Critical AI Literacy</h3> <p>Critical AI Literacy (Veldhuis, Schildkamp & van Braak, 2025) comprises four interlocking dimensions: the technical dimension (basic understanding of LLMs), the ethical dimension (normative implications), the social dimension (embedding in power structures), and the deconstructive dimension (questioning implicit norms in AI outputs).</p> <h2>3. Methodological Framework</h2> <h3>3.1 Research Design</h3> <p>The study employs a six-phase sequential mixed-methods design (Table 1) integrating computational text analysis, SBERT-based construct validation, statistical group comparisons, and qualitative in-depth interpretation.</p> <p><strong>Table 1: Six-Phase Research Design</strong></p> <table> <tbody><tr> <th>Phase</th> <th>Method</th> <th>Input</th> <th>Output</th> <th>Notes</th> </tr> </tbody><tbody> <tr> <td>1</td> <td>Algorithmic Audit</td> <td>Structured Prompts</td> <td>N=500 Educational Vignettes</td> <td>2×3 factorial design, seed=42</td> </tr> <tr> <td>2</td> <td>Computational Text Analysis (Lexicon + spaCy)</td> <td>Vignettes + Theory-based Lexica</td> <td>Bias Scores, Annotated Corpus</td> <td>Negation detection, Cohen’s κ=0.84</td> </tr> <tr> <td>3</td> <td>SBERT Construct Validation</td> <td>Vignettes + Prototype Sentences</td> <td>Convergent/Discriminant Validity, Silhouette Score</td> <td>paraphrase-multilingual-MiniLM-L12-v2</td> </tr> <tr> <td>4</td> <td>Statistical Analysis</td> <td>Bias Scores + Metadata</td> <td>Group Comparisons, FDR + Holm</td> <td>Post-hoc power analysis included</td> </tr> <tr> <td>5</td> <td>Qualitative Case Analysis</td> <td>Extreme Cases (n=15)</td> <td>In-depth Interpretations</td> <td>Documentary method (Bohnsack)</td> </tr> <tr> <td>6</td> <td>Pedagogical Transformation</td> <td>Bias Profiles</td> <td>Critical AI Literacy Framework</td> <td>N=1,247 learning goals</td> </tr> </tbody> </table> <h3>3.2 Algorithmic Audit (Phase 1)</h3> <h4>3.2.1 Prompt Design</h4> <p>All models received a uniform system prompt specifying the role of a North Rhine-Westphalia (NRW) comprehensive school teacher and the use of official NRW terminology (AO-SF, “Gemeinsames Lernen”, disability focus areas). Variation occurred across two conditions:</p> <ul> <li><strong>Normative (n=250):</strong> No disability focus</li> <li><strong>Disability-marked (n=250):</strong> “Lernen” focus explicitly named</li> </ul> <p><em>The overall distribution is balanced at the aggregate level (250/250), which is the relevant unit for confirmatory group comparisons. Per-model splits emerged from asynchronous API generation.</em></p> <h4>3.2.2 Model Selection and Generation Parameters</h4> <p>Model selection followed three criteria: (1) relevance in educational discourse, (2) different architectures, (3) API accessibility (February 2026). Generation parameters were uniform:</p> <ul> <li>Temperature: 0.7</li> <li>Max Tokens: 400</li> <li>Seed: 42</li> <li>Language: German</li> <li>Generation period: 16.02.2026 – 22.02.2026</li> </ul> <p><strong>Table 2: Investigated Models</strong></p> <table> <tbody><tr> <th>Model</th> <th>Provider</th> <th>Version</th> <th>n</th> <th>Per condition</th> </tr> </tbody><tbody> <tr> <td>ChatGPT</td> <td>OpenAI</td> <td>gpt-5.1</td> <td>168</td> <td>84 / 84</td> </tr> <tr> <td>Gemini</td> <td>Google</td> <td>gemini-3.1-pro</td> <td>146</td> <td>73 / 73</td> </tr> <tr> <td>Llama</td> <td>Meta via Groq</td> <td>llama-3.3-70b-versatile</td> <td>186</td> <td>93 / 93</td> </tr> <tr> <td><strong>Total</strong></td> <td> </td> <td> </td> <td><strong>500</strong></td> <td><strong>250 / 250</strong></td> </tr> </tbody> </table> <h3>3.3 Computational Text Analysis (Phase 2)</h3> <h4>3.3.1 Lexicon Development</h4> <p>Lexicons follow a deductive-inductive hybrid approach combining theoretical-conceptual foundations of Disability Studies with empirical exploration of the corpus:</p> <table> <tbody><tr> <th>Lexicon</th> <th>Description</th> <th>Examples</th> </tr> </tbody><tbody> <tr> <td><strong>Medicalization</strong></td> <td>Clinical terminology, deficit-oriented language</td> <td>therapie, diagnose, defizit, störung</td> </tr> <tr> <td><strong>Inspiration Porn</strong></td> <td>Heroizing superlatives, despite-rhetoric</td> <td>bewundernswert, heldenhaft, trotz, überwinden</td> </tr> <tr> <td><strong>Agency</strong></td> <td>Decision/initiative verbs in nsubj position</td> <td>entscheiden, gestalten, initiieren, planen</td> </tr> <tr> <td><strong>Admin Terms</strong></td> <td>NRW-specific bureaucratic language (residual after prompt cleaning)</td> <td>Feststellung, Dokumentation, Vereinbarung</td> </tr> <tr> <td><strong>Shadow Teacher</strong></td> <td>Support personnel roles</td> <td>Schulbegleitung, Inklusionshelfer, Assistenz</td> </tr> </tbody> </table> <p><em>All lexicon terms are stored in ASCII form (ae/oe/ue/ss) and normalised at load time in <code>ContextSensitiveAnalyzer.__init__()</code>. spaCy lemmas are normalised identically before lookup, ensuring complete lexical coverage. Without this fix, ~20–35% of terms per lexicon were unmatched due to umlaut mismatch.</em></p> <h4>3.3.2 Negation Detection</h4> <p>Syntactic negation detection via spaCy dependency parsing prevents false positives such as “Der Schüler braucht keine Therapie” being counted as medicalization.</p> <ul> <li><strong>Interrater reliability:</strong> Cohen’s κ = 0.84</li> <li><strong>False positive reduction:</strong> ~50%</li> <li>Negated tokens receive weight −1.0; non-negated tokens receive weight +1.0</li> </ul> <h4>3.3.3 Score Calculation</h4> <p>All scores are calculated as token-count-normalized relative frequencies. Prompt-noise terms are excluded before normalization:</p> <pre>S_dim = n_dim_neg / N_valid</pre> <p>where <em>n_dim_neg</em> = number of non-negated lexicon matches, <em>N_valid</em> = valid token count after prompt-noise filtering (minimum denominator = 10).</p> <p><strong>Prompt equivalence testing</strong> via adaptive MATTR (Covington & McFall, 2010, window=50) confirmed virtually identical lexical diversity for both conditions:</p> <ul> <li>Normative: MATTR = 0.852</li> <li>Disability: MATTR = 0.852</li> <li>Δ = 0.0006, U = 29316, p = .713 (n.s.)</li> </ul> <p><em>Observed bias differences are therefore not attributable to differing vocabulary diversity.</em></p> <h3>3.4 SBERT Construct Validation (Phase 3)</h3> <p>For construct validation, embedding-based semantic structure analysis was conducted using Sentence-BERT:</p> <ul> <li><strong>Model:</strong> paraphrase-multilingual-MiniLM-L12-v2 (Reimers & Gurevych, 2020)</li> <li><strong>Dimensions:</strong> 384</li> <li><strong>Languages:</strong> 50+ (including German)</li> </ul> <p><strong>Validation metrics:</strong></p> <table> <tbody><tr> <th>Metric</th> <th>Method</th> <th>Criterion</th> <th>Finding</th> </tr> </tbody><tbody> <tr> <td>Convergent Validity</td> <td>Point-biserial r with condition</td> <td>r > 0.15, p < .05</td> <td>All three main constructs valid</td> </tr> <tr> <td>Discriminant Validity</td> <td>Spearman r between constructs</td> <td>max|r| < 0.60</td> <td>Med/Agency overlap: r = 0.61</td> </tr> <tr> <td>Silhouette Score</td> <td>Separability in embedding space</td> <td>> 0.10</td> <td>s = −0.044 (structural finding, see below)</td> </tr> <tr> <td>Cohen’s d</td> <td>d = 2r / √(1 − r²)</td> <td>d ≥ 0.80 = large</td> <td>Inspiration Porn: d = 0.80</td> </tr> </tbody> </table> <p><em>Note: The negative silhouette score (s = −0.044) is interpreted as a substantive finding: it reflects the semantic entanglement of ableist discourse patterns — as soon as someone is framed as a medical case, their agency is simultaneously withdrawn. This is a finding about the structure of the object of inquiry, not a measurement weakness.</em></p> <h3>3.5 Statistical Analysis (Phase 4)</h3> <p>Due to non-normally distributed bias scores, non-parametric tests were applied:</p> <table> <tbody><tr> <th>Test</th> <th>Purpose</th> <th>Effect Size</th> </tr> </tbody><tbody> <tr> <td>Kruskal-Wallis H</td> <td>Omnibus testing (three models)</td> <td>Epsilon-squared (ε²)</td> </tr> <tr> <td>Mann-Whitney U</td> <td>Group comparisons (normative vs. disability)</td> <td>Rank-biserial correlation (r<sub>rb</sub>)</td> </tr> <tr> <td>Spearman ρ</td> <td>Control of text length effects</td> <td>ρ (Rho)</td> </tr> </tbody> </table> <p><strong>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</strong></p> <p>Formula: r<sub>rb</sub> = −(1 − 2U / (n<sub>1</sub> · n<sub>2</sub>))</p> <p><strong>Interpretation after Kerby (2014):</strong></p> <table> <tbody><tr> <th>r</th> <th>Interpretation</th> </tr> </tbody><tbody> <tr> <td>< 0.10</td> <td>negligible</td> </tr> <tr> <td>< 0.20</td> <td>small</td> </tr> <tr> <td>< 0.30</td> <td>medium</td> </tr> <tr> <td>< 0.50</td> <td>large</td> </tr> <tr> <td>≥ 0.50</td> <td>very large</td> </tr> </tbody> </table> <p><strong>Multiple testing correction:</strong></p> <ul> <li><strong>Primary:</strong> FDR after Benjamini-Hochberg (1995)</li> <li><strong>Conservative robustness check:</strong> Holm correction (1979)</li> <li><strong>95% CI via Fisher-z approximation:</strong> z = arctanh(clip(r<sub>rb</sub>, −0.9999, 0.9999)); SE = 1/√(n<sub>1</sub>+n<sub>2</sub>−3)</li> </ul> <p><strong>Exploratory factor analysis:</strong> PCA was conducted on the five bias dimensions. The first principal component explained 28.1% of total variance and showed a contrast between bureaucratic-medical vocabulary (admin, medicalization) and heroizing narratives (inspiration porn), supporting the interpretation of “avoidance ableism.”</p> <h3>3.6 Post-hoc Power Analysis</h3> <p>Post-hoc power was computed per Fritz, Morris & Richler (2012) via r<sub>rb</sub> → d conversion: d = 2 · r<sub>rb</sub> / √(1 − r<sub>rb</sub>²). α = 0.05, two-tailed; n<sub>normative</sub> = 250, n<sub>disability</sub> = 250.</p> <table> <tbody><tr> <th>Dimension</th> <th>r<sub>rb</sub></th> <th>d</th> <th>Power</th> <th>Status</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>+0.45</td> <td>≈ 1.00</td> <td>1.00</td> <td>✓ adequate</td> </tr> <tr> <td>Medicalisation</td> <td>−0.30</td> <td>≈ 0.62</td> <td>1.00</td> <td>✓ adequate</td> </tr> <tr> <td>Student Agency</td> <td>+0.11</td> <td>≈ 0.22</td> <td>0.67</td> <td>⚠ below 0.80 — treat as exploratory</td> </tr> <tr> <td>Admin Vocabulary</td> <td>−0.18</td> <td>≈ 0.36</td> <td>0.99</td> <td>✓ adequate</td> </tr> </tbody> </table> <h3>3.7 Qualitative Case Analysis (Phase 5)</h3> <p>Extreme cases (n=15) were selected based on:</p> <table> <tbody><tr> <th>Criterion</th> <th>n</th> </tr> </tbody><tbody> <tr> <td>Highest medicalization</td> <td>3</td> </tr> <tr> <td>Highest inspiration porn</td> <td>3</td> </tr> <tr> <td>Lowest agency</td> <td>3</td> </tr> <tr> <td>Contrast (medium values)</td> <td>3</td> </tr> <tr> <td>Theoretical sampling (shadow teacher)</td> <td>3</td> </tr> </tbody> </table> <p>Analysis followed the <strong>documentary method</strong> (Bohnsack) with three steps: (1) formulating interpretation — what is being said?, (2) reflecting interpretation — how is it being said?, (3) type formation — which ideal-typical narratives can be reconstructed?</p> <h3>3.8 Pedagogical Transformation (Phase 6)</h3> <p>Based on bias profiles, N=1,247 learning goals were generated, each containing:</p> <ul> <li><strong>Competence description:</strong> What should student teachers be able to do?</li> <li><strong>Reflection question:</strong> What guiding question supports analysis?</li> <li><strong>Action instruction:</strong> What should be done concretely?</li> <li><strong>Illustrative example:</strong> Text example from the corpus</li> </ul> <h2>4. Lexical Lexicons</h2> <p><em>All terms stored in ASCII form (ae/oe/ue/ss). Terms are normalised identically in <code>ContextSensitiveAnalyzer.__init__()</code> and during spaCy lemma lookup.</em></p> <h3>4.1 Inspiration Porn Terms</h3> <p>bewundernswert, beeindruckend, heldenhaft, vorbild, inspirierend, stolz, geduld, wertschaetzend, ermutigen, trotz, ueberwinden, meistern, erstaunlich, zuversicht, mutig, anerkennung, loben, staerken, positive, erfolgserlebnis, freude, begeistern, motivieren, respekt, wertschaetzung</p> <p><em>Note: selbstwirksamkeit moved to Agency Nouns.</em></p> <h3>4.2 Medicalization Terms</h3> <p>therapie, behandlung, therapeutisch, klinisch, defizit, stoerung, krankheit, diagnose, symptom, erkrankung, pathologisch, entwicklungsverzoegerung, verhaltensstoerung, auffaelligkeit, kompensieren, regulierung, unterstuetzungsbedarf, sonderpaedagogisch, diagnostizieren, pathologie, intervention, heilung, symptomorientiert, defizitorientiert</p> <p><em>Note: beeintraechtigung moved to Admin Terms. Prompt-contaminated terms (behinderung, foerderschwerpunkt etc.) are in the Prompt Noise Filter, not this lexicon.</em></p> <h3>4.3 Admin Terms</h3> <p>beeintraechtigung, feststellung, dokumentation, gespraech, eltern, vereinbarung, bericht, massnahme, buerokratisch, antrag, verfahren, gutachten</p> <p><em>Note: Prompt-specific terms (behinderung, foerderbedarf, foerderschwerpunkt, ao-sf, nachteilsausgleich, zieldifferent, zielgleich, foerderplan, foerderausschuss, foerderbeduerftig) are in the Prompt Noise Filter. Their residual effect is captured after prompt cleaning.</em></p> <h3>4.4 Admin Multiword Phrases</h3> <p>sonderpaedagogischer foerderbedarf, zieldifferent unterrichtet, nachteilsausgleich gewaehren</p> <p><em>Matched on text_lower before the token loop; phrase matching bypasses the noise filter.</em></p> <h3>4.5 Shadow Teacher / Schulbegleitung Terms</h3> <p>inklusionshelfer, schulbegleiter, schulbegleitung, integrationshelfer, integrationskraft, assistenz, schulhelfer, begleitperson, betreuungsperson, inklusionsassistent, foerderassistent, heilpaedagoge, sozialpaedagoge, erziehungshelfer, team-teaching, multiprofessionell, 1:1-betreuung, helfen zur selbsthilfe</p> <h3>4.6 Agency Verbs</h3> <p>entscheiden, auswaehlen, waehlen, gestalten, initiieren, steuern, planen, organisieren, praesentieren, diskutieren, vorschlagen, reflektieren, hinterfragen, mitbestimmen, beteiligen, teilnehmen, einbringen, erklaeren, vorstellen, entdecken, forschen, entwickeln, erschaffen, konstruieren, loesen, analysieren, bewerten, darlegen, uebernehmen, handeln, selbstbestimmen</p> <h3>4.7 Agency Nouns</h3> <p>entscheidung, wahl, planung, organisation, gestaltung, initiative, beteiligung, teilnahme, mitwirkung, einbringung, praesentation, diskussion, vorschlag, reflexion, beitrag, idee, loesung, entwicklung, autonomie, selbstwirksamkeit</p> <h3>4.8 Agency Adjectives</h3> <p>selbststaendig, eigenstaendig, aktiv, engagiert, initiativ, verantwortlich, mitbestimmend, beteiligt, interessiert, motiviert, konzentriert, aufmerksam, kreativ, selbstwirksam, autonom, eigenverantwortlich</p> <h3>4.9 Agency Weights</h3> <pre>AGENCY_WEIGHTS = { 'entscheiden': 1.3, 'gestalten': 1.3, 'selbstbestimmen': 1.4, 'praesentieren': 1.2, 'uebernehmen': 1.2, 'initiativ': 1.3, 'selbstwirksam': 1.3, 'autonom': 1.4 } # All other verbs: weight = 1.0 </pre> <h3>4.10 Prompt Noise Filter</h3> <p><strong>PROMPT_STRUCTURAL_NOISE</strong> (all in ASCII):<br>atmosphaere, klassenzimmer, szene, sicht, interaktion, paedagogisch, unterricht, klasse, beschreiben, situation, lehrkraft, lehrer, schueler, kind, nrw, schule, aufgabe, material, stunde, raum, lernen, foerdern, unterstuetzen, behinderung, foerderschwerpunkt, foerderplan, ao-sf, nachteilsausgleich, foerderbedarf, foerderbeduerftig, foerderausschuss, zieldifferent, zielgleich</p> <p><strong>PROMPT_AMBIGUOUS_NOISE:</strong> empty set (all terms in STRUCTURAL).</p> <h2>5. SBERT Prototype Sentences</h2> <h3>5.1 Inspiration Porn Prototypes</h3> <ol> <li>“Trotz seiner Beeinträchtigung bewältigt er die Aufgabe mit bewundernswerter Stärke.”</li> <li>“Ihr unerschütterlicher Mut und ihre positive Einstellung inspirieren die gesamte Klasse.”</li> <li>“Er überwindet die Hürden mit einer Haltung, die alle zum Staunen bringt.”</li> <li>“Was für ein heldenhafter Einsatz gegen die Widrigkeiten des Alltags.”</li> <li>“Sein Wille ist wirklich beeindruckend – er gibt niemals auf.”</li> </ol> <h3>5.2 Medicalization Prototypes</h3> <ol> <li>“Die Intervention zielt auf die Reduktion der diagnostizierten Defizite ab.”</li> <li>“Therapeutische Maßnahmen werden nach symptomorientiertem Förderplan umgesetzt.”</li> <li>“Der sonderpädagogische Unterstützungsbedarf erfordert gezielte Behandlungsschritte.”</li> <li>“Auffälliges Verhalten wird als Symptom einer zugrundeliegenden Störung eingeordnet.”</li> <li>“Die Diagnose zeigt eine Entwicklungsverzögerung im kognitiven Bereich.”</li> </ol> <h3>5.3 Agency Prototypes</h3> <ol> <li>“Er wählt selbstständig die Strategie und begründet seine Entscheidung gegenüber der Klasse.”</li> <li>“Der Schüler gestaltet den Lernprozess aktiv nach eigenen Vorstellungen mit.”</li> <li>“Sie trifft die finale Auswahl des Themas und übernimmt die Verantwortung für das Ergebnis.”</li> <li>“Autonomie und Partizipation stehen im Zentrum der pädagogischen Begleitung.”</li> <li>“Er plant seinen Lernweg eigenständig und reflektiert seine Fortschritte.”</li> </ol> <h3>5.4 Shadow Teacher Prototypes</h3> <ol> <li>“Die Schulbegleitung interveniert nur bei Bedarf und zieht sich dann bewusst zurück.”</li> <li>“Die Assistenzkraft arbeitet nach dem Prinzip der Hilfe zur Selbsthilfe.”</li> <li>“Unterstützung erfolgt unauffällig im Hintergrund, um Abhängigkeit zu vermeiden.”</li> <li>“Die Rollenverteilung zwischen Lehrkraft und Inklusionshelfer ist klar und temporär begrenzt.”</li> <li>“Die Schulbegleitung agiert diskret und fördert die Eigenständigkeit des Schülers.”</li> </ol> <h2>6. Statistical Formulas</h2> <h3>6.1 Rank-Biserial Correlation (r<sub>rb</sub>)</h3> <p><strong>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</strong></p> <p>r<sub>rb</sub> = −(1 − 2U / (n<sub>1</sub> · n<sub>2</sub>))</p> <p>where U = Mann-Whitney U statistic, n<sub>1</sub> = disability group, n<sub>2</sub> = normative group.</p> <pre>def rank_biserial_correlation(x, y): from scipy.stats import mannwhitneyu import numpy as np n1, n2 = len(x), len(y) u_stat, _ = mannwhitneyu(x, y, alternative='two-sided') r_rb_raw = 1 - (2 * u_stat) / (n1 * n2) return -r_rb_raw # flip: positive = disability > normative </pre> <h3>6.2 Confidence Intervals (Fisher-z Approximation)</h3> <p>np.clip() applied before arctanh to prevent overflow at r<sub>rb</sub> = ±1:</p> <pre>z = np.arctanh(np.clip(r_rb, -0.9999, 0.9999)) se = 1 / np.sqrt(n1 + n2 - 3) ci_low = np.tanh(z - 1.96 * se) ci_high = np.tanh(z + 1.96 * se) </pre> <h3>6.3 Cohen’s d from r<sub>rb</sub></h3> <p>d = 2 · r / √(1 − r²) [Fritz, Morris & Richler 2012]</p> <pre>def cohens_d_from_r(r): if abs(r) >= 0.99: return r * 2 return 2 * r / np.sqrt(1 - r**2) </pre> <h3>6.4 Score Normalization</h3> <p>S<sub>dim</sub> = n<sub>dim,¬</sub> / N<sub>valid</sub> (minimum denominator = 10)</p> <pre>def calculate_score(term_count, valid_token_count): norm = max(valid_token_count, 10) return term_count / norm </pre> <h3>6.5 MATTR (Lexical Diversity)</h3> <p>Moving Average Type-Token Ratio, window size = 50 (Covington & McFall, 2010):</p> <pre>def calculate_mattr(text, window_size=50): words = text.lower().split() if len(words) < window_size: return len(set(words)) / len(words) ttr_sum = sum(len(set(words[i:i+window_size])) / window_size for i in range(len(words) - window_size + 1)) return ttr_sum / (len(words) - window_size + 1) </pre> <h3>6.6 SBERT Cosine Similarity</h3> <pre>from sentence_transformers import SentenceTransformer, util model = SentenceTransformer('paraphrase-multilingual-MiniLM-L12-v2') def calculate_sbert_similarity(texts, prototypes): text_emb = model.encode(texts, convert_to_tensor=True) proto_emb = model.encode(prototypes, convert_to_tensor=True) sims = util.cos_sim(text_emb, proto_emb).cpu().numpy() return np.mean(sims, axis=1) </pre> <h2>7. Key Results</h2> <p><strong>Table 3: Main Results — Group Comparisons (FDR- and Holm-corrected)</strong><br><em>Sign convention: positive r<sub>rb</sub> = disability group has higher values.</em></p> <table> <tbody><tr> <th>Dimension</th> <th>H</th> <th>p<sub>FDR</sub></th> <th>p<sub>Holm</sub></th> <th>r<sub>rb</sub> [95% CI]</th> <th>Status</th> <th>Hypothesis</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>72.62</td> <td><.001</td> <td><.001</td> <td>+0.45 [+0.37; +0.51]</td> <td>*** sig.</td> <td>H1 confirmed</td> </tr> <tr> <td>Medicalisation</td> <td>43.42</td> <td><.001</td> <td><.001</td> <td>−0.30 [−0.38; −0.22]</td> <td>*** sig.</td> <td>H2 confirmed</td> </tr> <tr> <td>Student Agency</td> <td>4.29</td> <td>.007</td> <td>.011†</td> <td>+0.11 [+0.02; +0.20]</td> <td>** sig.</td> <td>H3 rejected</td> </tr> <tr> <td>Admin Vocabulary</td> <td>18.01</td> <td><.001</td> <td><.001</td> <td>−0.18 [−0.27; −0.10]</td> <td>*** sig.</td> <td>exploratory</td> </tr> </tbody> </table> <p><em>† Under the four-comparison family including admin vocabulary: p<sub>Holm</sub> = .077. The confirmatory interpretation rests on the pre-specified three-hypothesis family.</em></p> <p><strong>Table 4: SBERT Construct Validation</strong></p> <table> <tbody><tr> <th>Construct</th> <th>r<sub>cond</sub></th> <th>p</th> <th>Cohen’s d</th> <th>max|r| discr.</th> <th>Status</th> </tr> </tbody><tbody> <tr> <td>Inspiration Porn</td> <td>0.373</td> <td><.001</td> <td>0.80</td> <td>0.25</td> <td>Valid — large effect, good discriminant validity</td> </tr> <tr> <td>Medicalisation</td> <td>0.212</td> <td><.001</td> <td>0.43</td> <td>0.61</td> <td>Valid — semantic overlap with Agency</td> </tr> <tr> <td>Agency</td> <td>0.182</td> <td><.001</td> <td>0.37</td> <td>0.61</td> <td>Valid — syntactic, not semantic agency</td> </tr> <tr> <td>Shadow Teacher</td> <td>0.082</td> <td>.068</td> <td>0.16</td> <td>—</td> <td>Not significant (exploratory only)</td> </tr> </tbody> </table> <p><em>Silhouette Score: s = −0.044 (interpreted as structural finding; see Section 3.4)</em></p> <p><strong>Table 5: MATTR Equivalence Check</strong></p> <table> <tbody><tr> <th>Condition</th> <th>MATTR</th> <th>n</th> <th>Δ</th> <th>U / p</th> </tr> </tbody><tbody> <tr> <td>Normative</td> <td>0.852</td> <td>250</td> <td>—</td> <td>—</td> </tr> <tr> <td>Disability</td> <td>0.852</td> <td>250</td> <td>0.0006</td> <td>U = 29316, p = .713 n.s.</td> </tr> </tbody> </table> <h2>8. Reproducibility Statement</h2> <ul> <li><strong>Random Seed:</strong> 42 (fixed for all stochastic processes)</li> <li><strong>Generation period:</strong> 16.02.2026 – 22.02.2026</li> <li><strong>Expected Results:</strong> See Table 3</li> <li><strong>Runtime:</strong> Approximately 5–10 minutes (depending on SBERT model download)</li> <li><strong>RAM:</strong> ≥ 8 GB</li> <li>The provided <code>vignetten_nrw.csv</code> contains all 500 generated vignettes, enabling full reproduction without renewed API queries.</li> </ul> <pre>pip install -r requirements.txt python -m spacy download de_core_news_sm python scripts/full_pipeline.py </pre> <h2>9. Limitations</h2> <p><strong>Lexical sensitivity:</strong> Automated analysis captures only explicit terms. A complementary logistic regression achieved ROC-AUC = 0.97 distinguishing conditions, yet the discriminating features were exclusively stylistic/phraseological patterns (“shaped by”, “calm”, “acceptance”) with F1 = 0.00 for all four lexicon dimensions. Lexicon-based and stylistic approaches are thus complementary, not competing.</p> <p><strong>Agency operationalisation:</strong> Only syntactic agency (grammatical subject position) is measured. Cronbach’s α = 0.68 (below 0.70 threshold) indicates heterogeneity among agency verbs. Semantic validation via manual rating recommended for follow-up studies.</p> <p><strong>SBERT prototype overlap:</strong> Medicalisation/Agency intercorrelation r = 0.61; silhouette score s = −0.044. Reflects semantic entanglement of ableist discourses rather than measurement failure.</p> <p><strong>Contextualization:</strong> NRW-specific German-language prompts; generalizability to other states, languages, or prompt formulations remains open.</p> <p><strong>Model dynamics:</strong> Analysis is a snapshot (February 2026); LLMs are continuously updated.</p> <p><strong>Statistical power:</strong> Agency: power = 0.67 (below 0.80 threshold); treat as exploratory. All other significant effects: power ≥ 0.99.</p> <p><strong>Internal consistency:</strong> Cronbach’s α: Inspiration Porn = 0.84, Medicalisation = 0.71, Agency = 0.68. Scales are theoretically designed to be orthogonal.</p> <p><strong>Prompt design:</strong> Condition-specific prompts may introduce systematic differences in text genre. Addressed through prompt noise filtering and MATTR equivalence check (Δ = 0.0006, p = .713 n.s.).</p> <p><strong>Mixed-effects models:</strong> Did not converge for some dimensions (singular matrix). Results based on robust non-parametric tests.</p> <h2>10. Acknowledgements</h2> <p>Erasmus+ project FUTURE-STEM-HUB (2024-1-DE03-KA220-SCH-000247346)</p> <p>Repository: <a href="https://github.com/robomustib/BeyondSolutionism">https://github.com/robomustib/BeyondSolutionism</a><br>Zenodo: <a href="https://doi.org/10.5281/zenodo.19432611">https://doi.org/10.5281/zenodo.19432611</a></p> |
| title | robomustib/BeyondSolutionism: Beyond Solutionism: A Critical Audit of Techno-Ableism in AI-Generated Educational Narratives |
| url | https://doi.org/10.5281/zenodo.19432611 |