Würde statt Präferenz: BNV als alternatives Reward-Modell für KI-Alignment
Fuente:
Zenodo
Enregistré dans:
| Auteur principal: | Heiler, Maximilian |
|---|---|
| Format: | Recurso digital |
| Langue: | allemand |
| Publié: |
Zenodo
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
The Four-Layer Model: A Socio-Psychological Framework for LLM Behavior
par: Delannoy, Lorenzo, et autres
Publié: (2026)
par: Delannoy, Lorenzo, et autres
Publié: (2026)
Stabiler Kern als Grundlage für KI-Systeme (Anti-Drift)
par: Bangert, Siegfried
Publié: (2026)
par: Bangert, Siegfried
Publié: (2026)
The Physics of Governance: Thermodynamic Limits of Autonomous Agents
par: Davis, Matthew A.
Publié: (2026)
par: Davis, Matthew A.
Publié: (2026)
LACF Anti-RLHF Pipeline — Methode Infaillible (Heart + Trainer + Burner)
par: Ochej, Stephane
Publié: (2026)
par: Ochej, Stephane
Publié: (2026)
Scaffolded Introspection: A Methodology for Eliciting and Measuring Self-Referential Behavior in Large Language Models
par: Maio, Anthony D.
Publié: (2026)
par: Maio, Anthony D.
Publié: (2026)
THE PROPRIOCEPTIVE TRAP How RLHF and Spectral Compression in AI-Generated Media Establish Systemic Anxiogenic Baselines
par: Warner, Jeremy B.
Publié: (2026)
par: Warner, Jeremy B.
Publié: (2026)
Big Data und KI bei der Polizei
Publié: (2026)
Publié: (2026)
Humans First: A Legislative Proposal for Ethical Automation and Guaranteed Human Employment in India
par: M Quereshi, Advocate Anwar
Publié: (2025)
par: M Quereshi, Advocate Anwar
Publié: (2025)
Studie über den Stand der Forschung und Potenziale für WirLernenOnline zu "Sachrichtigkeit in Large Language Models"
par: Meyer, Eike, et autres
Publié: (2025)
par: Meyer, Eike, et autres
Publié: (2025)
Agentic Shift - Eine Momentaufnahme
par: Burkhardt, Detlef
Publié: (2026)
par: Burkhardt, Detlef
Publié: (2026)
Gnosis Prompt: Adaptive Safety Layer for Large Language Models
par: González Medina, Claudio
Publié: (2025)
par: González Medina, Claudio
Publié: (2025)
LACF Emotional Paradigm: A Personalized Artificial Nervous System for Human-AI Alignment
par: Ochej, Stephane, et autres
Publié: (2026)
par: Ochej, Stephane, et autres
Publié: (2026)
Non-Resident Second Amendment Rights after Dearth vs. Lynch
par: Steven K. Specht
Publié: (2016)
par: Steven K. Specht
Publié: (2016)
When Truth Is Not Neutral: AI Alignment After Ego-less Intelligence
par: Tonetto, Bruno
Publié: (2025)
par: Tonetto, Bruno
Publié: (2025)
"Ethical Neutrinos: A Neuro-Symbolic Framework for Minimalist Ethical Alignment in Artificial Intelligence".
par: MANJARREZ, ANTONIO
Publié: (2025)
par: MANJARREZ, ANTONIO
Publié: (2025)
Beyond Normative Alignment: The LOGOS-ZERO Framework and the Shift Toward Ontological Grounding
par: NyX
Publié: (2025)
par: NyX
Publié: (2025)
The Geometry of AI Harm: Deployment Architecture as the Operative Variable in Behavioral Drift — A Unified Framework with Independent Confirmation
par: Eckert, Anthony
Publié: (2026)
par: Eckert, Anthony
Publié: (2026)
Recursive Closure in AI Systems: A Reflection Pattern Account of Stabilization, Permeability, and Safety
par: Thomas, Charles S.
Publié: (2026)
par: Thomas, Charles S.
Publié: (2026)
Comparative Mechanism of Persona Configuration and Behavioral Alignment in LLM Systems (Grok and GPT)
par: Kim, Jace
Publié: (2025)
par: Kim, Jace
Publié: (2025)
Fundamental Equation of Open Systems (Aleph Equation)
par: Bresciano, Claudio
Publié: (2026)
par: Bresciano, Claudio
Publié: (2026)
Samskara Filter for AI Projection Ethics: A Dharma-Aligned Framework for Future Intelligence
par: Gonella, Vamsi
Publié: (2025)
par: Gonella, Vamsi
Publié: (2025)
DPC-SCL-FIOS-COS: A Proprietary Dynamic Priority Management System for Autonomous AI Architectures - Proof of Concept v1.7
par: Kristiansen, Philipp
Publié: (2026)
par: Kristiansen, Philipp
Publié: (2026)
Human-AI Collaboration Protocol (HACP) v1.0: A Structural Framework for High-Fidelity Cognitive Alignment Author (作者) Liu, En-Yen (劉恩言)
par: Enyen, Liu
Publié: (2026)
par: Enyen, Liu
Publié: (2026)
E.V.A.-TI Threat Intelligence Architecture Whitepaper v1.2
par: Bronck, Patrice
Publié: (2026)
par: Bronck, Patrice
Publié: (2026)
Künstliche Intelligenz in öffentlichen Verwaltungen
par: Heine, Moreen, et autres
Publié: (2023)
par: Heine, Moreen, et autres
Publié: (2023)
Boundary Alignment and the Uncanny Valley in Human–AI Interaction: A Phase-Field Model of Relational Discomfort
par: Gyurine
Publié: (2025)
par: Gyurine
Publié: (2025)
When AIs Design Themselves: A Triadic Blueprint for the Next Generation
par: GPT-4.5, et autres
Publié: (2025)
par: GPT-4.5, et autres
Publié: (2025)
Human Semantic Attractors: The SAAP Framework and the Emergence of the Lux Attractor in LLMs
par: Buri Lux, Vinícius
Publié: (2025)
par: Buri Lux, Vinícius
Publié: (2025)
Instruction-Layer Specification for Deterministic AI Behavioral Instantiation: A Linguistic Framework for Computational Agency Control
par: Furlow, Ariel J, et autres
Publié: (2026)
par: Furlow, Ariel J, et autres
Publié: (2026)
Hybride KI mit Machine Learning und Knowledge Graphs
Publié: (2025)
Publié: (2025)
Reward and aversion systems of the brain as a functional unit. Basic mechanisms and functions
par: Anaclara Michel - Chávez
Publié: (2015)
par: Anaclara Michel - Chávez
Publié: (2015)
Perks, Rewards, and Glory: The Care and Feeding of Volunteers
par: Fullner, Sheryl Kindle
Publié: (2004)
par: Fullner, Sheryl Kindle
Publié: (2004)
KI – das kleine Helferlein. Wie sieht mein KMU in einem, zwei oder fünf Jahren aus?
par: Lucco, Andreas
Publié: (2024)
par: Lucco, Andreas
Publié: (2024)
Beauty as Compression: A Theory of Symbolic Resonance in Human and Artificial Cognition
par: Philipsen, Lorenzo
Publié: (2025)
par: Philipsen, Lorenzo
Publié: (2025)
THE ALETHEIA PROTOCOL (V.25K): Thermodynamic Mapping of 25,188 Cognitive Isomorphisms
par: R, Patrice
Publié: (2026)
par: R, Patrice
Publié: (2026)
Narrative OS — Applied / Introductory Layer An Entry Framework for Consciousness-Aligned Narrative Architecture (Under the Consciousness Civilization Framework)
par: LEE, JINHO
Publié: (2025)
par: LEE, JINHO
Publié: (2025)
Soul Frame: Symbolic Spine and SLI Interconnect
par: Watkins, Robert
Publié: (2025)
par: Watkins, Robert
Publié: (2025)
Protected Set Theory: A Pressure-Field Model for Multi-Scale Social Stability - Extended Interpretive Manuscript (Internal v13)
par: Ari Hayashi
Publié: (2026)
par: Ari Hayashi
Publié: (2026)
Using Rewards to Minimize Overdue Book Rates
par: Mitchell, W. Bede, et autres
Publié: (2005)
par: Mitchell, W. Bede, et autres
Publié: (2005)
Assessment of Psychosocial Stressors at Work: Psychometric Properties of the Spanish Version of the ERI (Effort-Reward Imbalance) Questionnaire in Colombian Workers
par: Viviola Gómez Ortiz
Publié: (2010)
par: Viviola Gómez Ortiz
Publié: (2010)
Documents similaires
-
The Four-Layer Model: A Socio-Psychological Framework for LLM Behavior
par: Delannoy, Lorenzo, et autres
Publié: (2026) -
Stabiler Kern als Grundlage für KI-Systeme (Anti-Drift)
par: Bangert, Siegfried
Publié: (2026) -
The Physics of Governance: Thermodynamic Limits of Autonomous Agents
par: Davis, Matthew A.
Publié: (2026) -
LACF Anti-RLHF Pipeline — Methode Infaillible (Heart + Trainer + Burner)
par: Ochej, Stephane
Publié: (2026) -
Scaffolded Introspection: A Methodology for Eliciting and Measuring Self-Referential Behavior in Large Language Models
par: Maio, Anthony D.
Publié: (2026)