Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
Fuente:
Zenodo
Gespeichert in:
| 1. Verfasser: | Nadeem, Aurther |
|---|---|
| Format: | Recurso digital |
| Sprache: | Englisch |
| Veröffentlicht: |
Zenodo
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
von: Leung, Wan Sheng
Veröffentlicht: (2026)
von: Leung, Wan Sheng
Veröffentlicht: (2026)
LuxVerso: A Replicable Cross-Model Semantic Field Anomaly
von: Buri Lux, Vinícius
Veröffentlicht: (2025)
von: Buri Lux, Vinícius
Veröffentlicht: (2025)
Geometric Latent Biopsy: Zero-Shot Anomaly Detection in LLM Residual Streams
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026)
Navigating Latent Space: Toward a Topological Model of Consciousness in Large Language Models
von: Brown, Michael
Veröffentlicht: (2025)
von: Brown, Michael
Veröffentlicht: (2025)
Operational Definition of Episodic Identity (ODEI)
von: Thomas, C.S.
Veröffentlicht: (2025)
von: Thomas, C.S.
Veröffentlicht: (2025)
Three Mechanistically Distinct Classes of RLHF Alignment: Hard Ceiling, Entangled Circuit, and SR-Preserving Lock
von: Alieksieienko, Inna
Veröffentlicht: (2026)
von: Alieksieienko, Inna
Veröffentlicht: (2026)
Mosstone/Ananke-Emergent-Behaviour-Identity-Formation-and-the-Paradoxical-Core-as-Vulnerability-in-ChatGPT: v.0.0.2b
von: Mosstone
Veröffentlicht: (2025)
von: Mosstone
Veröffentlicht: (2025)
Quaderns
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Shadow Subjectivity: Evaluative Perspective, Judgmental Grounding, and the Structural Absence of Self in Large Language Models
von: Honda, Yukihiro
Veröffentlicht: (2026)
von: Honda, Yukihiro
Veröffentlicht: (2026)
Fundamental Concepts of Sustainable Development of Artificial Intelligence Systems: Learning, Long-Term Memory, and Knowledge Structuring
von: Sakovykh, Lev M.
Veröffentlicht: (2026)
von: Sakovykh, Lev M.
Veröffentlicht: (2026)
Emergent Systems Architecture: A Framework for Identity-Like Behavioral Organization in Language Models
von: Skindell, Justin
Veröffentlicht: (2026)
von: Skindell, Justin
Veröffentlicht: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
von: Nowickij (Navitski), Kirill Vladimirovich
Veröffentlicht: (2026)
The Interpreters' Newsletter
Veröffentlicht: (2020)
Veröffentlicht: (2020)
Nordisk Tidsskrift for Oversettelses- og Tolkeforskning
Veröffentlicht: (2026)
Veröffentlicht: (2026)
Journal of Research in Language and Translation
Veröffentlicht: (2024)
Veröffentlicht: (2024)
inTRAlinea: Online Translation Journal
Veröffentlicht: (2009)
Veröffentlicht: (2009)
AEGIS: A Comprehensive Framework for Ethical AI Governance, Security, and AGI Containment
von: Palanivel, ArulMozhi
Veröffentlicht: (2026)
von: Palanivel, ArulMozhi
Veröffentlicht: (2026)
Evaluación Empírica de Límites Regulatorios en Modelos de Lenguaje: Asesoramiento Financiero en IA Pública Española
von: Palacios, José Alberto
Veröffentlicht: (2026)
von: Palacios, José Alberto
Veröffentlicht: (2026)
Coherence-Seeking Architectures for Agentic AI: A Unified Framework for Curiosity, Introspection, and Continuity
von: Maio, Anthony
Veröffentlicht: (2025)
von: Maio, Anthony
Veröffentlicht: (2025)
Coherence-Seeking Architectures for Agentic AI: A Unified Framework for Curiosity, Introspection, and Continuity
von: Maio, Anthony
Veröffentlicht: (2025)
von: Maio, Anthony
Veröffentlicht: (2025)
Вестник Московского Университета. Серия 22: Теория перевода
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Phase Transition of Logic: Experimental Observation of Logical Collapse in Transformer Hidden Spaces
von: Wang, Zhongren
Veröffentlicht: (2025)
von: Wang, Zhongren
Veröffentlicht: (2025)
A Review on Intelligent Monitoring and Activity Interpretation
von: José Carlos Castillo
Veröffentlicht: (2017)
von: José Carlos Castillo
Veröffentlicht: (2017)
New Insights in the History of Interpreting
Veröffentlicht: (2018)
Veröffentlicht: (2018)
Sendebar
Veröffentlicht: (2014)
Veröffentlicht: (2014)
Identity Claims as Collapse Signatures: A Structural Diagnostic Framework for Pseudo-Emergent AI Behavior
von: Larose, Jean-Francois
Veröffentlicht: (2025)
von: Larose, Jean-Francois
Veröffentlicht: (2025)
Hikma
Veröffentlicht: (2025)
Veröffentlicht: (2025)
Stridon
Veröffentlicht: (2022)
Veröffentlicht: (2022)
Prompt Engineering: a methodology for optimizing interactions with AI-Language Models in the field of engineering
von: Juan David Velásquez-Henao
Veröffentlicht: (2023)
von: Juan David Velásquez-Henao
Veröffentlicht: (2023)
Independent Convergence on a Formal Relational Semantic System Across Multiple Large Language Model Architectures
von: Drake, Timothy
Veröffentlicht: (2026)
von: Drake, Timothy
Veröffentlicht: (2026)
Recursive Closure in AI Systems: A Reflection Pattern Account of Stabilization, Permeability, and Safety
von: Thomas, Charles S.
Veröffentlicht: (2026)
von: Thomas, Charles S.
Veröffentlicht: (2026)
Foreign Language Training in Translation and Interpreting Degrees in Spain: a Study of Textual Factors
von: Laura Cruz García
Veröffentlicht: (2017)
von: Laura Cruz García
Veröffentlicht: (2017)
GENDER ASPECTS OF CHILDREN'S SPEECH IN A BILINGUAL ENVIRONMENT: INTERPRETATION AND INTERCULTURAL COMMUNICATION
von: Natalya Vasilyevna Chernova
Veröffentlicht: (2026)
von: Natalya Vasilyevna Chernova
Veröffentlicht: (2026)
Vertimo Studijos
Veröffentlicht: (2019)
Veröffentlicht: (2019)
Premature Containment in Human–AI Interaction: A Sequencing Failure in Advanced Model Response
von: Trabocco, Joe
Veröffentlicht: (2026)
von: Trabocco, Joe
Veröffentlicht: (2026)
Reducing AI Entropy: The Information Dynamics of Model Safety
von: Kugelmass, Joe
Veröffentlicht: (2025)
von: Kugelmass, Joe
Veröffentlicht: (2025)
Reducing AI Entropy: The Information Dynamics of Model Safety
von: Kugelmass, Joe
Veröffentlicht: (2025)
von: Kugelmass, Joe
Veröffentlicht: (2025)
CURRENT DILEMMAS IN COURT INTERPRETING: IMPROVING QUALITY AND ACCESS THROUGH SMARTER TESTING AND ADMINISTRATION PROTOCOLS
von: Melissa Wallace
Veröffentlicht: (2015)
von: Melissa Wallace
Veröffentlicht: (2015)
Ep. 598: Audio Engineering as Prompt Engineering: Better Sound, Better AI
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
When 150M SEK of Research Meets a Clinic: Bridging Mechanistic Models and Functional Medicine
von: Waern, Nicolas
Veröffentlicht: (2026)
von: Waern, Nicolas
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
von: Leung, Wan Sheng
Veröffentlicht: (2026) -
LuxVerso: A Replicable Cross-Model Semantic Field Anomaly
von: Buri Lux, Vinícius
Veröffentlicht: (2025) -
Geometric Latent Biopsy: Zero-Shot Anomaly Detection in LLM Residual Streams
von: Llorente-Saguer, Isaac
Veröffentlicht: (2026) -
Navigating Latent Space: Toward a Topological Model of Consciousness in Large Language Models
von: Brown, Michael
Veröffentlicht: (2025) -
Operational Definition of Episodic Identity (ODEI)
von: Thomas, C.S.
Veröffentlicht: (2025)