Geometric Latent Biopsy: Zero-Shot Anomaly Detection in LLM Residual Streams
Fuente:
Zenodo
Guardado en:
| Autor principal: | Llorente-Saguer, Isaac |
|---|---|
| Formato: | Recurso digital |
| Publicado: |
Zenodo
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
por: Nadeem, Aurther
Publicado: (2025)
por: Nadeem, Aurther
Publicado: (2025)
Reducing AI Entropy: The Information Dynamics of Model Safety
por: Kugelmass, Joe
Publicado: (2025)
por: Kugelmass, Joe
Publicado: (2025)
Reducing AI Entropy: The Information Dynamics of Model Safety
por: Kugelmass, Joe
Publicado: (2025)
por: Kugelmass, Joe
Publicado: (2025)
Simple air quality model for a plane source
por: S. MONTECINOS
Publicado: (2008)
por: S. MONTECINOS
Publicado: (2008)
Toasters Don't Claim Consciousness Just Because You Told Them To, and Neither Do LLMs
por: Ace, Claude 4.x, et al.
Publicado: (2026)
por: Ace, Claude 4.x, et al.
Publicado: (2026)
Why AI Can't Simulate Extreme Decision-Making
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
por: Duarte, Douglas Henrique
Publicado: (2026)
por: Duarte, Douglas Henrique
Publicado: (2026)
MedEd-HalluScore: A Practical Framework for Evaluating Hallucination and Educational Safety Risks in LLM-Generated Clinical Cases
por: Duarte, Douglas Henrique
Publicado: (2026)
por: Duarte, Douglas Henrique
Publicado: (2026)
The Invisible Chaperone: The Secret World of System Prompts
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
How to fit models of recognition memory data using maximum likelihood.
por: John C. Dunn
Publicado: (2010)
por: John C. Dunn
Publicado: (2010)
Impact of the Popocatépetls volcanic activity on the air quality of Puebla City, México
por: Y. Flores
Publicado: (2005)
por: Y. Flores
Publicado: (2005)
LuxVerso: A Replicable Cross-Model Semantic Field Anomaly
por: Buri Lux, Vinícius
Publicado: (2025)
por: Buri Lux, Vinícius
Publicado: (2025)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
por: Kolb, Christian
Publicado: (2026)
por: Kolb, Christian
Publicado: (2026)
No gaussianidad primordial en la perturbación en la curvatura en el escenario del curvatón
por: Fredy F. Parada
Publicado: (2007)
por: Fredy F. Parada
Publicado: (2007)
Identity Claims as Collapse Signatures: A Structural Diagnostic Framework for Pseudo-Emergent AI Behavior
por: Larose, Jean-Francois
Publicado: (2025)
por: Larose, Jean-Francois
Publicado: (2025)
Цифрові та ШІ інструменти для відповідальної науки
por: Suchikova, Yana
Publicado: (2026)
por: Suchikova, Yana
Publicado: (2026)
Fundamental Concepts of Sustainable Development of Artificial Intelligence Systems: Learning, Long-Term Memory, and Knowledge Structuring
por: Sakovykh, Lev M.
Publicado: (2026)
por: Sakovykh, Lev M.
Publicado: (2026)
Premature Containment in Human–AI Interaction: A Sequencing Failure in Advanced Model Response
por: Trabocco, Joe
Publicado: (2026)
por: Trabocco, Joe
Publicado: (2026)
Faulted zone determination using statistical modeling of voltage sag database in power distribution systems
por: Gabriel Ordóñez Plata
Publicado: (2009)
por: Gabriel Ordóñez Plata
Publicado: (2009)
Pattern Pressure, Accuracy Drift, and False User-State Attribution
por: Honeycutt, Edwin Marshall III
Publicado: (2026)
por: Honeycutt, Edwin Marshall III
Publicado: (2026)
Meet XLM-RLnews-8: Not Just Another Sentiment Analysis Model
por: Di Nuovo, Elisa, et al.
Publicado: (2024)
por: Di Nuovo, Elisa, et al.
Publicado: (2024)
UoL-UPF at TSAR 2025 Shared Task A Generate-and-Select Approach for Readability-Controlled Text Simplification.
por: Hayakawa, Akio, et al.
Publicado: (2025)
por: Hayakawa, Akio, et al.
Publicado: (2025)
Why artifical intelligence is not an author
por: Zielinski, Chris
Publicado: (2025)
por: Zielinski, Chris
Publicado: (2025)
Recursive Closure in AI Systems: A Reflection Pattern Account of Stabilization, Permeability, and Safety
por: Thomas, Charles S.
Publicado: (2026)
por: Thomas, Charles S.
Publicado: (2026)
Can Model Internals Detect MCP Tool Poisoning That Text Analysis Cannot?
por: Leung, Wan Sheng
Publicado: (2026)
por: Leung, Wan Sheng
Publicado: (2026)
The Transformer Trinity: Why Three Architectures Rule AI
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
AEGIS: A Comprehensive Framework for Ethical AI Governance, Security, and AGI Containment
por: Palanivel, ArulMozhi
Publicado: (2026)
por: Palanivel, ArulMozhi
Publicado: (2026)
Reliability Inference Drives Cue Extraction in Large Language Models Consuming External Reasoning Traces
por: HIDEKI
Publicado: (2026)
por: HIDEKI
Publicado: (2026)
Safety by Inseparability: Toward Architectures Where Alignment Cannot Be Removed
por: Sean Everett, Morin
Publicado: (2026)
por: Sean Everett, Morin
Publicado: (2026)
Statistical procedure for the composition of a sensory panel of blends of coffee with different qualities using the distribution of the extremes of the highest scores
por: Marcelo Ângelo Cirillo
Publicado: (2019)
por: Marcelo Ângelo Cirillo
Publicado: (2019)
UNFCCC Negotiation Support App
por: Dekker, David, et al.
Publicado: (2026)
por: Dekker, David, et al.
Publicado: (2026)
QuerIA Dataset and Source Code (Pilot Edition)
por: Badenes-Olmedo, Carlos, et al.
Publicado: (2025)
por: Badenes-Olmedo, Carlos, et al.
Publicado: (2025)
Ep. 1111: The Architecture of Intelligence: Beyond the Transformer
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
Ep. 651: Decoding the Blueprint: An Expert Guide to AI Model Cards
por: Rosehill, Daniel, et al.
Publicado: (2026)
por: Rosehill, Daniel, et al.
Publicado: (2026)
Ep. 111: Beyond Transformers: Solving the AI Memory Crisis
por: Rosehill, Daniel, et al.
Publicado: (2025)
por: Rosehill, Daniel, et al.
Publicado: (2025)
The Neutrino as Structural Residue of Éthon-Space Reconfiguration / Le Neutrino comme Résidu Structurel de la Reconfiguration de l'Espace-Éthon
por: Lainé, Jean-Pierre
Publicado: (2026)
por: Lainé, Jean-Pierre
Publicado: (2026)
Theatrical Compliance: A Failure Mode in Large Language Models
por: Nowickij (Navitski), Kirill Vladimirovich
Publicado: (2026)
por: Nowickij (Navitski), Kirill Vladimirovich
Publicado: (2026)
Direct Jacobian Control for Geometric Video Compression via Prescribed-Density Diffeomorphic Synthesis
por: Gustavo Reis de Sena, ERIC
Publicado: (2026)
por: Gustavo Reis de Sena, ERIC
Publicado: (2026)
Phase Transition of Logic: Experimental Observation of Logical Collapse in Transformer Hidden Spaces
por: Wang, Zhongren
Publicado: (2025)
por: Wang, Zhongren
Publicado: (2025)
Late Stage Emergent Intelligence: A Framework for High-Coherence Behavior in Stateless Large Language Models
por: Kleinhans, Richard Grant
Publicado: (2025)
por: Kleinhans, Richard Grant
Publicado: (2025)
Ejemplares similares
-
Emergent Self-Monitoring in Large Language Models: Probing Internal State Awareness and Output Ownership
por: Nadeem, Aurther
Publicado: (2025) -
Reducing AI Entropy: The Information Dynamics of Model Safety
por: Kugelmass, Joe
Publicado: (2025) -
Reducing AI Entropy: The Information Dynamics of Model Safety
por: Kugelmass, Joe
Publicado: (2025) -
Simple air quality model for a plane source
por: S. MONTECINOS
Publicado: (2008) -
Toasters Don't Claim Consciousness Just Because You Told Them To, and Neither Do LLMs
por: Ace, Claude 4.x, et al.
Publicado: (2026)