Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders
Fuente:
arXiv
Saved in:
| Main Authors: | Kuznetsov, Kristian, Kushnareva, Laida, Druzhinina, Polina, Razzhigaev, Anton, Voznyuk, Anastasia, Piontkovskaya, Irina, Burnaev, Evgeny, Barannikov, Serguei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Quantifying Logical Consistency in Transformers via Query-Key Alignment
by: Tulchinskii, Eduard, et al.
Published: (2025)
by: Tulchinskii, Eduard, et al.
Published: (2025)
Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
by: Tulchinskii, Eduard, et al.
Published: (2023)
by: Tulchinskii, Eduard, et al.
Published: (2023)
Listening to the Wise Few: Select-and-Copy Attention Heads for Multiple-Choice QA
by: Tulchinskii, Eduard, et al.
Published: (2024)
by: Tulchinskii, Eduard, et al.
Published: (2024)
Robust AI-Generated Text Detection by Restricted Embeddings
by: Kuznetsov, Kristian, et al.
Published: (2024)
by: Kuznetsov, Kristian, et al.
Published: (2024)
AI-generated text boundary detection with RoFT
by: Kushnareva, Laida, et al.
Published: (2023)
by: Kushnareva, Laida, et al.
Published: (2023)
Improving Interpretability and Robustness for the Detection of AI-Generated Images
by: Gaintseva, Tatiana, et al.
Published: (2024)
by: Gaintseva, Tatiana, et al.
Published: (2024)
MindShift: Analyzing Language Models' Reactions to Psychological Prompts
by: Vasiliuk, Anton, et al.
Published: (2025)
by: Vasiliuk, Anton, et al.
Published: (2025)
AudioSAE: Towards Understanding of Audio-Processing Models with Sparse AutoEncoders
by: Aparin, Georgii, et al.
Published: (2026)
by: Aparin, Georgii, et al.
Published: (2026)
Disentanglement Learning via Topology
by: Balabin, Nikita, et al.
Published: (2023)
by: Balabin, Nikita, et al.
Published: (2023)
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
by: Pedashenko, Vladislav, et al.
Published: (2025)
by: Pedashenko, Vladislav, et al.
Published: (2025)
The Density of Cross-Persistence Diagrams and Its Applications
by: Mironenko, Alexander, et al.
Published: (2026)
by: Mironenko, Alexander, et al.
Published: (2026)
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
by: Galichin, Andrey, et al.
Published: (2025)
by: Galichin, Andrey, et al.
Published: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
by: Rahmatullaev, Temurbek, et al.
Published: (2025)
by: Rahmatullaev, Temurbek, et al.
Published: (2025)
Loss Barcode: A Topological Measure of Escapability in Loss Landscapes
by: Barannikov, Serguei, et al.
Published: (2020)
by: Barannikov, Serguei, et al.
Published: (2020)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
Barcodes as Summary of Loss Function Topology
by: Barannikov, Serguei, et al.
Published: (2019)
by: Barannikov, Serguei, et al.
Published: (2019)
RTD-Lite: Scalable Topological Analysis for Comparing Weighted Graphs in Learning Tasks
by: Tulchinskii, Eduard, et al.
Published: (2025)
by: Tulchinskii, Eduard, et al.
Published: (2025)
A Method for Auto-Differentiation of the Voronoi Tessellation
by: Shumilin, Sergei, et al.
Published: (2023)
by: Shumilin, Sergei, et al.
Published: (2023)
Scalar Function Topology Divergence: Comparing Topology of 3D Objects
by: Trofimov, Ilya, et al.
Published: (2024)
by: Trofimov, Ilya, et al.
Published: (2024)
Edge-wise Topological Divergence Gaps: Guiding Search in Combinatorial Optimization
by: Trofimov, Ilya, et al.
Published: (2025)
by: Trofimov, Ilya, et al.
Published: (2025)
DeepPavlov at SemEval-2024 Task 8: Leveraging Transfer Learning for Detecting Boundaries of Machine-Generated Texts
by: Voznyuk, Anastasia, et al.
Published: (2024)
by: Voznyuk, Anastasia, et al.
Published: (2024)
The Shape of Learning: Anisotropy and Intrinsic Dimensions in Transformer-Based Models
by: Razzhigaev, Anton, et al.
Published: (2023)
by: Razzhigaev, Anton, et al.
Published: (2023)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
by: Sevriugov, Egor, et al.
Published: (2024)
by: Sevriugov, Egor, et al.
Published: (2024)
Sparse Autoencoders for Sequential Recommendation Models: Interpretation and Flexible Control
by: Klenitskiy, Anton, et al.
Published: (2025)
by: Klenitskiy, Anton, et al.
Published: (2025)
Addressing Hallucinations in Language Models with Knowledge Graph Embeddings as an Additional Modality
by: Chekalina, Viktoriia, et al.
Published: (2024)
by: Chekalina, Viktoriia, et al.
Published: (2024)
Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
Are AI Detectors Good Enough? A Survey on Quality of Datasets With Machine-Generated Texts
by: Gritsai, German, et al.
Published: (2024)
by: Gritsai, German, et al.
Published: (2024)
Advacheck at GenAI Detection Task 1: AI Detection Powered by Domain-Aware Multi-Tasking
by: Gritsai, German, et al.
Published: (2024)
by: Gritsai, German, et al.
Published: (2024)
Understanding Internal Representations of Recommendation Models with Sparse Autoencoders
by: Wang, Jiayin, et al.
Published: (2024)
by: Wang, Jiayin, et al.
Published: (2024)
Learning Retrieval Models with Sparse Autoencoders
by: Formal, Thibault, et al.
Published: (2026)
by: Formal, Thibault, et al.
Published: (2026)
Harnessing Feature Clustering For Enhanced Anomaly Detection With Variational Autoencoder And Dynamic Threshold
by: Ale, Tolulope, et al.
Published: (2024)
by: Ale, Tolulope, et al.
Published: (2024)
Encode Me If You Can: Learning Universal User Representations via Event Sequence Autoencoding
by: Klenitskiy, Anton, et al.
Published: (2025)
by: Klenitskiy, Anton, et al.
Published: (2025)
From Knots to Knobs: Towards Steerable Collaborative Filtering Using Sparse Autoencoders
by: Spišák, Martin, et al.
Published: (2026)
by: Spišák, Martin, et al.
Published: (2026)
Real-World Transferable Adversarial Attack on Face-Recognition Systems
by: Kaznacheev, Andrey, et al.
Published: (2025)
by: Kaznacheev, Andrey, et al.
Published: (2025)
Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval
by: Park, Seongwan, et al.
Published: (2025)
by: Park, Seongwan, et al.
Published: (2025)
CCT-Code: Cross-Consistency Training for Multilingual Clone Detection and Code Search
by: Tikhonov, Anton, et al.
Published: (2023)
by: Tikhonov, Anton, et al.
Published: (2023)
PersonalAI: A Systematic Comparison of Knowledge Graph Storage and Retrieval Approaches for Personalized LLM agents
by: Menschikov, Mikhail, et al.
Published: (2025)
by: Menschikov, Mikhail, et al.
Published: (2025)
Enhancing Fake-News Detection with Node-Level Topological Features
by: Xu, Kaiyuan
Published: (2025)
by: Xu, Kaiyuan
Published: (2025)
Variational Autoencoder for Channel Estimation: Real-World Measurement Insights
by: Baur, Michael, et al.
Published: (2023)
by: Baur, Michael, et al.
Published: (2023)
Commute Your Domains: Trajectory Optimality Criterion for Multi-Domain Learning
by: Rukhovich, Alexey, et al.
Published: (2025)
by: Rukhovich, Alexey, et al.
Published: (2025)
Similar Items
-
Quantifying Logical Consistency in Transformers via Query-Key Alignment
by: Tulchinskii, Eduard, et al.
Published: (2025) -
Intrinsic Dimension Estimation for Robust Detection of AI-Generated Texts
by: Tulchinskii, Eduard, et al.
Published: (2023) -
Listening to the Wise Few: Select-and-Copy Attention Heads for Multiple-Choice QA
by: Tulchinskii, Eduard, et al.
Published: (2024) -
Robust AI-Generated Text Detection by Restricted Embeddings
by: Kuznetsov, Kristian, et al.
Published: (2024) -
AI-generated text boundary detection with RoFT
by: Kushnareva, Laida, et al.
Published: (2023)