Cyclic Ablation: Testing Concept Localization against Functional Regeneration in AI
Fuente:
arXiv
Guardado en:
| Autor principal: | Kapelko, Eduard |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Active Inference Agency Formalization, Metrics, and Convergence Assessments
por: Kapelko, Eduard
Publicado: (2026)
por: Kapelko, Eduard
Publicado: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
por: Casademunt, Helena, et al.
Publicado: (2025)
por: Casademunt, Helena, et al.
Publicado: (2025)
Causality $\neq$ Invariance: Function and Concept Vectors in LLMs
por: Opiełka, Gustaw, et al.
Publicado: (2026)
por: Opiełka, Gustaw, et al.
Publicado: (2026)
Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures
por: Borobia, Hector, et al.
Publicado: (2026)
por: Borobia, Hector, et al.
Publicado: (2026)
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
por: Li, Maximilian, et al.
Publicado: (2023)
por: Li, Maximilian, et al.
Publicado: (2023)
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
por: Na, Clara, et al.
Publicado: (2024)
por: Na, Clara, et al.
Publicado: (2024)
Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering
por: Barbu, Eduard, et al.
Publicado: (2025)
por: Barbu, Eduard, et al.
Publicado: (2025)
MedConceptsQA: Open Source Medical Concepts QA Benchmark
por: Shoham, Ofir Ben, et al.
Publicado: (2024)
por: Shoham, Ofir Ben, et al.
Publicado: (2024)
Kernelized Concept Erasure
por: Ravfogel, Shauli, et al.
Publicado: (2022)
por: Ravfogel, Shauli, et al.
Publicado: (2022)
Vector Quantized Latent Concepts: A Scalable Alternative to Clustering-Based Concept Discovery
por: Yu, Xuemin, et al.
Publicado: (2026)
por: Yu, Xuemin, et al.
Publicado: (2026)
LLM Pretraining with Continuous Concepts
por: Tack, Jihoon, et al.
Publicado: (2025)
por: Tack, Jihoon, et al.
Publicado: (2025)
Towards Compositionality in Concept Learning
por: Stein, Adam, et al.
Publicado: (2024)
por: Stein, Adam, et al.
Publicado: (2024)
Linear Adversarial Concept Erasure
por: Ravfogel, Shauli, et al.
Publicado: (2022)
por: Ravfogel, Shauli, et al.
Publicado: (2022)
LCSB: Layer-Cyclic Selective Backpropagation for Memory-Efficient On-Device LLM Fine-Tuning
por: Park, Juneyoung, et al.
Publicado: (2026)
por: Park, Juneyoung, et al.
Publicado: (2026)
Concept Bottleneck Large Language Models
por: Sun, Chung-En, et al.
Publicado: (2024)
por: Sun, Chung-En, et al.
Publicado: (2024)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
por: Gadgil, Soham, et al.
Publicado: (2024)
por: Gadgil, Soham, et al.
Publicado: (2024)
AlignSAE: Concept-Aligned Sparse Autoencoders
por: Yang, Minglai, et al.
Publicado: (2025)
por: Yang, Minglai, et al.
Publicado: (2025)
Leveraging AI Graders for Missing Score Imputation to Achieve Accurate Ability Estimation in Constructed-Response Tests
por: Uto, Masaki, et al.
Publicado: (2025)
por: Uto, Masaki, et al.
Publicado: (2025)
Simple Mechanisms for Representing, Indexing and Manipulating Concepts
por: Li, Yuanzhi, et al.
Publicado: (2023)
por: Li, Yuanzhi, et al.
Publicado: (2023)
Beyond Individual Facts: Investigating Categorical Knowledge Locality of Taxonomy and Meronomy Concepts in GPT Models
por: Burger, Christopher, et al.
Publicado: (2024)
por: Burger, Christopher, et al.
Publicado: (2024)
Nonlinear Concept Erasure: a Density Matching Approach
por: Saillenfest, Antoine, et al.
Publicado: (2025)
por: Saillenfest, Antoine, et al.
Publicado: (2025)
Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
por: Karvonen, Adam, et al.
Publicado: (2024)
por: Karvonen, Adam, et al.
Publicado: (2024)
Khattat: Enhancing Readability and Concept Representation of Semantic Typography
por: Hussein, Ahmed, et al.
Publicado: (2024)
por: Hussein, Ahmed, et al.
Publicado: (2024)
Medical Concept Normalization in a Low-Resource Setting
por: Patzelt, Tim
Publicado: (2024)
por: Patzelt, Tim
Publicado: (2024)
Learning Machines: In Search of a Concept Oriented Language
por: Gunes, Veyis
Publicado: (2024)
por: Gunes, Veyis
Publicado: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
por: Rosati, Domenic, et al.
Publicado: (2024)
por: Rosati, Domenic, et al.
Publicado: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
por: Nakkiran, Preetum, et al.
Publicado: (2025)
por: Nakkiran, Preetum, et al.
Publicado: (2025)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
por: Yan, Xinyuan, et al.
Publicado: (2025)
por: Yan, Xinyuan, et al.
Publicado: (2025)
Concept Algebra for (Score-Based) Text-Controlled Generative Models
por: Wang, Zihao, et al.
Publicado: (2023)
por: Wang, Zihao, et al.
Publicado: (2023)
Can LLMs Learn New Concepts Incrementally without Forgetting?
por: Zheng, Junhao, et al.
Publicado: (2024)
por: Zheng, Junhao, et al.
Publicado: (2024)
CLUE: Concept-Level Uncertainty Estimation for Large Language Models
por: Wang, Yu-Hsiang, et al.
Publicado: (2024)
por: Wang, Yu-Hsiang, et al.
Publicado: (2024)
TempTest: Local Normalization Distortion and the Detection of Machine-generated Text
por: Kempton, Tom, et al.
Publicado: (2025)
por: Kempton, Tom, et al.
Publicado: (2025)
MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
por: von Oswald, Johannes, et al.
Publicado: (2025)
por: von Oswald, Johannes, et al.
Publicado: (2025)
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
por: Kudriashov, Sergei, et al.
Publicado: (2025)
por: Kudriashov, Sergei, et al.
Publicado: (2025)
Automated Knowledge Concept Annotation and Question Representation Learning for Knowledge Tracing
por: Ozyurt, Yilmazcan, et al.
Publicado: (2024)
por: Ozyurt, Yilmazcan, et al.
Publicado: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
por: Hong, Joey, et al.
Publicado: (2024)
por: Hong, Joey, et al.
Publicado: (2024)
Concept Tokens: Learning Behavioral Embeddings Through Concept Definitions
por: Sastre, Ignacio, et al.
Publicado: (2026)
por: Sastre, Ignacio, et al.
Publicado: (2026)
Adversarial Evasion Attack Efficiency against Large Language Models
por: Vitorino, João, et al.
Publicado: (2024)
por: Vitorino, João, et al.
Publicado: (2024)
Analogical Reasoning Inside Large Language Models: Concept Vectors and the Limits of Abstraction
por: Opiełka, Gustaw, et al.
Publicado: (2025)
por: Opiełka, Gustaw, et al.
Publicado: (2025)
Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
por: Yamashita, Tomoya, et al.
Publicado: (2025)
por: Yamashita, Tomoya, et al.
Publicado: (2025)
Ejemplares similares
-
Active Inference Agency Formalization, Metrics, and Convergence Assessments
por: Kapelko, Eduard
Publicado: (2026) -
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
por: Casademunt, Helena, et al.
Publicado: (2025) -
Causality $\neq$ Invariance: Function and Concept Vectors in LLMs
por: Opiełka, Gustaw, et al.
Publicado: (2026) -
Functional Component Ablation Reveals Specialization Patterns in Hybrid Language Model Architectures
por: Borobia, Hector, et al.
Publicado: (2026) -
Circuit Breaking: Removing Model Behaviors with Targeted Ablation
por: Li, Maximilian, et al.
Publicado: (2023)