DAPI: Domain Adaptive Toxicity Probe Vector Intervention for Fine-Grained Detoxification
Fuente:
arXiv
Saved in:
| Main Authors: | Hyeonsu, Cho, Kim, Dooyoung, Ko, Youngjoong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue Generation
by: Shin, Seungmin, et al.
Published: (2025)
by: Shin, Seungmin, et al.
Published: (2025)
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
by: Lai-Lopez, Nicole, et al.
Published: (2025)
by: Lai-Lopez, Nicole, et al.
Published: (2025)
Dependency Parsing with the Structuralized Prompt Template
by: Kim, Keunha, et al.
Published: (2025)
by: Kim, Keunha, et al.
Published: (2025)
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025)
by: Kang, Joonseong, et al.
Published: (2025)
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
by: Lu, Huimin, et al.
Published: (2025)
by: Lu, Huimin, et al.
Published: (2025)
SS-MPC: A Sequence-Structured Multi-Party Conversation System
by: Jang, Yoonjin, et al.
Published: (2025)
by: Jang, Yoonjin, et al.
Published: (2025)
ixi-GEN: Efficient Industrial sLLMs through Domain Adaptive Continual Pretraining
by: Kim, Seonwu, et al.
Published: (2025)
by: Kim, Seonwu, et al.
Published: (2025)
Test-Time Detoxification without Training or Learning Anything
by: Saglam, Baturay, et al.
Published: (2026)
by: Saglam, Baturay, et al.
Published: (2026)
Relational Knowledge Distillation Using Fine-tuned Function Vectors
by: Kang, Andrea, et al.
Published: (2026)
by: Kang, Andrea, et al.
Published: (2026)
Beyond Steering Vector: Flow-based Activation Steering for Inference-Time Intervention
by: Jin, Zehao, et al.
Published: (2026)
by: Jin, Zehao, et al.
Published: (2026)
How Reliable are Causal Probing Interventions?
by: Canby, Marc, et al.
Published: (2024)
by: Canby, Marc, et al.
Published: (2024)
Beyond Independent Passages: Adaptive Passage Combination Retrieval for Retrieval Augmented Open-Domain Question Answering
by: Ko, Ting-Wen, et al.
Published: (2025)
by: Ko, Ting-Wen, et al.
Published: (2025)
Revisiting Adaptive Rounding with Vectorized Reparameterization for LLM Quantization
by: Zhou, Yuli, et al.
Published: (2026)
by: Zhou, Yuli, et al.
Published: (2026)
Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
by: Yu, Jing, et al.
Published: (2025)
by: Yu, Jing, et al.
Published: (2025)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
by: Kim, Seungone, et al.
Published: (2023)
by: Kim, Seungone, et al.
Published: (2023)
Decoding Dense Embeddings: Sparse Autoencoders for Interpreting and Discretizing Dense Retrieval
by: Park, Seongwan, et al.
Published: (2025)
by: Park, Seongwan, et al.
Published: (2025)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
by: Lee, Changhun, et al.
Published: (2024)
by: Lee, Changhun, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
Instruction Fine-Tuning: Does Prompt Loss Matter?
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
On the Loss of Context-awareness in General Instruction Fine-tuning
by: Wang, Yihan, et al.
Published: (2024)
by: Wang, Yihan, et al.
Published: (2024)
Domain-Adaptive Pre-Training for Arabic Aspect-Based Sentiment Analysis: A Comparative Study of Domain Adaptation and Fine-Tuning Strategies
by: Alyami, Salha, et al.
Published: (2025)
by: Alyami, Salha, et al.
Published: (2025)
SVFT: Parameter-Efficient Fine-Tuning with Singular Vectors
by: Lingam, Vijay, et al.
Published: (2024)
by: Lingam, Vijay, et al.
Published: (2024)
CHAMP: A Competition-level Dataset for Fine-Grained Analyses of LLMs' Mathematical Reasoning Capabilities
by: Mao, Yujun, et al.
Published: (2024)
by: Mao, Yujun, et al.
Published: (2024)
RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning
by: Ji, Deyi, et al.
Published: (2025)
by: Ji, Deyi, et al.
Published: (2025)
Fine-tuning Large Language Models for Domain-specific Machine Translation
by: Zheng, Jiawei, et al.
Published: (2024)
by: Zheng, Jiawei, et al.
Published: (2024)
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
by: Li, Ziheng, et al.
Published: (2026)
by: Li, Ziheng, et al.
Published: (2026)
Revealing Fine-Grained Values and Opinions in Large Language Models
by: Wright, Dustin, et al.
Published: (2024)
by: Wright, Dustin, et al.
Published: (2024)
Can We Use Probing to Better Understand Fine-tuning and Knowledge Distillation of the BERT NLU?
by: Hościłowicz, Jakub, et al.
Published: (2023)
by: Hościłowicz, Jakub, et al.
Published: (2023)
MEUV: Achieving Fine-Grained Capability Activation in Large Language Models via Mutually Exclusive Unlock Vectors
by: Tong, Xin, et al.
Published: (2025)
by: Tong, Xin, et al.
Published: (2025)
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
by: Bronzini, Marco, et al.
Published: (2025)
by: Bronzini, Marco, et al.
Published: (2025)
Fine-Grained Modeling of Narrative Context: A Coherence Perspective via Retrospective Questions
by: Xu, Liyan, et al.
Published: (2024)
by: Xu, Liyan, et al.
Published: (2024)
On the Utility of Domain-Adjacent Fine-Tuned Model Ensembles for Few-shot Problems
by: Alam, Md Ibrahim Ibne, et al.
Published: (2024)
by: Alam, Md Ibrahim Ibne, et al.
Published: (2024)
Scaling Laws for Fine-Grained Mixture of Experts
by: Krajewski, Jakub, et al.
Published: (2024)
by: Krajewski, Jakub, et al.
Published: (2024)
Domain-Adaptive Continued Pre-Training of Small Language Models
by: Faroz, Salman
Published: (2025)
by: Faroz, Salman
Published: (2025)
XMoE: Sparse Models with Fine-grained and Adaptive Expert Selection
by: Yang, Yuanhang, et al.
Published: (2024)
by: Yang, Yuanhang, et al.
Published: (2024)
DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
by: Jung, Sunghee, et al.
Published: (2025)
by: Jung, Sunghee, et al.
Published: (2025)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
by: Jang, Yoonjin, et al.
Published: (2026)
by: Jang, Yoonjin, et al.
Published: (2026)
Domain-Adaptive Small Language Models for Structured Tax Code Prediction
by: Nath, Souvik, et al.
Published: (2025)
by: Nath, Souvik, et al.
Published: (2025)
Similar Items
-
ECO Decoding: Entropy-Based Control for Controllability and Fluency in Controllable Dialogue Generation
by: Shin, Seungmin, et al.
Published: (2025) -
ylmmcl at Multilingual Text Detoxification 2025: Lexicon-Guided Detoxification and Classifier-Gated Rewriting
by: Lai-Lopez, Nicole, et al.
Published: (2025) -
Dependency Parsing with the Structuralized Prompt Template
by: Kim, Keunha, et al.
Published: (2025) -
Adaptive Task Vectors for Large Language Models
by: Kang, Joonseong, et al.
Published: (2025) -
DetoxLLM: A Framework for Detoxification with Explanations
by: Khondaker, Md Tawkat Islam, et al.
Published: (2024)