Contrast-CAT: Contrasting Activations for Enhanced Interpretability in Transformer-based Text Classifiers
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Sungmin, Lee, Jeonghyun, Lee, Sangkyun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Context-Enhanced Contrastive Search for Improved LLM Text Generation
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
di: Sen, Jaydip, et al.
Pubblicazione: (2025)
Efficient Process Reward Modeling via Contrastive Mutual Information
di: Lee, Nakyung, et al.
Pubblicazione: (2026)
di: Lee, Nakyung, et al.
Pubblicazione: (2026)
Contrastive Token-level Explanations for Graph-based Rumour Detection
di: Chin, Daniel Wai Kit, et al.
Pubblicazione: (2025)
di: Chin, Daniel Wai Kit, et al.
Pubblicazione: (2025)
An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis
di: Dao, Phuong Q., et al.
Pubblicazione: (2025)
di: Dao, Phuong Q., et al.
Pubblicazione: (2025)
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
Enhancing Context Through Contrast
di: Ambilduke, Kshitij, et al.
Pubblicazione: (2024)
di: Ambilduke, Kshitij, et al.
Pubblicazione: (2024)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
di: Phan, Phuc, et al.
Pubblicazione: (2024)
di: Phan, Phuc, et al.
Pubblicazione: (2024)
Contrastive Instruction Tuning
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2024)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2024)
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation
di: Hsu, I-Hung, et al.
Pubblicazione: (2024)
di: Hsu, I-Hung, et al.
Pubblicazione: (2024)
DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning
di: Guo, Xun, et al.
Pubblicazione: (2024)
di: Guo, Xun, et al.
Pubblicazione: (2024)
Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification
di: Yu, Simon, et al.
Pubblicazione: (2023)
di: Yu, Simon, et al.
Pubblicazione: (2023)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
di: Lin, Zicheng, et al.
Pubblicazione: (2024)
IITK at SemEval-2024 Task 1: Contrastive Learning and Autoencoders for Semantic Textual Relatedness in Multilingual Texts
di: Basak, Udvas, et al.
Pubblicazione: (2024)
di: Basak, Udvas, et al.
Pubblicazione: (2024)
Automatic Pair Construction for Contrastive Post-training
di: Xu, Canwen, et al.
Pubblicazione: (2023)
di: Xu, Canwen, et al.
Pubblicazione: (2023)
Evolutionary Contrastive Distillation for Language Model Alignment
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
di: Katz-Samuels, Julian, et al.
Pubblicazione: (2024)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
di: Yin, Yueqin, et al.
Pubblicazione: (2024)
di: Yin, Yueqin, et al.
Pubblicazione: (2024)
Contrastive Identification and Generation in the Limit
di: Li, Xiaoyu, et al.
Pubblicazione: (2026)
di: Li, Xiaoyu, et al.
Pubblicazione: (2026)
HU at SemEval-2024 Task 8A: Can Contrastive Learning Learn Embeddings to Detect Machine-Generated Text?
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2024)
CLIPO: Contrastive Learning in Policy Optimization Generalizes RLVR
di: Cui, Sijia, et al.
Pubblicazione: (2026)
di: Cui, Sijia, et al.
Pubblicazione: (2026)
Circuit Insights: Towards Interpretability Beyond Activations
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
di: Golimblevskaia, Elena, et al.
Pubblicazione: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
di: Luo, Haozheng, et al.
Pubblicazione: (2026)
Anchored Preference Optimization and Contrastive Revisions: Addressing Underspecification in Alignment
di: D'Oosterlinck, Karel, et al.
Pubblicazione: (2024)
di: D'Oosterlinck, Karel, et al.
Pubblicazione: (2024)
CELL your Model: Contrastive Explanations for Large Language Models
di: Luss, Ronny, et al.
Pubblicazione: (2024)
di: Luss, Ronny, et al.
Pubblicazione: (2024)
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings
di: Hallee, Logan, et al.
Pubblicazione: (2024)
di: Hallee, Logan, et al.
Pubblicazione: (2024)
Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling
di: Zhuang, Chengxu, et al.
Pubblicazione: (2024)
di: Zhuang, Chengxu, et al.
Pubblicazione: (2024)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
di: Sriram, Aniruddh, et al.
Pubblicazione: (2024)
di: Sriram, Aniruddh, et al.
Pubblicazione: (2024)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
di: Tsymbalov, Aleksandr, et al.
Pubblicazione: (2025)
di: Tsymbalov, Aleksandr, et al.
Pubblicazione: (2025)
Debiasing Text Safety Classifiers through a Fairness-Aware Ensemble
di: Sturman, Olivia, et al.
Pubblicazione: (2024)
di: Sturman, Olivia, et al.
Pubblicazione: (2024)
Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts
di: Zhang, Yifan, et al.
Pubblicazione: (2024)
di: Zhang, Yifan, et al.
Pubblicazione: (2024)
Towards LLM-guided Causal Explainability for Black-box Text Classifiers
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2023)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2023)
Measuring the Depth of LLM Unlearning via Activation Patching
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
di: Lee, Jaeung, et al.
Pubblicazione: (2026)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
di: Lin, Hender
Pubblicazione: (2025)
di: Lin, Hender
Pubblicazione: (2025)
Confidence Regularized Masked Language Modeling using Text Length
di: Ji, Seunghyun, et al.
Pubblicazione: (2025)
di: Ji, Seunghyun, et al.
Pubblicazione: (2025)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
di: Ko, Jongwoo, et al.
Pubblicazione: (2025)
di: Ko, Jongwoo, et al.
Pubblicazione: (2025)
Contrastive Decoding for Synthetic Data Generation in Low-Resource Language Modeling
di: Ulm, Jannek, et al.
Pubblicazione: (2025)
di: Ulm, Jannek, et al.
Pubblicazione: (2025)
Sequence-level Large Language Model Training with Contrastive Preference Optimization
di: Feng, Zhili, et al.
Pubblicazione: (2025)
di: Feng, Zhili, et al.
Pubblicazione: (2025)
ConLID: Supervised Contrastive Learning for Low-Resource Language Identification
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
di: Foroutan, Negar, et al.
Pubblicazione: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
What's New in My Data? Novelty Exploration via Contrastive Generation
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Context-Enhanced Contrastive Search for Improved LLM Text Generation
di: Sen, Jaydip, et al.
Pubblicazione: (2025) -
Efficient Process Reward Modeling via Contrastive Mutual Information
di: Lee, Nakyung, et al.
Pubblicazione: (2026) -
Contrastive Token-level Explanations for Graph-based Rumour Detection
di: Chin, Daniel Wai Kit, et al.
Pubblicazione: (2025) -
An Enhanced Dual Transformer Contrastive Network for Multimodal Sentiment Analysis
di: Dao, Phuong Q., et al.
Pubblicazione: (2025) -
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)