Consistency Training while Mitigating Obfuscation via Rate Matching
Fuente:
arXiv
Saved in:
| Main Authors: | Imran, Sohaib, Gupta, Prakhar, Elstner, Jannes, Africa, David Demitri |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026)
by: Ivanov, Igor, et al.
Published: (2026)
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
by: Rivera, Joshua Fonseca, et al.
Published: (2025)
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025)
by: Weiss, Yuval, et al.
Published: (2025)
Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
by: Imran, Sohaib, et al.
Published: (2025)
by: Imran, Sohaib, et al.
Published: (2025)
Meta-Pretraining for Zero-Shot Cross-Lingual Named Entity Recognition in Low-Resource Philippine Languages
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
by: Wollschläger, Tom, et al.
Published: (2025)
by: Wollschläger, Tom, et al.
Published: (2025)
Mitigating Self-Preference by Authorship Obfuscation
by: Mahbub, Taslim, et al.
Published: (2025)
by: Mahbub, Taslim, et al.
Published: (2025)
Pico: A Modular Framework for Hypothesis-Driven Small Language Model Research
by: Martinez, Richard Diehl, et al.
Published: (2025)
by: Martinez, Richard Diehl, et al.
Published: (2025)
Are LLM Belief Updates Consistent with Bayes' Theorem?
by: Imran, Sohaib, et al.
Published: (2025)
by: Imran, Sohaib, et al.
Published: (2025)
No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probes
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
by: Cencerrado, Iván Vicente Moreno, et al.
Published: (2025)
Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
by: Tan, Daniel, et al.
Published: (2025)
by: Tan, Daniel, et al.
Published: (2025)
Identifying a Circuit for Verb Conjugation in GPT-2
by: Africa, David Demitri
Published: (2025)
by: Africa, David Demitri
Published: (2025)
Online Training of Large Language Models: Learn while chatting
by: Liang, Juhao, et al.
Published: (2024)
by: Liang, Juhao, et al.
Published: (2024)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
by: Montalan, Jann Railey, et al.
Published: (2025)
by: Montalan, Jann Railey, et al.
Published: (2025)
Does Self-Evaluation Enable Wireheading in Language Models?
by: Africa, David Demitri, et al.
Published: (2025)
by: Africa, David Demitri, et al.
Published: (2025)
Beyond the Parameters: A Technical Survey of Contextual Enrichment in Large Language Models: From In-Context Prompting to Causal Retrieval-Augmented Generation
by: Bansal, Prakhar, et al.
Published: (2026)
by: Bansal, Prakhar, et al.
Published: (2026)
Personalized Author Obfuscation with Large Language Models
by: Shokri, Mohammad, et al.
Published: (2025)
by: Shokri, Mohammad, et al.
Published: (2025)
Evaluating and Understanding Scheming Propensity in LLM Agents
by: Hopman, Mia, et al.
Published: (2026)
by: Hopman, Mia, et al.
Published: (2026)
Leveraging Machine-Generated Rationales to Facilitate Social Meaning Detection in Conversations
by: Dutt, Ritam, et al.
Published: (2024)
by: Dutt, Ritam, et al.
Published: (2024)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
by: Paul, Indraneil, et al.
Published: (2025)
by: Paul, Indraneil, et al.
Published: (2025)
Reducing Political Manipulation with Consistency Training
by: Phan, Long, et al.
Published: (2026)
by: Phan, Long, et al.
Published: (2026)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024)
by: Mohseni, Seyedreza, et al.
Published: (2024)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
by: Khouja, Jude, et al.
Published: (2025)
by: Khouja, Jude, et al.
Published: (2025)
Evaluating Role-Consistency in LLMs for Counselor Training
by: Rudolph, Eric, et al.
Published: (2026)
by: Rudolph, Eric, et al.
Published: (2026)
Post-Training Language Models for Crosslingual Consistency
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
by: Liang, Yu, et al.
Published: (2026)
by: Liang, Yu, et al.
Published: (2026)
ArXiv-to-Model: A Practical Study of Scientific LM Training
by: Gupta, Anuj
Published: (2026)
by: Gupta, Anuj
Published: (2026)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit
by: Goddard, Charles, et al.
Published: (2025)
by: Goddard, Charles, et al.
Published: (2025)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
by: Hackl, Veronika, et al.
Published: (2023)
by: Hackl, Veronika, et al.
Published: (2023)
Improving Training Efficiency and Reducing Maintenance Costs via Language Specific Model Merging
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
Resource-Efficient Fine-Tuning of LLaMA-3.2-3B for Medical Chain-of-Thought Reasoning
by: Mansha, Imran
Published: (2025)
by: Mansha, Imran
Published: (2025)
"OK Aura, Be Fair With Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection
by: López, Fernando, et al.
Published: (2026)
by: López, Fernando, et al.
Published: (2026)
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
by: Chua, James, et al.
Published: (2024)
by: Chua, James, et al.
Published: (2024)
Revisiting In-Context Learning with Long Context Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
Using Machine Learning to Enhance the Detection of Obfuscated Abusive Words in Swahili: A Focus on Child Safety
by: Nabangi, Phyllis, et al.
Published: (2026)
by: Nabangi, Phyllis, et al.
Published: (2026)
Semantic Consistency for Assuring Reliability of Large Language Models
by: Raj, Harsh, et al.
Published: (2023)
by: Raj, Harsh, et al.
Published: (2023)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
by: Kumar, Rajeev, et al.
Published: (2025)
by: Kumar, Rajeev, et al.
Published: (2025)
Similar Items
-
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
by: Ivanov, Igor, et al.
Published: (2026) -
Steering Awareness: Detecting Activation Steering from Within
by: Rivera, Joshua Fonseca, et al.
Published: (2025) -
Learning Dynamics of Meta-Learning in Small Model Pretraining
by: Africa, David Demitri, et al.
Published: (2025) -
Investigating ReLoRA: Effects on the Learning Dynamics of Small Language Models
by: Weiss, Yuval, et al.
Published: (2025) -
Out-of-Context Abduction: LLMs Make Inferences About Procedural Data Leveraging Declarative Facts in Earlier Training Data
by: Imran, Sohaib, et al.
Published: (2025)