Language steering in latent space to mitigate unintended code-switching
Fuente:
arXiv
Saved in:
| Main Authors: | Goncharov, Andrey, Kondusov, Nikolai, Zaytsev, Alexey |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When an LLM is apprehensive about its answers -- and when its uncertainty is justified
by: Sychev, Petr, et al.
Published: (2025)
by: Sychev, Petr, et al.
Published: (2025)
Complexity-aware fine-tuning
by: Goncharov, Andrey, et al.
Published: (2025)
by: Goncharov, Andrey, et al.
Published: (2025)
Beyond Simple Averaging: Improving NLP Ensemble Performance with Topological-Data-Analysis-Based Weighting
by: Proskura, Polina, et al.
Published: (2024)
by: Proskura, Polina, et al.
Published: (2024)
MedSyn: LLM-based Synthetic Medical Text Generation Framework
by: Kumichev, Gleb, et al.
Published: (2024)
by: Kumichev, Gleb, et al.
Published: (2024)
ALIEN: Aligned Entropy Head for Improving Uncertainty Estimation of LLMs
by: Zabolotnyi, Artem, et al.
Published: (2025)
by: Zabolotnyi, Artem, et al.
Published: (2025)
AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection
by: Shingi, Geet, et al.
Published: (2021)
by: Shingi, Geet, et al.
Published: (2021)
Low-resource Machine Translation for Code-switched Kazakh-Russian Language Pair
by: Borisov, Maksim, et al.
Published: (2025)
by: Borisov, Maksim, et al.
Published: (2025)
Toward universal steering and monitoring of AI models
by: Beaglehole, Daniel, et al.
Published: (2025)
by: Beaglehole, Daniel, et al.
Published: (2025)
Enhancing the Reliability of Medical AI through Expert-guided Uncertainty Modeling
by: Khalin, Aleksei, et al.
Published: (2026)
by: Khalin, Aleksei, et al.
Published: (2026)
A gentle push funziona benissimo: making instructed models in Italian via contrastive activation steering
by: Scalena, Daniel, et al.
Published: (2024)
by: Scalena, Daniel, et al.
Published: (2024)
The evaluation of a code-switched Sepedi-English automatic speech recognition system
by: Phaladi, Amanda, et al.
Published: (2024)
by: Phaladi, Amanda, et al.
Published: (2024)
Can sparse autoencoders be used to decompose and interpret steering vectors?
by: Mayne, Harry, et al.
Published: (2024)
by: Mayne, Harry, et al.
Published: (2024)
Holistic Uncertainty Estimation For Open-Set Recognition
by: Erlygin, Leonid, et al.
Published: (2024)
by: Erlygin, Leonid, et al.
Published: (2024)
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Efficient semantic uncertainty quantification in language models via diversity-steered sampling
by: Park, Ji Won, et al.
Published: (2025)
by: Park, Ji Won, et al.
Published: (2025)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
by: Li, Pengyi, et al.
Published: (2025)
by: Li, Pengyi, et al.
Published: (2025)
Label Attention Network for Temporal Sets Prediction: You Were Looking at a Wrong Self-Attention
by: Kovtun, Elizaveta, et al.
Published: (2023)
by: Kovtun, Elizaveta, et al.
Published: (2023)
Uniting contrastive and generative learning for event sequences models
by: Yugay, Aleksandr, et al.
Published: (2024)
by: Yugay, Aleksandr, et al.
Published: (2024)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
Commute Your Domains: Trajectory Optimality Criterion for Multi-Domain Learning
by: Rukhovich, Alexey, et al.
Published: (2025)
by: Rukhovich, Alexey, et al.
Published: (2025)
Learn Your Reference Model for Real Good Alignment
by: Gorbatovski, Alexey, et al.
Published: (2024)
by: Gorbatovski, Alexey, et al.
Published: (2024)
Uncertainty Estimation for the Open-Set Text Classification systems
by: Erlygin, Leonid, et al.
Published: (2026)
by: Erlygin, Leonid, et al.
Published: (2026)
Foundation for unbiased cross-validation of spatio-temporal models for species distribution modeling
by: Koldasbayeva, Diana, et al.
Published: (2025)
by: Koldasbayeva, Diana, et al.
Published: (2025)
Projectable Models: One-Shot Generation of Small Specialized Transformers from Large Ones
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
Encode, Think, Decode: Scaling test-time reasoning with recursive latent thoughts
by: Koishekenov, Yeskendir, et al.
Published: (2025)
by: Koishekenov, Yeskendir, et al.
Published: (2025)
A comparison of latent semantic analysis and correspondence analysis of document-term matrices
by: Qi, Qianqian, et al.
Published: (2021)
by: Qi, Qianqian, et al.
Published: (2021)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
by: Rozanov, Nikolai, et al.
Published: (2024)
by: Rozanov, Nikolai, et al.
Published: (2024)
On Creating an English-Thai Code-switched Machine Translation in Medical Domain
by: Pengpun, Parinthapat, et al.
Published: (2024)
by: Pengpun, Parinthapat, et al.
Published: (2024)
Parameter-Efficient Neural CDEs via Implicit Function Jacobians
by: Kuleshov, Ilya, et al.
Published: (2025)
by: Kuleshov, Ilya, et al.
Published: (2025)
ESQA: Event Sequences Question Answering
by: Abdullaeva, Irina, et al.
Published: (2024)
by: Abdullaeva, Irina, et al.
Published: (2024)
Contextually Guided Transformers via Low-Rank Adaptation
by: Zhmoginov, Andrey, et al.
Published: (2025)
by: Zhmoginov, Andrey, et al.
Published: (2025)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
by: Li, Pengyi, et al.
Published: (2026)
by: Li, Pengyi, et al.
Published: (2026)
Where is the signal in tokenization space?
by: Geh, Renato Lui, et al.
Published: (2024)
by: Geh, Renato Lui, et al.
Published: (2024)
C-ing Clearly: Enhanced Binary Code Explanations using C code
by: Poncu, Teodor, et al.
Published: (2025)
by: Poncu, Teodor, et al.
Published: (2025)
Zero-knowledge LLM hallucination detection and mitigation through fine-grained cross-model consistency
by: Goel, Aman, et al.
Published: (2025)
by: Goel, Aman, et al.
Published: (2025)
Linear probes rely on textual evidence: Results from leakage mitigation studies in language models
by: Boxo, Gerard, et al.
Published: (2025)
by: Boxo, Gerard, et al.
Published: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
by: Sevriugov, Egor, et al.
Published: (2024)
by: Sevriugov, Egor, et al.
Published: (2024)
RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache Compression
by: Behnam, Payman, et al.
Published: (2025)
by: Behnam, Payman, et al.
Published: (2025)
The Unreasonable Ineffectiveness of the Deeper Layers
by: Gromov, Andrey, et al.
Published: (2024)
by: Gromov, Andrey, et al.
Published: (2024)
code_transformed: The Influence of Large Language Models on Code
by: Xu, Yuliang, et al.
Published: (2025)
by: Xu, Yuliang, et al.
Published: (2025)
Similar Items
-
When an LLM is apprehensive about its answers -- and when its uncertainty is justified
by: Sychev, Petr, et al.
Published: (2025) -
Complexity-aware fine-tuning
by: Goncharov, Andrey, et al.
Published: (2025) -
Beyond Simple Averaging: Improving NLP Ensemble Performance with Topological-Data-Analysis-Based Weighting
by: Proskura, Polina, et al.
Published: (2024) -
MedSyn: LLM-based Synthetic Medical Text Generation Framework
by: Kumichev, Gleb, et al.
Published: (2024) -
ALIEN: Aligned Entropy Head for Improving Uncertainty Estimation of LLMs
by: Zabolotnyi, Artem, et al.
Published: (2025)