Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
Fuente:
arXiv
Guardado en:
| Autores principales: | Salnikov, Mikhail, Korzh, Dmitrii, Lazichny, Ivan, Karimov, Elvir, Iudin, Artyom, Oseledets, Ivan, Rogov, Oleg Y., Loukachevitch, Natalia, Panchenko, Alexander, Tutubalina, Elena |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Certification of Speaker Recognition Models to Additive Perturbations
por: Korzh, Dmitrii, et al.
Publicado: (2024)
por: Korzh, Dmitrii, et al.
Publicado: (2024)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
por: Korzh, Dmitrii, et al.
Publicado: (2025)
por: Korzh, Dmitrii, et al.
Publicado: (2025)
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
por: Karimov, Elvir, et al.
Publicado: (2025)
por: Karimov, Elvir, et al.
Publicado: (2025)
LLM-Guided Prompt Evolution for Password Guessing
por: Mazin, Vladimir A., et al.
Publicado: (2026)
por: Mazin, Vladimir A., et al.
Publicado: (2026)
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
por: Dvirniak, Artem, et al.
Publicado: (2026)
por: Dvirniak, Artem, et al.
Publicado: (2026)
CLEAR: Character Unlearning in Textual and Visual Modalities
por: Dontsov, Alexey, et al.
Publicado: (2024)
por: Dontsov, Alexey, et al.
Publicado: (2024)
General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized Smoothing
por: Korzh, Dmitrii, et al.
Publicado: (2023)
por: Korzh, Dmitrii, et al.
Publicado: (2023)
Harnessing non-adversarial robustness in large language models
por: Zhou, Qinghua, et al.
Publicado: (2026)
por: Zhou, Qinghua, et al.
Publicado: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
por: Korznikov, Anton, et al.
Publicado: (2025)
por: Korznikov, Anton, et al.
Publicado: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
por: Korznikov, Anton, et al.
Publicado: (2026)
por: Korznikov, Anton, et al.
Publicado: (2026)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
por: Korznikov, Anton, et al.
Publicado: (2025)
por: Korznikov, Anton, et al.
Publicado: (2025)
Probabilistically Robust Watermarking of Neural Networks
por: Pautov, Mikhail, et al.
Publicado: (2024)
por: Pautov, Mikhail, et al.
Publicado: (2024)
Probabilistic Verification of Voice Anti-Spoofing Models
por: Kushnir, Evgeny, et al.
Publicado: (2026)
por: Kushnir, Evgeny, et al.
Publicado: (2026)
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
por: Galichin, Andrey, et al.
Publicado: (2025)
por: Galichin, Andrey, et al.
Publicado: (2025)
Spread them Apart: Towards Robust Watermarking of Generated Content
por: Pautov, Mikhail, et al.
Publicado: (2025)
por: Pautov, Mikhail, et al.
Publicado: (2025)
One Task Vector is not Enough: A Large-Scale Study for In-Context Learning
por: Tikhonov, Pavel, et al.
Publicado: (2025)
por: Tikhonov, Pavel, et al.
Publicado: (2025)
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
por: Galichin, Andrey V., et al.
Publicado: (2024)
por: Galichin, Andrey V., et al.
Publicado: (2024)
Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
por: Somov, Oleg, et al.
Publicado: (2026)
por: Somov, Oleg, et al.
Publicado: (2026)
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
por: Borodin, Kirill, et al.
Publicado: (2024)
por: Borodin, Kirill, et al.
Publicado: (2024)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
por: Seleznyov, Mikhail, et al.
Publicado: (2025)
HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
por: Krylov, Alexey, et al.
Publicado: (2025)
por: Krylov, Alexey, et al.
Publicado: (2025)
Konstruktor: A Strong Baseline for Simple Knowledge Graph Question Answering
por: Lysyuk, Maria, et al.
Publicado: (2024)
por: Lysyuk, Maria, et al.
Publicado: (2024)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
por: Alekseev, Artem, et al.
Publicado: (2025)
por: Alekseev, Artem, et al.
Publicado: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
por: Seleznyov, Mikhail, et al.
Publicado: (2026)
por: Seleznyov, Mikhail, et al.
Publicado: (2026)
Token-Level Density-Based Uncertainty Quantification Methods for Eliciting Truthfulness of Large Language Models
por: Vazhentsev, Artem, et al.
Publicado: (2025)
por: Vazhentsev, Artem, et al.
Publicado: (2025)
Exploring Prompt-Based Methods for Zero-Shot Hypernym Prediction with Large Language Models
por: Tikhomirov, Mikhail, et al.
Publicado: (2024)
por: Tikhomirov, Mikhail, et al.
Publicado: (2024)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
por: Vazhentsev, Artem, et al.
Publicado: (2026)
por: Vazhentsev, Artem, et al.
Publicado: (2026)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
por: Afonin, Nikita, et al.
Publicado: (2025)
por: Afonin, Nikita, et al.
Publicado: (2025)
A limited global perspective on what makes anatomical public engagement good or bad
por: Victoria Gomez, et al.
Publicado: (2025)
por: Victoria Gomez, et al.
Publicado: (2025)
Orwell: the good and the bad
por: Lyndsey Jenkins
Publicado: (2024)
por: Lyndsey Jenkins
Publicado: (2024)
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
por: Borisiuk, Anna, et al.
Publicado: (2026)
por: Borisiuk, Anna, et al.
Publicado: (2026)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
por: Braslavski, Pavel, et al.
Publicado: (2026)
por: Braslavski, Pavel, et al.
Publicado: (2026)
SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers
por: Chekalina, Viktoriia, et al.
Publicado: (2024)
por: Chekalina, Viktoriia, et al.
Publicado: (2024)
The good and the bad: fungi in Africa
por: R. H. Kurtzman Jr.
Publicado: (2011)
por: R. H. Kurtzman Jr.
Publicado: (2011)
Confidence Estimation for Error Detection in Text-to-SQL Systems
por: Somov, Oleg, et al.
Publicado: (2025)
por: Somov, Oleg, et al.
Publicado: (2025)
On the efficient preconditioning of the Stokes equations in tight geometries
por: Vladislav Pimanov, et al.
Publicado: (2024)
por: Vladislav Pimanov, et al.
Publicado: (2024)
On the Spatial Structure of Mixture-of-Experts in Transformers
por: Bershatsky, Daniel, et al.
Publicado: (2025)
por: Bershatsky, Daniel, et al.
Publicado: (2025)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
por: Mezentsev, Gleb, et al.
Publicado: (2025)
por: Mezentsev, Gleb, et al.
Publicado: (2025)
Associative memory and dead neurons
por: Fanaskov, Vladimir, et al.
Publicado: (2024)
por: Fanaskov, Vladimir, et al.
Publicado: (2024)
Inverted Activations: Reducing Memory Footprint in Neural Network Training
por: Novikov, Georgii, et al.
Publicado: (2024)
por: Novikov, Georgii, et al.
Publicado: (2024)
Ejemplares similares
-
Certification of Speaker Recognition Models to Additive Perturbations
por: Korzh, Dmitrii, et al.
Publicado: (2024) -
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
por: Korzh, Dmitrii, et al.
Publicado: (2025) -
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
por: Karimov, Elvir, et al.
Publicado: (2025) -
LLM-Guided Prompt Evolution for Password Guessing
por: Mazin, Vladimir A., et al.
Publicado: (2026) -
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
por: Dvirniak, Artem, et al.
Publicado: (2026)