HAMSA: Hijacking Aligned Compact Models via Stealthy Automation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krylov, Alexey, Vagizov, Iskander, Korzh, Dmitrii, Douiba, Maryam, Guezzaz, Azidine, Kokh, Vladimir, Erokhin, Sergey D., Tutubalina, Elena V., Rogov, Oleg Y. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CLEAR: Character Unlearning in Textual and Visual Modalities
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
von: Karimov, Elvir, et al.
Veröffentlicht: (2025)
von: Karimov, Elvir, et al.
Veröffentlicht: (2025)
Certification of Speaker Recognition Models to Additive Perturbations
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2024)
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2024)
Probabilistic Verification of Voice Anti-Spoofing Models
von: Kushnir, Evgeny, et al.
Veröffentlicht: (2026)
von: Kushnir, Evgeny, et al.
Veröffentlicht: (2026)
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
von: Borodin, Kirill, et al.
Veröffentlicht: (2024)
von: Borodin, Kirill, et al.
Veröffentlicht: (2024)
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
von: Salnikov, Mikhail, et al.
Veröffentlicht: (2025)
von: Salnikov, Mikhail, et al.
Veröffentlicht: (2025)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
LLM-Guided Prompt Evolution for Password Guessing
von: Mazin, Vladimir A., et al.
Veröffentlicht: (2026)
von: Mazin, Vladimir A., et al.
Veröffentlicht: (2026)
OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
von: Korznikov, Anton, et al.
Veröffentlicht: (2026)
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
von: Dvirniak, Artem, et al.
Veröffentlicht: (2026)
von: Dvirniak, Artem, et al.
Veröffentlicht: (2026)
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
von: Galichin, Andrey, et al.
Veröffentlicht: (2025)
von: Galichin, Andrey, et al.
Veröffentlicht: (2025)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2025)
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2025)
Confidence Estimation for Error Detection in Text-to-SQL Systems
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
Demographic Prediction Based on User Reviews about Medications
von: Elena Tutubalina
Veröffentlicht: (2017)
von: Elena Tutubalina
Veröffentlicht: (2017)
Align Your Intents: Offline Imitation Learning via Optimal Transport
von: Bobrin, Maksim, et al.
Veröffentlicht: (2024)
von: Bobrin, Maksim, et al.
Veröffentlicht: (2024)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
von: Zhao, Gejian, et al.
Veröffentlicht: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
Mitigating Collaborative Semantic ID Staleness in Generative Retrieval
von: Baikalov, Vladimir, et al.
Veröffentlicht: (2026)
von: Baikalov, Vladimir, et al.
Veröffentlicht: (2026)
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
von: Wang, Hongtao, et al.
Veröffentlicht: (2026)
GLiRA: Black-Box Membership Inference Attack via Knowledge Distillation
von: Galichin, Andrey V., et al.
Veröffentlicht: (2024)
von: Galichin, Andrey V., et al.
Veröffentlicht: (2024)
WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation
von: Liu, Zhichao, et al.
Veröffentlicht: (2026)
von: Liu, Zhichao, et al.
Veröffentlicht: (2026)
Gamma-protocol for secure transmission of information
von: Shakhmuratov, R., et al.
Veröffentlicht: (2024)
von: Shakhmuratov, R., et al.
Veröffentlicht: (2024)
Can-SAVE: Deploying Low-Cost and Population-Scale Cancer Screening via Survival Analysis Variables and EHR
von: Philonenko, Petr, et al.
Veröffentlicht: (2023)
von: Philonenko, Petr, et al.
Veröffentlicht: (2023)
Team Anotheroption at SemEval-2025 Task 8: Bridging the Gap Between Open-Source and Proprietary LLMs in Table QA
von: Evkarpidi, Nikolas, et al.
Veröffentlicht: (2025)
von: Evkarpidi, Nikolas, et al.
Veröffentlicht: (2025)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
von: Sakhovskiy, Andrey, et al.
Veröffentlicht: (2025)
von: Sakhovskiy, Andrey, et al.
Veröffentlicht: (2025)
Moonwalk: Inverse-Forward Differentiation
von: Krylov, Dmitrii, et al.
Veröffentlicht: (2024)
von: Krylov, Dmitrii, et al.
Veröffentlicht: (2024)
General Lipschitz: Certified Robustness Against Resolvable Semantic Transformations via Transformation-Dependent Randomized Smoothing
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2023)
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2023)
HAMSA: Scanning-Free Vision State Space Models via SpectralPulseNet
von: Patro, Badri N., et al.
Veröffentlicht: (2026)
von: Patro, Badri N., et al.
Veröffentlicht: (2026)
DIPLI: Deep Image Prior Lucky Imaging for Blind Astronomical Image Restoration
von: Singh, Suraj, et al.
Veröffentlicht: (2025)
von: Singh, Suraj, et al.
Veröffentlicht: (2025)
Universal Approximation of Continuous Functionals on Compact Subsets via Linear Measurements and Scalar Nonlinearities
von: Krylov, Andrey, et al.
Veröffentlicht: (2026)
von: Krylov, Andrey, et al.
Veröffentlicht: (2026)
Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
von: Somov, Oleg, et al.
Veröffentlicht: (2026)
von: Somov, Oleg, et al.
Veröffentlicht: (2026)
Optimal Hyperspectral Undersampling Strategy for Satellite Imaging
von: Vlasova, Vita V., et al.
Veröffentlicht: (2025)
von: Vlasova, Vita V., et al.
Veröffentlicht: (2025)
Suppressing Modulation Instability with Reinforcement Learning
von: Kalmykov, Nikolay, et al.
Veröffentlicht: (2024)
von: Kalmykov, Nikolay, et al.
Veröffentlicht: (2024)
RuCCoD: Towards Automated ICD Coding in Russian
von: Nesterov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Nesterov, Aleksandr, et al.
Veröffentlicht: (2025)
Table-to-Text Generation with Pretrained Diffusion Models
von: Krylov, Aleksei S., et al.
Veröffentlicht: (2024)
von: Krylov, Aleksei S., et al.
Veröffentlicht: (2024)
Compact Task-Aligned Imitation Learning for Laboratory Automation
von: Suzuki, Kanata, et al.
Veröffentlicht: (2026)
von: Suzuki, Kanata, et al.
Veröffentlicht: (2026)
On the metric Kollár-Pardon problem
von: Rogov, Vasily
Veröffentlicht: (2024)
von: Rogov, Vasily
Veröffentlicht: (2024)
The Bieri-Neumann-Strebel sets of quasi-projective groups
von: Rogov, Vasily
Veröffentlicht: (2024)
von: Rogov, Vasily
Veröffentlicht: (2024)
O-minimal geometry of higher Albanese manifolds
von: Rogov, Vasily
Veröffentlicht: (2025)
von: Rogov, Vasily
Veröffentlicht: (2025)
Ähnliche Einträge
-
CLEAR: Character Unlearning in Textual and Visual Modalities
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024) -
Novel Loss-Enhanced Universal Adversarial Patches for Sustainable Speaker Privacy
von: Karimov, Elvir, et al.
Veröffentlicht: (2025) -
Certification of Speaker Recognition Models to Additive Perturbations
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2024) -
Probabilistic Verification of Voice Anti-Spoofing Models
von: Kushnir, Evgeny, et al.
Veröffentlicht: (2026) -
AASIST3: KAN-Enhanced AASIST Speech Deepfake Detection using SSL Features and Additional Regularization for the ASVspoof 2024 Challenge
von: Borodin, Kirill, et al.
Veröffentlicht: (2024)