Representation Engineering for Large-Language Models: Survey and Research Challenges
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bartoszcze, Lukasz, Munshi, Sarthak, Sukidi, Bryan, Yen, Jennifer, Yang, Zejia, Williams-King, David, Le, Linh, Asuzu, Kosi, Maple, Carsten |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Representation Noising: A Defence Mechanism Against Harmful Finetuning
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Immunization against harmful fine-tuning attacks
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
von: Williams-King, David, et al.
Veröffentlicht: (2025)
von: Williams-King, David, et al.
Veröffentlicht: (2025)
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
von: Wehner, Jan, et al.
Veröffentlicht: (2025)
von: Wehner, Jan, et al.
Veröffentlicht: (2025)
Individualised Counterfactual Examples Using Conformal Prediction Intervals
von: Adams, James M., et al.
Veröffentlicht: (2025)
von: Adams, James M., et al.
Veröffentlicht: (2025)
Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success
von: Maple, Carsten, et al.
Veröffentlicht: (2026)
von: Maple, Carsten, et al.
Veröffentlicht: (2026)
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
Justified Evidence Collection for Argument-based AI Fairness Assurance
von: Sabuncuoglu, Alpay, et al.
Veröffentlicht: (2025)
von: Sabuncuoglu, Alpay, et al.
Veröffentlicht: (2025)
Towards Robust Federated Analytics via Differentially Private Measurements of Statistical Heterogeneity
von: Scott, Mary, et al.
Veröffentlicht: (2024)
von: Scott, Mary, et al.
Veröffentlicht: (2024)
Private Federated Multiclass Post-hoc Calibration
von: Maddock, Samuel, et al.
Veröffentlicht: (2025)
von: Maddock, Samuel, et al.
Veröffentlicht: (2025)
FLAIM: AIM-based Synthetic Data Generation in the Federated Setting
von: Maddock, Samuel, et al.
Veröffentlicht: (2023)
von: Maddock, Samuel, et al.
Veröffentlicht: (2023)
DriveSafe: A Hierarchical Risk Taxonomy for Safety-Critical LLM-Based Driving Assistants
von: Kumar, Abhishek, et al.
Veröffentlicht: (2026)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2026)
Audio Computer-Assisted Self Interview Compared to Traditional Interview in an HIV-Related Behavioral Survey in Vietnam
von: Linh Cu Le
Veröffentlicht: (2012)
von: Linh Cu Le
Veröffentlicht: (2012)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
von: Le, Linh, et al.
Veröffentlicht: (2026)
von: Le, Linh, et al.
Veröffentlicht: (2026)
Threat, Risk and Mitigation Taxonomy for Digital Identity Systems
von: SHEIK, AL TARIQ, et al.
Veröffentlicht: (2024)
von: SHEIK, AL TARIQ, et al.
Veröffentlicht: (2024)
LearnedCache: An eBPF-Integrated Perceptron-Based Eviction Policy for the Linux Page Cache
von: Qi, Zejia
Veröffentlicht: (2026)
von: Qi, Zejia
Veröffentlicht: (2026)
Towards Smart Healthcare: Challenges and Opportunities in IoT and ML
von: Saifuzzaman, Munshi, et al.
Veröffentlicht: (2023)
von: Saifuzzaman, Munshi, et al.
Veröffentlicht: (2023)
Large Language Models and the Rationalist Empiricist Debate
von: King, David
Veröffentlicht: (2024)
von: King, David
Veröffentlicht: (2024)
Differentially Private Health Tokens for Estimating COVID-19 Risk
von: Butler, David, et al.
Veröffentlicht: (2020)
von: Butler, David, et al.
Veröffentlicht: (2020)
Data-Agnostic Face Image Synthesis Detection Using Bayesian CNNs
von: Leyva, Roberto, et al.
Veröffentlicht: (2024)
von: Leyva, Roberto, et al.
Veröffentlicht: (2024)
Operationalising Artificial Intelligence Bills of Materials (AIBOMs) for Verifiable AI Provenance and Lifecycle Assurance
von: Radanliev, Petar, et al.
Veröffentlicht: (2026)
von: Radanliev, Petar, et al.
Veröffentlicht: (2026)
Distributed, communication-efficient, and differentially private estimation of KL divergence
von: Scott, Mary, et al.
Veröffentlicht: (2024)
von: Scott, Mary, et al.
Veröffentlicht: (2024)
SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation
von: Radanliev, Petar, et al.
Veröffentlicht: (2026)
von: Radanliev, Petar, et al.
Veröffentlicht: (2026)
Field-Localized Forgery Detection for Digital Identity Documents
von: Kumar, Abhishek, et al.
Veröffentlicht: (2026)
von: Kumar, Abhishek, et al.
Veröffentlicht: (2026)
Detecting Face Synthesis Using a Concealed Fusion Model
von: Leyva, Roberto, et al.
Veröffentlicht: (2024)
von: Leyva, Roberto, et al.
Veröffentlicht: (2024)
SRA: Span Representation Alignment for Large Language Model Distillation
von: Dao, Quoc Phong, et al.
Veröffentlicht: (2026)
von: Dao, Quoc Phong, et al.
Veröffentlicht: (2026)
acad_recuperation_joueurs_exclus_fr-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
acad_self_excluded_player_recovery_en-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
acad_dispute_resolution_handbook_bilingual_fr-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
acad_rg_resource_compendium_bilingual_fr-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
acad_withdrawal_caps_high_rollers_en-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
acad_trustpilot_casinos_quebecois_fr-ca
von: Maple, Kevon
Veröffentlicht: (2026)
von: Maple, Kevon
Veröffentlicht: (2026)
Refugee Reception in Southern Africa
von: Maple, Nicholas
Veröffentlicht: (2024)
von: Maple, Nicholas
Veröffentlicht: (2024)
Manifold of Failure: Behavioral Attraction Basins in Language Models
von: Munshi, Sarthak, et al.
Veröffentlicht: (2026)
von: Munshi, Sarthak, et al.
Veröffentlicht: (2026)
Empirical and Sustainability Aspects of Software Engineering Research in the Era of Large Language Models: A Reflection
von: Williams, David, et al.
Veröffentlicht: (2025)
von: Williams, David, et al.
Veröffentlicht: (2025)
ACSE-Eval: Can LLMs threat model real-world cloud infrastructure?
von: Munshi, Sarthak, et al.
Veröffentlicht: (2025)
von: Munshi, Sarthak, et al.
Veröffentlicht: (2025)
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
von: Rao, Hongzhou, et al.
Veröffentlicht: (2025)
von: Rao, Hongzhou, et al.
Veröffentlicht: (2025)
A Game-Theoretic Approach for PMU Deployment Against False Data Injection Attacks
von: Maleki, Sajjad, et al.
Veröffentlicht: (2024)
von: Maleki, Sajjad, et al.
Veröffentlicht: (2024)
A privacy preserving querying mechanism with high utility for electric vehicles
von: Atmaca, Ugur Ilker, et al.
Veröffentlicht: (2022)
von: Atmaca, Ugur Ilker, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Representation Noising: A Defence Mechanism Against Harmful Finetuning
von: Rosati, Domenic, et al.
Veröffentlicht: (2024) -
Immunization against harmful fine-tuning attacks
von: Rosati, Domenic, et al.
Veröffentlicht: (2024) -
Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity
von: Williams-King, David, et al.
Veröffentlicht: (2025) -
Taxonomy, Opportunities, and Challenges of Representation Engineering for Large Language Models
von: Wehner, Jan, et al.
Veröffentlicht: (2025) -
Individualised Counterfactual Examples Using Conformal Prediction Intervals
von: Adams, James M., et al.
Veröffentlicht: (2025)