Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xing, Xie, Huiyuan, Wang, Yiyan, Xiao, Chaojun, Chen, Huimin, Sargeant, Holli, Steffek, Felix, Shao, Jie, Liu, Zhiyuan, Sun, Maosong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
by: Sargeant, Holli, et al.
Published: (2024)
by: Sargeant, Holli, et al.
Published: (2024)
LLM vs. Lawyers: Identifying a Subset of Summary Judgments in a Large UK Case Law Dataset
by: Izzidien, Ahmed, et al.
Published: (2024)
by: Izzidien, Ahmed, et al.
Published: (2024)
The Cambridge Law Corpus: A Dataset for Legal AI Research
by: Östling, Andreas, et al.
Published: (2023)
by: Östling, Andreas, et al.
Published: (2023)
When Should Algorithms Resign? A Proposal for AI Governance
by: Bhatt, Umang, et al.
Published: (2024)
by: Bhatt, Umang, et al.
Published: (2024)
Formalising Anti-Discrimination Law in Automated Decision Systems
by: Sargeant, Holli, et al.
Published: (2024)
by: Sargeant, Holli, et al.
Published: (2024)
Automatic Information Extraction From Employment Tribunal Judgements Using Large Language Models
by: de Faria, Joana Ribeiro, et al.
Published: (2024)
by: de Faria, Joana Ribeiro, et al.
Published: (2024)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs
by: Huang, Yuxiang, et al.
Published: (2025)
by: Huang, Yuxiang, et al.
Published: (2025)
From Estimation to Discrimination: Algorithmic Bias, Predictive Uncertainty, and Anti‐Discrimination Law
by: Holli Sargeant
Published: (2026)
by: Holli Sargeant
Published: (2026)
The CLC-UKET Dataset: Benchmarking Case Outcome Prediction for the UK Employment Tribunal
by: Xie, Huiyuan, et al.
Published: (2024)
by: Xie, Huiyuan, et al.
Published: (2024)
Unequal Uncertainty: Rethinking Algorithmic Interventions for Mitigating Discrimination from AI
by: Sargeant, Holli, et al.
Published: (2025)
by: Sargeant, Holli, et al.
Published: (2025)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
Enhancing Legal Case Retrieval via Scaling High-quality Synthetic Query-Candidate Pairs
by: Gao, Cheng, et al.
Published: (2024)
by: Gao, Cheng, et al.
Published: (2024)
Exploring the Benefit of Activation Sparsity in Pre-training
by: Zhang, Zhengyan, et al.
Published: (2024)
by: Zhang, Zhengyan, et al.
Published: (2024)
Locret: Enhancing Eviction in Long-Context LLM Inference with Trained Retaining Heads on Consumer-Grade Devices
by: Huang, Yuxiang, et al.
Published: (2024)
by: Huang, Yuxiang, et al.
Published: (2024)
LexChain: Modeling Legal Reasoning Chains for Chinese Tort Case Analysis
by: Xie, Huiyuan, et al.
Published: (2025)
by: Xie, Huiyuan, et al.
Published: (2025)
Variator: Accelerating Pre-trained Models with Plug-and-Play Compression Modules
by: Xiao, Chaojun, et al.
Published: (2023)
by: Xiao, Chaojun, et al.
Published: (2023)
Densing Law of LLMs
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
by: Gao, Cheng, et al.
Published: (2026)
by: Gao, Cheng, et al.
Published: (2026)
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
by: Song, Chenyang, et al.
Published: (2025)
by: Song, Chenyang, et al.
Published: (2025)
ComplicitSplat: Downstream Models are Vulnerable to Blackbox Attacks by 3D Gaussian Splat Camouflages
by: Hull, Matthew, et al.
Published: (2025)
by: Hull, Matthew, et al.
Published: (2025)
Ouroboros: Generating Longer Drafts Phrase by Phrase for Faster Speculative Decoding
by: Zhao, Weilin, et al.
Published: (2024)
by: Zhao, Weilin, et al.
Published: (2024)
LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases
by: Cai, Yida, et al.
Published: (2025)
by: Cai, Yida, et al.
Published: (2025)
Legal$Δ$: Enhancing Legal Reasoning in LLMs via Reinforcement Learning with Chain-of-Thought Guided Information Gain
by: Dai, Xin, et al.
Published: (2025)
by: Dai, Xin, et al.
Published: (2025)
mCoT: Multilingual Instruction Tuning for Reasoning Consistency in Language Models
by: Lai, Huiyuan, et al.
Published: (2024)
by: Lai, Huiyuan, et al.
Published: (2024)
Robust and Scalable Model Editing for Large Language Models
by: Chen, Yingfa, et al.
Published: (2024)
by: Chen, Yingfa, et al.
Published: (2024)
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
by: Wang, Xiaorong, et al.
Published: (2025)
by: Wang, Xiaorong, et al.
Published: (2025)
LegalDuet: Learning Fine-grained Representations for Legal Judgment Prediction via a Dual-View Contrastive Learning
by: Xu, Buqiang, et al.
Published: (2024)
by: Xu, Buqiang, et al.
Published: (2024)
Exploring Format Consistency for Instruction Tuning
by: Liang, Shihao, et al.
Published: (2023)
by: Liang, Shihao, et al.
Published: (2023)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
by: Chen, Weize, et al.
Published: (2025)
by: Chen, Weize, et al.
Published: (2025)
Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
by: Selvaganapathy, Sanjeeevan, et al.
Published: (2025)
by: Selvaganapathy, Sanjeeevan, et al.
Published: (2025)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
by: He, Bingxiang, et al.
Published: (2025)
by: He, Bingxiang, et al.
Published: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
by: Luo, Yuqi, et al.
Published: (2024)
by: Luo, Yuqi, et al.
Published: (2024)
Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning
by: Yao, Rujing, et al.
Published: (2025)
by: Yao, Rujing, et al.
Published: (2025)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
by: Chen, Junhao, et al.
Published: (2024)
by: Chen, Junhao, et al.
Published: (2024)
CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
by: Grundmann, Paul, et al.
Published: (2025)
by: Grundmann, Paul, et al.
Published: (2025)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
by: Bai, Yuzhuo, et al.
Published: (2025)
by: Bai, Yuzhuo, et al.
Published: (2025)
LinguaGame: A Linguistically Grounded Game-Theoretic Paradigm for Multi-Agent Dialogue Generation
by: Ye, Yuxiao, et al.
Published: (2026)
by: Ye, Yuxiao, et al.
Published: (2026)
Similar Items
-
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
by: Sargeant, Holli, et al.
Published: (2024) -
LLM vs. Lawyers: Identifying a Subset of Summary Judgments in a Large UK Case Law Dataset
by: Izzidien, Ahmed, et al.
Published: (2024) -
The Cambridge Law Corpus: A Dataset for Legal AI Research
by: Östling, Andreas, et al.
Published: (2023) -
When Should Algorithms Resign? A Proposal for AI Governance
by: Bhatt, Umang, et al.
Published: (2024) -
Formalising Anti-Discrimination Law in Automated Decision Systems
by: Sargeant, Holli, et al.
Published: (2024)