Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
Fuente:
arXiv
Salvato in:
| Autori principali: | Koo, Hamin, Kim, Minseon, Kim, Jaehyung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
di: Koo, Hamin, et al.
Pubblicazione: (2025)
di: Koo, Hamin, et al.
Pubblicazione: (2025)
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
di: Kim, Minseon, et al.
Pubblicazione: (2024)
di: Kim, Minseon, et al.
Pubblicazione: (2024)
Optimizing Query Generation for Enhanced Document Retrieval in RAG
di: Koo, Hamin, et al.
Pubblicazione: (2024)
di: Koo, Hamin, et al.
Pubblicazione: (2024)
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
di: Kim, Seoyeon, et al.
Pubblicazione: (2026)
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Structural Reasoning Improves Molecular Understanding of LLM
di: Jang, Yunhui, et al.
Pubblicazione: (2024)
di: Jang, Yunhui, et al.
Pubblicazione: (2024)
Revisiting the Uniform Information Density Hypothesis in LLM Reasoning
di: Gwak, Minju, et al.
Pubblicazione: (2025)
di: Gwak, Minju, et al.
Pubblicazione: (2025)
Training-free LLM Verification via Recycling Few-shot Examples
di: Lee, Dongseok, et al.
Pubblicazione: (2025)
di: Lee, Dongseok, et al.
Pubblicazione: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
di: Kim, Dongyoung, et al.
Pubblicazione: (2024)
Efficient LLM Collaboration via Planning
di: Lee, Byeongchan, et al.
Pubblicazione: (2025)
di: Lee, Byeongchan, et al.
Pubblicazione: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
di: Kim, Minseon, et al.
Pubblicazione: (2025)
di: Kim, Minseon, et al.
Pubblicazione: (2025)
Personalized LLM Decoding via Contrasting Personal Preference
di: Bu, Hyungjune, et al.
Pubblicazione: (2025)
di: Bu, Hyungjune, et al.
Pubblicazione: (2025)
Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection
di: Kim, Hojin, et al.
Pubblicazione: (2026)
di: Kim, Hojin, et al.
Pubblicazione: (2026)
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
di: Li, Miles Q., et al.
Pubblicazione: (2026)
di: Li, Miles Q., et al.
Pubblicazione: (2026)
SelectLLM: Can LLMs Select Important Instructions to Annotate?
di: Parkar, Ritik Sachin, et al.
Pubblicazione: (2024)
di: Parkar, Ritik Sachin, et al.
Pubblicazione: (2024)
Efficient LLM-Jailbreaking via Multimodal-LLM Jailbreak
di: Ji, Haoxuan, et al.
Pubblicazione: (2024)
di: Ji, Haoxuan, et al.
Pubblicazione: (2024)
InterPol: De-anonymizing LM Arena via Interpolated Preference Learning
di: Cho, Minsung, et al.
Pubblicazione: (2026)
di: Cho, Minsung, et al.
Pubblicazione: (2026)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
di: Kim, Dongyoung, et al.
Pubblicazione: (2026)
Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
di: Choi, Junhyuk, et al.
Pubblicazione: (2026)
TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRA
di: Jung, Chanjoo, et al.
Pubblicazione: (2025)
di: Jung, Chanjoo, et al.
Pubblicazione: (2025)
Enhancing Instruction Following of LLMs via Activation Steering with Dynamic Rejection
di: Kang, Minjae, et al.
Pubblicazione: (2026)
di: Kang, Minjae, et al.
Pubblicazione: (2026)
MEEA: Mere Exposure Effect-Driven Confrontational Optimization for LLM Jailbreaking
di: Zhang, Jianyi, et al.
Pubblicazione: (2025)
di: Zhang, Jianyi, et al.
Pubblicazione: (2025)
Jailbreaking LLM-Controlled Robots
di: Robey, Alexander, et al.
Pubblicazione: (2024)
di: Robey, Alexander, et al.
Pubblicazione: (2024)
Protein Representation Learning by Capturing Protein Sequence-Structure-Function Relationship
di: Ko, Eunji, et al.
Pubblicazione: (2024)
di: Ko, Eunji, et al.
Pubblicazione: (2024)
Bi-Level Prompt Optimization for Multimodal LLM-as-a-Judge
di: Pan, Bo, et al.
Pubblicazione: (2026)
di: Pan, Bo, et al.
Pubblicazione: (2026)
LLM-Enhanced Black-Litterman Portfolio Optimization
di: Lee, Youngbin, et al.
Pubblicazione: (2025)
di: Lee, Youngbin, et al.
Pubblicazione: (2025)
Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
di: Wu, Tianhao, et al.
Pubblicazione: (2024)
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
di: Zhu, Wentian, et al.
Pubblicazione: (2025)
di: Zhu, Wentian, et al.
Pubblicazione: (2025)
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments
di: Li, Yuran, et al.
Pubblicazione: (2025)
di: Li, Yuran, et al.
Pubblicazione: (2025)
LLM Misalignment via Adversarial RLHF Platforms
di: Entezami, Erfan, et al.
Pubblicazione: (2025)
di: Entezami, Erfan, et al.
Pubblicazione: (2025)
Gap-K%: Measuring Top-1 Prediction Gap for Detecting Pretraining Data
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
di: Kwak, Minseo, et al.
Pubblicazione: (2026)
Few-shot Personalization of LLMs with Mis-aligned Responses
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
LLM-as-Judge for Semantic Judging of Powerline Segmentation in UAV Inspection
di: Hossain, Akram, et al.
Pubblicazione: (2026)
di: Hossain, Akram, et al.
Pubblicazione: (2026)
Judge Reliability Harness: Stress Testing the Reliability of LLM Judges
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
di: Dev, Sunishchal, et al.
Pubblicazione: (2026)
Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
di: Barnes, Jeremy, et al.
Pubblicazione: (2025)
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
di: Pan, Tianjun, et al.
Pubblicazione: (2026)
M-Prometheus: A Suite of Open Multilingual LLM Judges
di: Pombal, José, et al.
Pubblicazione: (2025)
di: Pombal, José, et al.
Pubblicazione: (2025)
Learning to Correct for QA Reasoning with Black-box LLMs
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
di: Kim, Jaehyung, et al.
Pubblicazione: (2024)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
di: Li, Xiaochuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
di: Koo, Hamin, et al.
Pubblicazione: (2025) -
Automatic Jailbreaking of the Text-to-Image Generative AI Systems
di: Kim, Minseon, et al.
Pubblicazione: (2024) -
Optimizing Query Generation for Enhanced Document Retrieval in RAG
di: Koo, Hamin, et al.
Pubblicazione: (2024) -
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
di: Kim, Seoyeon, et al.
Pubblicazione: (2026) -
Revisiting the UID Hypothesis in LLM Reasoning Traces
di: Gwak, Minju, et al.
Pubblicazione: (2025)