Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Rui, Wang, Peiyi, Ma, Jingyuan, Zhang, Di, Sha, Lei, Sui, Zhifang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Plug-and-Play Training Framework for Preference Optimization
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
Towards Harmonized Uncertainty Estimation for Large Language Models
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template Decomposition
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
Chain-of-Thought Tokens are Computer Program Variables
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
di: Zhu, Fangwei, et al.
Pubblicazione: (2025)
HauntAttack: When Attack Follows Reasoning as a Shadow
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)
Not All Demonstration Examples are Equally Beneficial: Reweighting Demonstration Examples for In-Context Learning
di: Yang, Zhe, et al.
Pubblicazione: (2023)
di: Yang, Zhe, et al.
Pubblicazione: (2023)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
Large Language Models Struggle with Unreasonability in Math Problems
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
di: Ma, Jingyuan, et al.
Pubblicazione: (2024)
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning
di: Li, Zheng, et al.
Pubblicazione: (2025)
di: Li, Zheng, et al.
Pubblicazione: (2025)
AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming
di: Diao, Muxi, et al.
Pubblicazione: (2025)
di: Diao, Muxi, et al.
Pubblicazione: (2025)
How Far are LLMs from Being Our Digital Twins? A Benchmark for Persona-Based Behavior Chain Simulation
di: Li, Rui, et al.
Pubblicazione: (2025)
di: Li, Rui, et al.
Pubblicazione: (2025)
CoLT: Reasoning with Chain of Latent Tool Calls
di: Zhu, Fangwei, et al.
Pubblicazione: (2026)
di: Zhu, Fangwei, et al.
Pubblicazione: (2026)
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding
di: Xia, Heming, et al.
Pubblicazione: (2024)
di: Xia, Heming, et al.
Pubblicazione: (2024)
Harnessing the Plug-and-Play Controller by Prompting
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
Detoxification for LLM: From Dataset Itself
di: Shao, Wei, et al.
Pubblicazione: (2026)
di: Shao, Wei, et al.
Pubblicazione: (2026)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
di: Freenor, Michael, et al.
Pubblicazione: (2025)
di: Freenor, Michael, et al.
Pubblicazione: (2025)
FSM: A Finite State Machine Based Zero-Shot Prompting Paradigm for Multi-Hop Question Answering
di: Wang, Xiaochen, et al.
Pubblicazione: (2024)
di: Wang, Xiaochen, et al.
Pubblicazione: (2024)
Towards Better RL Training Data Utilization via Second-Order Rollout
di: Yang, Zhe, et al.
Pubblicazione: (2026)
di: Yang, Zhe, et al.
Pubblicazione: (2026)
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
di: Wang, Peiyi, et al.
Pubblicazione: (2023)
Exploring Activation Patterns of Parameters in Language Models
di: Wang, Yudong, et al.
Pubblicazione: (2024)
di: Wang, Yudong, et al.
Pubblicazione: (2024)
SelfCP: Compressing Over-Limit Prompt via the Frozen Large Language Model Itself
di: Gao, Jun, et al.
Pubblicazione: (2024)
di: Gao, Jun, et al.
Pubblicazione: (2024)
Language Models Encode the Value of Numbers Linearly
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
Enhancing Reliability across Short and Long-Form QA via Reinforcement Learning
di: Wang, Yudong, et al.
Pubblicazione: (2025)
di: Wang, Yudong, et al.
Pubblicazione: (2025)
A Survey on In-context Learning
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
di: Dong, Qingxiu, et al.
Pubblicazione: (2022)
Red Teaming Visual Language Models
di: Li, Mukai, et al.
Pubblicazione: (2024)
di: Li, Mukai, et al.
Pubblicazione: (2024)
PeriodicLoRA: Breaking the Low-Rank Bottleneck in LoRA Optimization
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
di: Meng, Xiangdi, et al.
Pubblicazione: (2024)
FERRET: Framework for Expansion Reliant Red Teaming
di: Mehrabi, Ninareh, et al.
Pubblicazione: (2026)
di: Mehrabi, Ninareh, et al.
Pubblicazione: (2026)
Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
di: Chin, Zhi-Yi, et al.
Pubblicazione: (2023)
di: Chin, Zhi-Yi, et al.
Pubblicazione: (2023)
Red Teaming Multimodal Language Models: Evaluating Harm Across Prompt Modalities and Models
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
di: Van Doren, Madison, et al.
Pubblicazione: (2025)
Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents
di: Mao, Yanxu, et al.
Pubblicazione: (2026)
di: Mao, Yanxu, et al.
Pubblicazione: (2026)
Against The Achilles' Heel: A Survey on Red Teaming for Generative Models
di: Lin, Lizhi, et al.
Pubblicazione: (2024)
di: Lin, Lizhi, et al.
Pubblicazione: (2024)
Resource Consumption Red-Teaming for Large Vision-Language Models
di: Gao, Haoran, et al.
Pubblicazione: (2025)
di: Gao, Haoran, et al.
Pubblicazione: (2025)
Self-Guard: Empower the LLM to Safeguard Itself
di: Wang, Zezhong, et al.
Pubblicazione: (2023)
di: Wang, Zezhong, et al.
Pubblicazione: (2023)
Prompting ChatGPT for Translation: A Comparative Analysis of Translation Brief and Persona Prompts
di: He, Sui
Pubblicazione: (2024)
di: He, Sui
Pubblicazione: (2024)
SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine
di: Wang, Xiaochen, et al.
Pubblicazione: (2024)
di: Wang, Xiaochen, et al.
Pubblicazione: (2024)
A Probabilistic Inference Scaling Theory for LLM Self-Correction
di: Yang, Zhe, et al.
Pubblicazione: (2025)
di: Yang, Zhe, et al.
Pubblicazione: (2025)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
di: Yang, Zhe, et al.
Pubblicazione: (2024)
di: Yang, Zhe, et al.
Pubblicazione: (2024)
HistLens: Mapping Idea Change across Concepts and Corpora
di: Jing, Yi, et al.
Pubblicazione: (2026)
di: Jing, Yi, et al.
Pubblicazione: (2026)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
di: Zhan, Weidong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Plug-and-Play Training Framework for Preference Optimization
di: Ma, Jingyuan, et al.
Pubblicazione: (2024) -
Towards Harmonized Uncertainty Estimation for Large Language Models
di: Li, Rui, et al.
Pubblicazione: (2025) -
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template Decomposition
di: Zhu, Fangwei, et al.
Pubblicazione: (2024) -
Chain-of-Thought Tokens are Computer Program Variables
di: Zhu, Fangwei, et al.
Pubblicazione: (2025) -
HauntAttack: When Attack Follows Reasoning as a Shadow
di: Ma, Jingyuan, et al.
Pubblicazione: (2025)