Curiosity-driven Red-teaming for Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Zhang-Wei, Shenfeld, Idan, Wang, Tsun-Hsuan, Chuang, Yung-Sung, Pareja, Aldo, Glass, James, Srivastava, Akash, Agrawal, Pulkit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Value Augmented Sampling for Language Model Alignment and Personalization
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
Language Model Personalization via Reward Factorization
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2023)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2023)
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2023)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2023)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
Self-Distillation Enables Continual Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2026)
JUICER: Data-Efficient Imitation Learning for Robotic Assembly
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
Don't "Overthink" Passage Reranking: Is Reasoning Truly Necessary?
von: Jedidi, Nour, et al.
Veröffentlicht: (2025)
von: Jedidi, Nour, et al.
Veröffentlicht: (2025)
Zero-Shot Dense Retrieval with Embeddings from Relevance Feedback
von: Jedidi, Nour, et al.
Veröffentlicht: (2024)
von: Jedidi, Nour, et al.
Veröffentlicht: (2024)
LAB: Large-Scale Alignment for ChatBots
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
von: Sudalairaj, Shivchander, et al.
Veröffentlicht: (2024)
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
von: Bahlous-Boldi, Ryan, et al.
Veröffentlicht: (2026)
von: Bahlous-Boldi, Ryan, et al.
Veröffentlicht: (2026)
From Imitation to Refinement -- Residual RL for Precise Assembly
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
von: Ankile, Lars, et al.
Veröffentlicht: (2024)
SelfCite: Self-Supervised Alignment for Context Attribution in Large Language Models
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2025)
Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models
von: Puri, Isha, et al.
Veröffentlicht: (2026)
von: Puri, Isha, et al.
Veröffentlicht: (2026)
Natural Language Embedded Programs for Hybrid Language Symbolic Reasoning
von: Zhang, Tianhua, et al.
Veröffentlicht: (2023)
von: Zhang, Tianhua, et al.
Veröffentlicht: (2023)
Aligning Language Models from User Interactions
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2026)
von: Buening, Thomas Kleine, et al.
Veröffentlicht: (2026)
Embodied Red Teaming for Auditing Robotic Foundation Models
von: Karnik, Sathwik, et al.
Veröffentlicht: (2024)
von: Karnik, Sathwik, et al.
Veröffentlicht: (2024)
Cleansing Jewel: A Neural Spelling Correction Model Built On Google OCR-ed Tibetan Manuscripts
von: Luo, Queenie, et al.
Veröffentlicht: (2023)
von: Luo, Queenie, et al.
Veröffentlicht: (2023)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
CALM: Curiosity-Driven Auditing for Large Language Models
von: Zheng, Xiang, et al.
Veröffentlicht: (2025)
von: Zheng, Xiang, et al.
Veröffentlicht: (2025)
Why Did Apple Fall: Evaluating Curiosity in Large Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Training Language Models via Neural Cellular Automata
von: Lee, Dan, et al.
Veröffentlicht: (2026)
von: Lee, Dan, et al.
Veröffentlicht: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
Learning How Hard to Think: Input-Adaptive Allocation of LM Computation
von: Damani, Mehul, et al.
Veröffentlicht: (2024)
von: Damani, Mehul, et al.
Veröffentlicht: (2024)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
von: Arif, Samee, et al.
Veröffentlicht: (2026)
von: Arif, Samee, et al.
Veröffentlicht: (2026)
HAMMER: Hamiltonian Curiosity Augmented Large Language Model Reinforcement
von: Yang, Ming, et al.
Veröffentlicht: (2025)
von: Yang, Ming, et al.
Veröffentlicht: (2025)
Learning to Attribute with Attention
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2025)
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2025)
Activation-Informed Merging of Large Language Models
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2025)
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2025)
Judging It, Washing It: Scoring and Greenwashing Corporate Climate Disclosures using Large Language Models
von: Chuang, Marianne, et al.
Veröffentlicht: (2025)
von: Chuang, Marianne, et al.
Veröffentlicht: (2025)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Quantifying Generalization Complexity for Large Language Models
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
von: Qi, Zhenting, et al.
Veröffentlicht: (2024)
Towards Audio Token Compression in Large Audio Language Models
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
von: Bhati, Saurabhchand, et al.
Veröffentlicht: (2025)
Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning
von: Ho, Duc Hieu, et al.
Veröffentlicht: (2025)
von: Ho, Duc Hieu, et al.
Veröffentlicht: (2025)
Beyond Single-Shot: Multi-step Tool Retrieval via Query Planning
von: Fang, Wei, et al.
Veröffentlicht: (2026)
von: Fang, Wei, et al.
Veröffentlicht: (2026)
Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games
von: Ma, Chengdong, et al.
Veröffentlicht: (2023)
von: Ma, Chengdong, et al.
Veröffentlicht: (2023)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
von: Dou, Yipu, et al.
Veröffentlicht: (2026)
Grounding Language Plans in Demonstrations Through Counterfactual Perturbations
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
von: Wang, Yanwei, et al.
Veröffentlicht: (2024)
Enhancing Large Language Models with Neurosymbolic Reasoning for Multilingual Tasks
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2025)
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2025)
Do Large Language Models know who did what to whom?
von: Denning, Joseph M., et al.
Veröffentlicht: (2025)
von: Denning, Joseph M., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Value Augmented Sampling for Language Model Alignment and Personalization
von: Han, Seungwook, et al.
Veröffentlicht: (2024) -
Language Model Personalization via Reward Factorization
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025) -
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025) -
TGRL: An Algorithm for Teacher Guided Reinforcement Learning
von: Shenfeld, Idan, et al.
Veröffentlicht: (2023) -
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2023)