Guardado en:
| Autores principales: | Kumar, Priyanshu, Lau, Elaine, Vijayakumar, Saranya, Trinh, Tu, Team, Scale Red, Chang, Elaine, Robinson, Vaughn, Hendryx, Sean, Zhou, Shuyan, Fredrikson, Matt, Yue, Summer, Wang, Zifan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2410.13886 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Jailbreaking to Jailbreak
por: Kritz, Jeremy, et al.
Publicado: (2025)
por: Kritz, Jeremy, et al.
Publicado: (2025)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
Reliable Weak-to-Strong Monitoring of LLM Agents
por: Kale, Neil, et al.
Publicado: (2025)
por: Kale, Neil, et al.
Publicado: (2025)
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
por: Knight, Christina Q., et al.
Publicado: (2025)
por: Knight, Christina Q., et al.
Publicado: (2025)
A Recipe for Improved Certifiable Robustness
por: Hu, Kai, et al.
Publicado: (2023)
por: Hu, Kai, et al.
Publicado: (2023)
VeriSplit: Secure and Practical Offloading of Machine Learning Inferences across IoT Devices
por: Zhang, Han, et al.
Publicado: (2024)
por: Zhang, Han, et al.
Publicado: (2024)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
por: Whitehead, Spencer, et al.
Publicado: (2024)
por: Whitehead, Spencer, et al.
Publicado: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
por: Palumbo, Nils, et al.
Publicado: (2024)
por: Palumbo, Nils, et al.
Publicado: (2024)
LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses
por: Lin, Weiran, et al.
Publicado: (2024)
por: Lin, Weiran, et al.
Publicado: (2024)
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
por: Yuan, Youliang, et al.
Publicado: (2024)
por: Yuan, Youliang, et al.
Publicado: (2024)
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
por: Chen, Luoyu, et al.
Publicado: (2026)
por: Chen, Luoyu, et al.
Publicado: (2026)
Jailbroken Frontier Models Retain Their Capabilities
por: Zhu, Daniel, et al.
Publicado: (2026)
por: Zhu, Daniel, et al.
Publicado: (2026)
LipNeXt: Scaling up Lipschitz-based Certified Robustness to Billion-parameter Models
por: Hu, Kai, et al.
Publicado: (2026)
por: Hu, Kai, et al.
Publicado: (2026)
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
por: Zhang, Hugh, et al.
Publicado: (2024)
por: Zhang, Hugh, et al.
Publicado: (2024)
Who Trains the Trainer? Library Staff are OPAC Users, Too.
por: Coppola, Elaine
Publicado: (1983)
por: Coppola, Elaine
Publicado: (1983)
Revisiting the Superficial Alignment Hypothesis
por: Raghavendra, Mohit, et al.
Publicado: (2024)
por: Raghavendra, Mohit, et al.
Publicado: (2024)
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks)
por: Priyanshu, Aman, et al.
Publicado: (2024)
por: Priyanshu, Aman, et al.
Publicado: (2024)
ADAPTACIÓN AL ESPAÑOL DE LA ESCALA DE AMBIENTE INVALIDANTE INFANTIL
por: Scale Martín M. Puddington
Publicado: (2017)
por: Scale Martín M. Puddington
Publicado: (2017)
Progress over Points: Reframing LM Benchmarks Around Scientific Objectives
por: Jin, Alwin, et al.
Publicado: (2025)
por: Jin, Alwin, et al.
Publicado: (2025)
Herz P1 Smart Scale More Than Just Weight Tracking
por: Herz P1 Smart Scale
Publicado: (2026)
por: Herz P1 Smart Scale
Publicado: (2026)
Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards
por: Da, Jeff, et al.
Publicado: (2025)
por: Da, Jeff, et al.
Publicado: (2025)
Going 3D with Technology: An Overarching Approach for Language Teachers
por: Jason D. Hendryx
Publicado: (2016)
por: Jason D. Hendryx
Publicado: (2016)
STARS (Secondary Training for Alaskan Rural Students): Communications. Draft Copy.
por: Griffin, Elaine, et al.
Publicado: (1977)
por: Griffin, Elaine, et al.
Publicado: (1977)
ToolComp: A Multi-Tool Reasoning & Process Supervision Benchmark
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
A Baseline Analysis of Reward Models' Ability To Accurately Analyze Foundation Models Under Distribution Shift
por: LeVine, Will, et al.
Publicado: (2023)
por: LeVine, Will, et al.
Publicado: (2023)
Does Refusal Training in LLMs Generalize to the Past Tense?
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
por: Andriushchenko, Maksym, et al.
Publicado: (2024)
Planning In Natural Language Improves LLM Search For Code Generation
por: Wang, Evan, et al.
Publicado: (2024)
por: Wang, Evan, et al.
Publicado: (2024)
Sex Differences in the Impact of Obesity on Immunity
por: Saranya Vijayakumar, et al.
Publicado: (2026)
por: Saranya Vijayakumar, et al.
Publicado: (2026)
Declaración europea por una nueva cultura del agua
por: Red Euwater
Publicado: (2005)
por: Red Euwater
Publicado: (2005)
The Functional Imperative: Structural Integrity and the Fallacy of Ornamental Soul
por: Red, Pax
Publicado: (2026)
por: Red, Pax
Publicado: (2026)
Revista de Prensa
por: Red Ires
Publicado: (2009)
por: Red Ires
Publicado: (2009)
The discrete empirical interpolation method in class identification and data summarization
por: Emily P. Hendryx Lyons
Publicado: (2024)
por: Emily P. Hendryx Lyons
Publicado: (2024)
Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation
por: Heiding, Fred, et al.
Publicado: (2025)
por: Heiding, Fred, et al.
Publicado: (2025)
Patrimônio Afro-Brasileiro no Contexto da Educação Escolar Quilombola
por: Elaine Monteiro
Publicado: (2019)
por: Elaine Monteiro
Publicado: (2019)
Imprensa pedagógica e o fazer historiográfico: o caso da Revista do Ensino (1929 – 1930)
por: Elaine Rodrigues
Publicado: (2015)
por: Elaine Rodrigues
Publicado: (2015)
Hijos de migrantes mexicanos en las escuelas de Estados Unidos
por: Elaine Levine
Publicado: (2006)
por: Elaine Levine
Publicado: (2006)
Tempo e performance
por: Elaine Conte
Publicado: (2016)
por: Elaine Conte
Publicado: (2016)
Homenagem revisita obra pioneira de Luiz Beltrão
por: Elaine Javorski
Publicado: (2010)
por: Elaine Javorski
Publicado: (2010)
O impacto do investidor institucional no preço das ações
por: Elaine Borges
Publicado: (2019)
por: Elaine Borges
Publicado: (2019)
Ejemplares similares
-
Jailbreaking to Jailbreak
por: Kritz, Jeremy, et al.
Publicado: (2025) -
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025) -
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025) -
Reliable Weak-to-Strong Monitoring of LLM Agents
por: Kale, Neil, et al.
Publicado: (2025) -
FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
por: Knight, Christina Q., et al.
Publicado: (2025)