LLM-Based Persuasion Enables Guardrail Override in Frontier LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Nogueira, Rodrigo, Almeida, Thales Sales, Bonás, Giovana Kerche, Roque, Andrea, Pires, Ramon, Abonizio, Hugo, Laitz, Thiago, Larcher, Celio, Junior, Roseval Malaquias, Piau, Marcos |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026)
by: Bonás, Giovana Kerche, et al.
Published: (2026)
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026)
by: Junior, Roseval Malaquias, et al.
Published: (2026)
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)
by: Laitz, Thiago, et al.
Published: (2026)
Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
by: Pires, Ramon, et al.
Published: (2026)
by: Pires, Ramon, et al.
Published: (2026)
Sabiá-3 Technical Report
by: Abonizio, Hugo, et al.
Published: (2024)
by: Abonizio, Hugo, et al.
Published: (2024)
Ticket-Bench: A Kickoff for Multilingual and Regionalized Agent Evaluation
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
TiEBe: Tracking Language Model Recall of Notable Worldwide Events Through Time
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
BRoverbs -- Measuring how much LLMs understand Portuguese proverbs
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
by: Santos, João Guilherme Alves, et al.
Published: (2025)
by: Santos, João Guilherme Alves, et al.
Published: (2025)
Sabiá-2: A New Generation of Portuguese Large Language Models
by: Almeida, Thales Sales, et al.
Published: (2024)
by: Almeida, Thales Sales, et al.
Published: (2024)
Automatic Legal Writing Evaluation of LLMs
by: Pires, Ramon, et al.
Published: (2025)
by: Pires, Ramon, et al.
Published: (2025)
PoETa v2: Toward More Robust Evaluation of Large Language Models in Portuguese
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
The interplay between domain specialization and model size
by: Junior, Roseval Malaquias, et al.
Published: (2025)
by: Junior, Roseval Malaquias, et al.
Published: (2025)
Juru: Legal Brazilian Large Language Model from Reputable Sources
by: Junior, Roseval Malaquias, et al.
Published: (2024)
by: Junior, Roseval Malaquias, et al.
Published: (2024)
Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime
by: Abonizio, Hugo, et al.
Published: (2025)
by: Abonizio, Hugo, et al.
Published: (2025)
SurveySum: A Dataset for Summarizing Multiple Scientific Articles into a Survey Section
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
by: Fernandes, Leandro Carísio, et al.
Published: (2024)
The Chronicles of RAG: The Retriever, the Chunk and the Generator
by: Finardi, Paulo, et al.
Published: (2024)
by: Finardi, Paulo, et al.
Published: (2024)
ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language
by: Piau, Marcos, et al.
Published: (2024)
by: Piau, Marcos, et al.
Published: (2024)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
by: Agarwal, Mahak, et al.
Published: (2025)
by: Agarwal, Mahak, et al.
Published: (2025)
Synthetic Rewriting as a Quality Multiplier: Evidence from Portuguese Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2026)
by: Almeida, Thales Sales, et al.
Published: (2026)
Building High-Quality Datasets for Portuguese LLMs: From Common Crawl Snapshots to Industrial-Grade Corpora
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
Curió-Edu 7B: Examining Data Selection Impacts in LLM Continued Pretraining
by: Almeida, Thales Sales, et al.
Published: (2025)
by: Almeida, Thales Sales, et al.
Published: (2025)
InRanker: Distilled Rankers for Zero-shot Information Retrieval
by: Laitz, Thiago, et al.
Published: (2024)
by: Laitz, Thiago, et al.
Published: (2024)
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs
by: Ferraretto, Fernando, et al.
Published: (2024)
by: Ferraretto, Fernando, et al.
Published: (2024)
Explainable LightGBM Approach for Predicting Myocardial Infarction Mortality
by: Vicente, Ana Letícia Garcez, et al.
Published: (2024)
by: Vicente, Ana Letícia Garcez, et al.
Published: (2024)
Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines
by: Abonizio, Hugo, et al.
Published: (2026)
by: Abonizio, Hugo, et al.
Published: (2026)
Guardrail Selection in Line Charts to Contextualize Persuasive Visualizations
by: Nadib, Khandaker Abrar, et al.
Published: (2026)
by: Nadib, Khandaker Abrar, et al.
Published: (2026)
HIV/AIDS Stigma and Discrimination among Nurses in Suriname
by: Winston Roseval
Published: (2007)
by: Winston Roseval
Published: (2007)
Measuring Cross-lingual Transfer in Bytes
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
by: de Souza, Leandro Rodrigues, et al.
Published: (2024)
CONFLITOS À MESA: Vegetarianos, consumo e identidade
by: Juliana Abonizio
Published: (2016)
by: Juliana Abonizio
Published: (2016)
Por uma quiromancia da vida urbana
by: Juliana Abonizio
Published: (2011)
by: Juliana Abonizio
Published: (2011)
Consumo alimentar e anticonsumismo: veganos e freeganos
by: Juliana Abonizio
Published: (2013)
by: Juliana Abonizio
Published: (2013)
INDEPENDÊNCIA, PODER JUDICIÁRIO E MINISTÉRIO PÚBLICO
by: Fábio Kerche
Published: (2018)
by: Fábio Kerche
Published: (2018)
MINISTÉRIO PÚBLICO, LAVA JATO E MÃOS LIMPAS: UMA ABORDAGEM INSTITUCIONAL
by: Fábio Kerche
Published: (2018)
by: Fábio Kerche
Published: (2018)
Os Conselhos Nacionais de Justiça e do Ministério Público no Brasil: instrumentos de accountability?
by: Fábio Kerche
Published: (2020)
by: Fábio Kerche
Published: (2020)
Autonomia e Discricionariedade do Ministério Público no Brasil
by: Fábio Kerche
Published: (2007)
by: Fábio Kerche
Published: (2007)
DE SEM-TERRA A SEM-TERRA: MEMÓRIAS E IDENTIDADES
by: Natália Kerche Alvaides
Published: (2013)
by: Natália Kerche Alvaides
Published: (2013)
Desenvolvimento, gestão e cooperação internacional: um estudo do projeto de desenvolvimento comunitário da bacia do Rio Gavião no sudoeste da Bahia
by: Weslei Gusmão Piau Santana
Published: (2013)
by: Weslei Gusmão Piau Santana
Published: (2013)
Similar Items
-
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026) -
CAPITU: A Benchmark for Evaluating Instruction-Following in Brazilian Portuguese with Literary Context
by: Bonás, Giovana Kerche, et al.
Published: (2026) -
Prosa: Rubric-Based Evaluation of LLMs on Real User Chats in Brazilian Portuguese
by: Junior, Roseval Malaquias, et al.
Published: (2026) -
MARCA: A Checklist-Based Benchmark for Multilingual Web Search
by: Almeida, Thales Sales, et al.
Published: (2026) -
Sabiá-4 Technical Report
by: Laitz, Thiago, et al.
Published: (2026)