Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Buszydlik, Aleksander, Dobiczek, Karol, Okoń, Michał Teodor, Skublicki, Konrad, Lippmann, Philip, Yang, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
von: Verma, Apurv, et al.
Veröffentlicht: (2024)
Towards Red Teaming in Multimodal and Multilingual Translation
von: Ropers, Christophe, et al.
Veröffentlicht: (2024)
von: Ropers, Christophe, et al.
Veröffentlicht: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
Towards Human Understanding of Paraphrase Types in Large Language Models
von: Meier, Dominik, et al.
Veröffentlicht: (2024)
von: Meier, Dominik, et al.
Veröffentlicht: (2024)
Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
von: Benkirane, Kenza, et al.
Veröffentlicht: (2024)
HACK: Hallucinations Along Certainty and Knowledge Axes
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
Task Contamination: Language Models May Not Be Few-Shot Anymore
von: Li, Changmao, et al.
Veröffentlicht: (2023)
von: Li, Changmao, et al.
Veröffentlicht: (2023)
Fine-Tuned Large Language Models for Logical Translation: Reducing Hallucinations with Lang2Logic
von: Pan, Muyu, et al.
Veröffentlicht: (2025)
von: Pan, Muyu, et al.
Veröffentlicht: (2025)
An Industrial-Scale Insurance LLM Achieving Verifiable Domain Mastery and Hallucination Control without Competence Trade-offs
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
von: Zhu, Qian, et al.
Veröffentlicht: (2026)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
EMO-KNOW: A Large Scale Dataset on Emotion and Emotion-cause
von: Nguyen, Mia Huong, et al.
Veröffentlicht: (2024)
von: Nguyen, Mia Huong, et al.
Veröffentlicht: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task
von: Evelo, Bart, et al.
Veröffentlicht: (2026)
von: Evelo, Bart, et al.
Veröffentlicht: (2026)
Precise Length Control in Large Language Models
von: Butcher, Bradley, et al.
Veröffentlicht: (2024)
von: Butcher, Bradley, et al.
Veröffentlicht: (2024)
Large Language Models for Biomedical Article Classification
von: Proboszcz, Jakub, et al.
Veröffentlicht: (2026)
von: Proboszcz, Jakub, et al.
Veröffentlicht: (2026)
Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language Models
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
von: Tian, Changxin, et al.
Veröffentlicht: (2025)
Strategy Adaptation in Large Language Model Werewolf Agents
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
von: Nakamori, Fuya, et al.
Veröffentlicht: (2025)
Socially Responsible Data for Large Multilingual Language Models
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
von: Smart, Andrew, et al.
Veröffentlicht: (2024)
RedHerring Attack: Testing the Reliability of Attack Detection
von: Rusert, Jonathan
Veröffentlicht: (2025)
von: Rusert, Jonathan
Veröffentlicht: (2025)
"AGI" team at SHROOM-CAP: Data-Centric Approach to Multilingual Hallucination Detection using XLM-RoBERTa
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
von: Rathva, Harsh, et al.
Veröffentlicht: (2025)
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2024)
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2024)
Qomhra: A Bilingual Irish and English Large Language Model
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
von: McInerney, Joseph, et al.
Veröffentlicht: (2025)
Dialect Normalization using Large Language Models and Morphological Rules
von: Dimakis, Antonios, et al.
Veröffentlicht: (2025)
von: Dimakis, Antonios, et al.
Veröffentlicht: (2025)
RUQuant: Towards Refining Uniform Quantization for Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
von: Riemenschneider, Frederick, et al.
Veröffentlicht: (2024)
von: Riemenschneider, Frederick, et al.
Veröffentlicht: (2024)
Knowledge Graphs, Large Language Models, and Hallucinations: An NLP Perspective
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2024)
von: Lavrinovics, Ernests, et al.
Veröffentlicht: (2024)
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
von: Peláez-González, Carlos, et al.
Veröffentlicht: (2025)
von: Peláez-González, Carlos, et al.
Veröffentlicht: (2025)
Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis
von: Nagao, Moe, et al.
Veröffentlicht: (2026)
von: Nagao, Moe, et al.
Veröffentlicht: (2026)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
Aligning Large Language Models for Faithful Integrity Against Opposing Argument
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
von: Zhao, Yong, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge Retrieval
von: Niu, Mengjia, et al.
Veröffentlicht: (2024)
von: Niu, Mengjia, et al.
Veröffentlicht: (2024)
Named Entity Recognition for Address Extraction in Speech-to-Text Transcriptions Using Synthetic Data
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
Intent Classification for Bank Chatbots through LLM Fine-Tuning
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
von: Lajčinová, Bibiána, et al.
Veröffentlicht: (2024)
How Do Large Language Models Acquire Factual Knowledge During Pretraining?
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
von: Chang, Hoyeon, et al.
Veröffentlicht: (2024)
LoRS: Efficient Low-Rank Adaptation for Sparse Large Language Model
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
von: Verma, Apurv, et al.
Veröffentlicht: (2024) -
Towards Red Teaming in Multimodal and Multilingual Translation
von: Ropers, Christophe, et al.
Veröffentlicht: (2024) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025) -
Towards Human Understanding of Paraphrase Types in Large Language Models
von: Meier, Dominik, et al.
Veröffentlicht: (2024)