Saved in:
| Main Authors: | Rivera, Mauricio, Godbout, Jean-François, Rabbany, Reihaneh, Pelrine, Kellin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2401.08694 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty Resolution in Misinformation Detection
by: Orlovskiy, Yury, et al.
Published: (2024)
by: Orlovskiy, Yury, et al.
Published: (2024)
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024)
by: Vergho, Tyler, et al.
Published: (2024)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
Web Retrieval Agents for Evidence-Based Misinformation Detection
by: Tian, Jacob-Junqi, et al.
Published: (2024)
by: Tian, Jacob-Junqi, et al.
Published: (2024)
Epistemic Integrity in Large Language Models
by: Ghafouri, Bijean, et al.
Published: (2024)
by: Ghafouri, Bijean, et al.
Published: (2024)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
by: Yang, Zachary, et al.
Published: (2023)
by: Yang, Zachary, et al.
Published: (2023)
Veracity: An Open-Source AI Fact-Checking System
by: Curtis, Taylor Lynn, et al.
Published: (2025)
by: Curtis, Taylor Lynn, et al.
Published: (2025)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
by: Gibbs, Tom, et al.
Published: (2024)
by: Gibbs, Tom, et al.
Published: (2024)
A Guide to Misinformation Detection Data and Evaluation
by: Thibault, Camille, et al.
Published: (2024)
by: Thibault, Camille, et al.
Published: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
From Intuition to Understanding: Using AI Peers to Overcome Physics Misconceptions
by: Weijers, Ruben, et al.
Published: (2025)
by: Weijers, Ruben, et al.
Published: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
The Structural Safety Generalization Problem
by: Broomfield, Julius, et al.
Published: (2025)
by: Broomfield, Julius, et al.
Published: (2025)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
by: Yang, Zachary, et al.
Published: (2025)
by: Yang, Zachary, et al.
Published: (2025)
Online Influence Campaigns: Strategies and Vulnerabilities
by: Musulan, Andreea, et al.
Published: (2024)
by: Musulan, Andreea, et al.
Published: (2024)
What do people want to fact-check?
by: Ghafouri, Bijean, et al.
Published: (2026)
by: Ghafouri, Bijean, et al.
Published: (2026)
A Simulation System Towards Solving Societal-Scale Manipulation
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
The $\textit{Silicon Society}$ Cookbook: Design Space of LLM-based Social Simulations
by: Bück-Kaeffer, Aurélien, et al.
Published: (2026)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2026)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training
by: Mohammadzadeh, Shahrad, et al.
Published: (2024)
by: Mohammadzadeh, Shahrad, et al.
Published: (2024)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Eliciting Uncertainty in Chain-of-Thought to Mitigate Bias against Forecasting Harmful User Behaviors
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
Semantic Density: Uncertainty Quantification for Large Language Models through Confidence Measurement in Semantic Space
by: Qiu, Xin, et al.
Published: (2024)
by: Qiu, Xin, et al.
Published: (2024)
PairBench: Are Vision-Language Models Reliable at Comparing What They See?
by: Feizi, Aarash, et al.
Published: (2025)
by: Feizi, Aarash, et al.
Published: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026)
by: Seleznyov, Mikhail, et al.
Published: (2026)
CrediBench: Building Web-Scale Network Datasets for Information Integrity
by: Kondrup, Emma, et al.
Published: (2025)
by: Kondrup, Emma, et al.
Published: (2025)
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
by: Kowal, Matthew, et al.
Published: (2025)
by: Kowal, Matthew, et al.
Published: (2025)
Evaluating Uncertainty Quantification Methods in Argumentative Large Language Models
by: Zhou, Kevin, et al.
Published: (2025)
by: Zhou, Kevin, et al.
Published: (2025)
Large language models can effectively convince people to believe conspiracies
by: Costello, Thomas H., et al.
Published: (2026)
by: Costello, Thomas H., et al.
Published: (2026)
Agentic Uncertainty Quantification
by: Zhang, Jiaxin, et al.
Published: (2026)
by: Zhang, Jiaxin, et al.
Published: (2026)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
by: Jing, Xiaonan, et al.
Published: (2024)
by: Jing, Xiaonan, et al.
Published: (2024)
OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
by: Livernoche, Victor, et al.
Published: (2025)
by: Livernoche, Victor, et al.
Published: (2025)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
by: Naderi, Nariman, et al.
Published: (2025)
by: Naderi, Nariman, et al.
Published: (2025)
MAQA: Evaluating Uncertainty Quantification in LLMs Regarding Data Uncertainty
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Low-Confidence Gold: Refining Low-Confidence Samples for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
Regional and Temporal Patterns of Partisan Polarization during the COVID-19 Pandemic in the United States and Canada
by: Yang, Zachary, et al.
Published: (2024)
by: Yang, Zachary, et al.
Published: (2024)
CSS: Contrastive Semantic Similarity for Uncertainty Quantification of LLMs
by: Ao, Shuang, et al.
Published: (2024)
by: Ao, Shuang, et al.
Published: (2024)
Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation
by: Lin, Qinhong, et al.
Published: (2024)
by: Lin, Qinhong, et al.
Published: (2024)
Comparing Uncertainty Measurement and Mitigation Methods for Large Language Models: A Systematic Review
by: Abbasli, Toghrul, et al.
Published: (2025)
by: Abbasli, Toghrul, et al.
Published: (2025)
Similar Items
-
Uncertainty Resolution in Misinformation Detection
by: Orlovskiy, Yury, et al.
Published: (2024) -
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024) -
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025) -
Web Retrieval Agents for Evidence-Based Misinformation Detection
by: Tian, Jacob-Junqi, et al.
Published: (2024) -
Epistemic Integrity in Large Language Models
by: Ghafouri, Bijean, et al.
Published: (2024)