Epistemic Integrity in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ghafouri, Bijean, Mohammadzadeh, Shahrad, Zhou, James, Nair, Pratheeksha, Tian, Jacob-Junqi, Tsujimura, Hikaru, Goel, Mayank, Krishna, Sukanya, Rabbany, Reihaneh, Godbout, Jean-François, Pelrine, Kellin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024)
by: Vergho, Tyler, et al.
Published: (2024)
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
by: Rivera, Mauricio, et al.
Published: (2024)
by: Rivera, Mauricio, et al.
Published: (2024)
Uncertainty Resolution in Misinformation Detection
by: Orlovskiy, Yury, et al.
Published: (2024)
by: Orlovskiy, Yury, et al.
Published: (2024)
What do people want to fact-check?
by: Ghafouri, Bijean, et al.
Published: (2026)
by: Ghafouri, Bijean, et al.
Published: (2026)
Weak Supervision for Real World Graphs
by: Nair, Pratheeksha, et al.
Published: (2025)
by: Nair, Pratheeksha, et al.
Published: (2025)
A Guide to Misinformation Detection Data and Evaluation
by: Thibault, Camille, et al.
Published: (2024)
by: Thibault, Camille, et al.
Published: (2024)
CrediBench: Building Web-Scale Network Datasets for Information Integrity
by: Kondrup, Emma, et al.
Published: (2025)
by: Kondrup, Emma, et al.
Published: (2025)
Web Retrieval Agents for Evidence-Based Misinformation Detection
by: Tian, Jacob-Junqi, et al.
Published: (2024)
by: Tian, Jacob-Junqi, et al.
Published: (2024)
Online Influence Campaigns: Strategies and Vulnerabilities
by: Musulan, Andreea, et al.
Published: (2024)
by: Musulan, Andreea, et al.
Published: (2024)
Veracity: An Open-Source AI Fact-Checking System
by: Curtis, Taylor Lynn, et al.
Published: (2025)
by: Curtis, Taylor Lynn, et al.
Published: (2025)
Ask before you Build: Rethinking AI-for-Good in Human Trafficking Interventions
by: Nair, Pratheeksha, et al.
Published: (2025)
by: Nair, Pratheeksha, et al.
Published: (2025)
$\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2025)
Hallucination Detox: Sensitivity Dropout (SenD) for Large Language Model Training
by: Mohammadzadeh, Shahrad, et al.
Published: (2024)
by: Mohammadzadeh, Shahrad, et al.
Published: (2024)
The Variance Paradox: How AI Reduces Diversity but Increases Novelty
by: Ghafouri, Bijean
Published: (2025)
by: Ghafouri, Bijean
Published: (2025)
Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility
by: Murphy, Brendan, et al.
Published: (2025)
by: Murphy, Brendan, et al.
Published: (2025)
From Intuition to Understanding: Using AI Peers to Overcome Physics Misconceptions
by: Weijers, Ruben, et al.
Published: (2025)
by: Weijers, Ruben, et al.
Published: (2025)
Lost Before Translation: Social Information Transmission and Survival in AI-AI Communication
by: Ghafouri, Bijean, et al.
Published: (2026)
by: Ghafouri, Bijean, et al.
Published: (2026)
Towards Detecting Contextual Real-Time Toxicity for In-Game Chat
by: Yang, Zachary, et al.
Published: (2023)
by: Yang, Zachary, et al.
Published: (2023)
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
by: Gibbs, Tom, et al.
Published: (2024)
by: Gibbs, Tom, et al.
Published: (2024)
The Structural Safety Generalization Problem
by: Broomfield, Julius, et al.
Published: (2025)
by: Broomfield, Julius, et al.
Published: (2025)
Regional and Temporal Patterns of Partisan Polarization during the COVID-19 Pandemic in the United States and Canada
by: Yang, Zachary, et al.
Published: (2024)
by: Yang, Zachary, et al.
Published: (2024)
Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks
by: Struppek, Lukas, et al.
Published: (2026)
by: Struppek, Lukas, et al.
Published: (2026)
RLHF May Not Reflect Genuine Preferences
by: Ghafouri, Bijean, et al.
Published: (2026)
by: Ghafouri, Bijean, et al.
Published: (2026)
A Simulation System Towards Solving Societal-Scale Manipulation
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
by: Touzel, Maximilian Puelma, et al.
Published: (2024)
Deepfakes in the 2025 Canadian Election: Prevalence, Partisanship, and Platform Dynamics
by: Livernoche, Victor, et al.
Published: (2025)
by: Livernoche, Victor, et al.
Published: (2025)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
by: Tsujimura, Hikaru, et al.
Published: (2025)
by: Tsujimura, Hikaru, et al.
Published: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
by: Livernoche, Victor, et al.
Published: (2025)
by: Livernoche, Victor, et al.
Published: (2025)
Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
by: Yang, Zachary, et al.
Published: (2025)
by: Yang, Zachary, et al.
Published: (2025)
Scavenging Hyena: Distilling Transformers into Long Convolution Models
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
by: Ralambomihanta, Tokiniaina Raharison, et al.
Published: (2024)
Exploiting Novel GPT-4 APIs
by: Pelrine, Kellin, et al.
Published: (2023)
by: Pelrine, Kellin, et al.
Published: (2023)
Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential Privacy
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
GPS-SSL: Guided Positive Sampling to Inject Prior Into Self-Supervised Learning
by: Feizi, Aarash, et al.
Published: (2024)
by: Feizi, Aarash, et al.
Published: (2024)
The $\textit{Silicon Society}$ Cookbook: Design Space of LLM-based Social Simulations
by: Bück-Kaeffer, Aurélien, et al.
Published: (2026)
by: Bück-Kaeffer, Aurélien, et al.
Published: (2026)
EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations
by: Sarangi, Sneheel, et al.
Published: (2026)
by: Sarangi, Sneheel, et al.
Published: (2026)
Large language models can effectively convince people to believe conspiracies
by: Costello, Thomas H., et al.
Published: (2026)
by: Costello, Thomas H., et al.
Published: (2026)
DropleX: Liquid sensing on tablet touchscreens
by: Zhang, Siqi, et al.
Published: (2025)
by: Zhang, Siqi, et al.
Published: (2025)
Beyond Prediction -- Structuring Epistemic Integrity in Artificial Reasoning Systems
by: Wright, Craig Steven
Published: (2025)
by: Wright, Craig Steven
Published: (2025)
It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics
by: Kowal, Matthew, et al.
Published: (2025)
by: Kowal, Matthew, et al.
Published: (2025)
PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a Smartwatch
by: Arakawa, Riku, et al.
Published: (2024)
by: Arakawa, Riku, et al.
Published: (2024)
Similar Items
-
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation
by: Vergho, Tyler, et al.
Published: (2024) -
Combining Confidence Elicitation and Sample-based Methods for Uncertainty Quantification in Misinformation Mitigation
by: Rivera, Mauricio, et al.
Published: (2024) -
Uncertainty Resolution in Misinformation Detection
by: Orlovskiy, Yury, et al.
Published: (2024) -
What do people want to fact-check?
by: Ghafouri, Bijean, et al.
Published: (2026) -
Weak Supervision for Real World Graphs
by: Nair, Pratheeksha, et al.
Published: (2025)