What About the Scene with the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency Via Adversarial Nudge
Fuente:
arXiv
Saved in:
| Main Authors: | Dutta, Arka, Dutta, Sujan, Magu, Rijul, Datta, Soumyajit, De Choudhury, Munmun, KhudaBukhsh, Ashiqur R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
by: Magu, Rijul, et al.
Published: (2025)
by: Magu, Rijul, et al.
Published: (2025)
Gender Representation and Bias in Indian Civil Service Mock Interviews
by: Banerjee, Somonnoy, et al.
Published: (2024)
by: Banerjee, Somonnoy, et al.
Published: (2024)
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
by: Dutta, Arka, et al.
Published: (2023)
by: Dutta, Arka, et al.
Published: (2023)
Investigating Vaccine Buyer's Remorse: Post-Vaccination Decision Regret in COVID-19 Social Media Using Politically Diverse Human Annotation
by: Stanley, Miles, et al.
Published: (2026)
by: Stanley, Miles, et al.
Published: (2026)
Infrastructure Ombudsman: Mining Future Failure Concerns from Structural Disaster Response
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
by: Weerasooriya, Tharindu Cyril, et al.
Published: (2023)
ARTICLE: Annotator Reliability Through In-Context Learning
by: Dutta, Sujan, et al.
Published: (2024)
by: Dutta, Sujan, et al.
Published: (2024)
Community Needs and Assets: A Computational Analysis of Community Conversations
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)
Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope Speech
by: Pofcher, Jonathan, et al.
Published: (2025)
by: Pofcher, Jonathan, et al.
Published: (2025)
Rater Cohesion and Quality from a Vicarious Perspective
by: Pandita, Deepak, et al.
Published: (2024)
by: Pandita, Deepak, et al.
Published: (2024)
When Neutral Summaries are not that Neutral: Quantifying Political Neutrality in LLM-Generated News Summaries
by: Vijay, Supriti, et al.
Published: (2024)
by: Vijay, Supriti, et al.
Published: (2024)
On the State of NLP Approaches to Modeling Depression in Social Media: A Post-COVID-19 Outlook
by: Bucur, Ana-Maria, et al.
Published: (2024)
by: Bucur, Ana-Maria, et al.
Published: (2024)
Datasets for Depression Modeling in Social Media: An Overview
by: Bucur, Ana-Maria, et al.
Published: (2025)
by: Bucur, Ana-Maria, et al.
Published: (2025)
Statistical inference for a multiscale stochastic model of enzyme kinetics via propagation of chaos
by: Ganguly, Arnab, et al.
Published: (2024)
by: Ganguly, Arnab, et al.
Published: (2024)
Asymptotic Analysis of the Total Quasi-Steady State Approximation for the Michaelis--Menten Enzyme Kinetic Reactions
by: Ganguly, Arnab, et al.
Published: (2025)
by: Ganguly, Arnab, et al.
Published: (2025)
Local times of self-intersection and sample path properties of Volterra Gaussian processes
by: Izyumtseva, Olga, et al.
Published: (2024)
by: Izyumtseva, Olga, et al.
Published: (2024)
Mixing time for an epidemic model on graphs with external sources of infection
by: KhudaBukhsh, Wasiur R., et al.
Published: (2025)
by: KhudaBukhsh, Wasiur R., et al.
Published: (2025)
Self-intersection local times for Volterra Gaussian processes in stochastic flows with interaction
by: Izyumtseva, Olga, et al.
Published: (2026)
by: Izyumtseva, Olga, et al.
Published: (2026)
Employing Social Media to Improve Mental Health Outcomes
by: De Choudhury, Munmun
Published: (2025)
by: De Choudhury, Munmun
Published: (2025)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Model-based clustering of time-dependent observations with common structural changes
by: Corradin, Riccardo, et al.
Published: (2024)
by: Corradin, Riccardo, et al.
Published: (2024)
Stochastic Analysis of Entanglement-assisted Quantum Communication Channels
by: Elsayed, Karim S., et al.
Published: (2024)
by: Elsayed, Karim S., et al.
Published: (2024)
Pairwise accelerated failure time regression models for infectious disease transmission in close-contact groups with external sources of infection
by: Sharker, Yushuf, et al.
Published: (2019)
by: Sharker, Yushuf, et al.
Published: (2019)
Pairwise Accelerated Failure Time Regression Models for Infectious Disease Transmission in Close‐Contact Groups With External Sources of Infection
by: Yushuf Sharker, et al.
Published: (2024)
by: Yushuf Sharker, et al.
Published: (2024)
Including frameworks of public health ethics in computational modelling of infectious disease interventions
by: Zarebski, Alexander E., et al.
Published: (2025)
by: Zarebski, Alexander E., et al.
Published: (2025)
Silent Signals, Loud Impact: LLMs for Word-Sense Disambiguation of Coded Dog Whistles
by: Kruk, Julia, et al.
Published: (2024)
by: Kruk, Julia, et al.
Published: (2024)
Understanding Online Discussion Across Difference: Insights from Gun Discourse on Reddit
by: Magu, Rijul, et al.
Published: (2024)
by: Magu, Rijul, et al.
Published: (2024)
Applying RLAIF for Code Generation with API-usage in Lightweight LLMs
by: Dutta, Sujan, et al.
Published: (2024)
by: Dutta, Sujan, et al.
Published: (2024)
Kernel-based estimators for functional causal effects
by: Raykov, Yordan P., et al.
Published: (2025)
by: Raykov, Yordan P., et al.
Published: (2025)
Understanding the Humans Behind Online Misinformation: An Observational Study Through the Lens of the COVID-19 Pandemic
by: Chandra, Mohit, et al.
Published: (2023)
by: Chandra, Mohit, et al.
Published: (2023)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
Benchmarking Federated Learning for Throughput Prediction in 5G Live Streaming Applications
by: Dutta, Yuvraj, et al.
Published: (2025)
by: Dutta, Yuvraj, et al.
Published: (2025)
Robust Multi-Modal Image Stitching for Improved Scene Understanding
by: Dutta, Aritra, et al.
Published: (2023)
by: Dutta, Aritra, et al.
Published: (2023)
The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health Support
by: Song, Inhwa, et al.
Published: (2024)
by: Song, Inhwa, et al.
Published: (2024)
JU-NLP at Touché: Covert Advertisement in Conversational AI-Generation and Detection Strategies
by: Dutta, Arka, et al.
Published: (2025)
by: Dutta, Arka, et al.
Published: (2025)
A Framework for Situating Innovations, Opportunities, and Challenges in Advancing Vertical Systems with Large AI Models
by: Verma, Gaurav, et al.
Published: (2025)
by: Verma, Gaurav, et al.
Published: (2025)
The Role of Partisan Culture in Mental Health Language Online
by: Pendse, Sachin R., et al.
Published: (2025)
by: Pendse, Sachin R., et al.
Published: (2025)
Solving Navier-Stokes Equations Using Data-free Physics-Informed Neural Networks With Hard Boundary Conditions
by: Pal, Ritik, et al.
Published: (2025)
by: Pal, Ritik, et al.
Published: (2025)
Linguistic Comparison of AI- and Human-Written Responses to Online Mental Health Queries
by: Saha, Koustuv, et al.
Published: (2025)
by: Saha, Koustuv, et al.
Published: (2025)
IAP: Invisible Adversarial Patch Attack through Perceptibility-Aware Localization and Perturbation Optimization
by: Dutta, Subrat Kishore, et al.
Published: (2025)
by: Dutta, Subrat Kishore, et al.
Published: (2025)
Similar Items
-
Navigating the Rabbit Hole: Emergent Biases in LLM-Generated Attack Narratives Targeting Mental Health Groups
by: Magu, Rijul, et al.
Published: (2025) -
Gender Representation and Bias in Indian Civil Service Mock Interviews
by: Banerjee, Somonnoy, et al.
Published: (2024) -
Down the Toxicity Rabbit Hole: A Novel Framework to Bias Audit Large Language Models
by: Dutta, Arka, et al.
Published: (2023) -
Investigating Vaccine Buyer's Remorse: Post-Vaccination Decision Regret in COVID-19 Social Media Using Politically Diverse Human Annotation
by: Stanley, Miles, et al.
Published: (2026) -
Infrastructure Ombudsman: Mining Future Failure Concerns from Structural Disaster Response
by: Chowdhury, Md Towhidul Absar, et al.
Published: (2024)