Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
Fuente:
arXiv
Saved in:
| Main Authors: | Joshi, Abhinav, Saha, Shaswati, Shukla, Divyaksh, Vema, Sriram, Jhamtani, Harsh, Gaur, Manas, Modi, Ashutosh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
by: Shukla, Divyaksh, et al.
Published: (2026)
by: Shukla, Divyaksh, et al.
Published: (2026)
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
by: Patel, Shubham, et al.
Published: (2024)
by: Patel, Shubham, et al.
Published: (2024)
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025)
by: Dwivedi, Ashutosh, et al.
Published: (2025)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Beyond Components: Singular Vector-Based Interpretability of Transformer Circuits
by: Ahmad, Areeb, et al.
Published: (2025)
by: Ahmad, Areeb, et al.
Published: (2025)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
Towards Measuring and Modeling "Culture" in LLMs: A Survey
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
by: Adilazuarda, Muhammad Farid, et al.
Published: (2024)
CoMuMDR: Code-mixed Multi-modal Multi-domain corpus for Discourse paRsing in conversations
by: Shukla, Divyaksh, et al.
Published: (2025)
by: Shukla, Divyaksh, et al.
Published: (2025)
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
COLD: Causal reasOning in cLosed Daily activities
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024)
by: Mohseni, Seyedreza, et al.
Published: (2024)
LoRMA: Low-Rank Multiplicative Adaptation for LLMs
by: Bihany, Harsh, et al.
Published: (2025)
by: Bihany, Harsh, et al.
Published: (2025)
Side Effects of Erasing Concepts from Diffusion Models
by: Saha, Shaswati, et al.
Published: (2025)
by: Saha, Shaswati, et al.
Published: (2025)
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
by: Rezkellah, Fatmazohra, et al.
Published: (2025)
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
by: Gyrard, Amelie, et al.
Published: (2024)
by: Gyrard, Amelie, et al.
Published: (2024)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
by: Gorti, Atmika, et al.
Published: (2024)
by: Gorti, Atmika, et al.
Published: (2024)
LM Agents for Coordinating Multi-User Information Gathering
by: Jhamtani, Harsh, et al.
Published: (2025)
by: Jhamtani, Harsh, et al.
Published: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Expressing Social Emotions: Misalignment Between LLMs and Human Cultural Emotion Norms
by: Bhattacharyya, Sree, et al.
Published: (2026)
by: Bhattacharyya, Sree, et al.
Published: (2026)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
by: Aggarwal, Yash, et al.
Published: (2026)
by: Aggarwal, Yash, et al.
Published: (2026)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
by: Cai, Yunna, et al.
Published: (2025)
by: Cai, Yunna, et al.
Published: (2025)
Robustness and Confounders in the Demographic Alignment of LLMs with Human Perceptions of Offensiveness
by: Alipour, Shayan, et al.
Published: (2024)
by: Alipour, Shayan, et al.
Published: (2024)
Focal Inferential Infusion Coupled with Tractable Density Discrimination for Implicit Hate Detection
by: Masud, Sarah, et al.
Published: (2023)
by: Masud, Sarah, et al.
Published: (2023)
Steering Large Language Models between Code Execution and Textual Reasoning
by: Chen, Yongchao, et al.
Published: (2024)
by: Chen, Yongchao, et al.
Published: (2024)
Robust Pronoun Fidelity with English LLMs: Are they Reasoning, Repeating, or Just Biased?
by: Gautam, Vagrant, et al.
Published: (2024)
by: Gautam, Vagrant, et al.
Published: (2024)
Eight Methods to Evaluate Robust Unlearning in LLMs
by: Lynch, Aengus, et al.
Published: (2024)
by: Lynch, Aengus, et al.
Published: (2024)
Meursault as a Data Point
by: Pratap, Abhinav
Published: (2025)
by: Pratap, Abhinav
Published: (2025)
IITK at SemEval-2024 Task 2: Exploring the Capabilities of LLMs for Safe Biomedical Natural Language Inference for Clinical Trials
by: Mandal, Shreyasi, et al.
Published: (2024)
by: Mandal, Shreyasi, et al.
Published: (2024)
Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs
by: Pratelli, Manuel, et al.
Published: (2025)
by: Pratelli, Manuel, et al.
Published: (2025)
LLMs are Biased Teachers: Evaluating LLM Bias in Personalized Education
by: Weissburg, Iain, et al.
Published: (2024)
by: Weissburg, Iain, et al.
Published: (2024)
Toward Robust Legal Text Formalization into Defeasible Deontic Logic using LLMs
by: Horner, Elias, et al.
Published: (2025)
by: Horner, Elias, et al.
Published: (2025)
Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2025)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
Beyond prompt brittleness: Evaluating the reliability and consistency of political worldviews in LLMs
by: Ceron, Tanise, et al.
Published: (2024)
by: Ceron, Tanise, et al.
Published: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
by: Hu, Michael Y., et al.
Published: (2025)
by: Hu, Michael Y., et al.
Published: (2025)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
by: Ranjit, Jaspreet, et al.
Published: (2024)
by: Ranjit, Jaspreet, et al.
Published: (2024)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
by: Joshi, Abhinav, et al.
Published: (2024)
by: Joshi, Abhinav, et al.
Published: (2024)
Similar Items
-
Calibration vs Decision Making: Revisiting the Reliability Paradox in Unlearned Language Models
by: Shukla, Divyaksh, et al.
Published: (2026) -
Towards Quantifying Commonsense Reasoning with Mechanistic Insights
by: Joshi, Abhinav, et al.
Published: (2025) -
IITK at SemEval-2024 Task 10: Who is the speaker? Improving Emotion Recognition and Flip Reasoning in Conversations via Speaker Embeddings
by: Patel, Shubham, et al.
Published: (2024) -
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
by: Dwivedi, Ashutosh, et al.
Published: (2025) -
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025)