Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
Fuente:
arXiv
Saved in:
| Main Authors: | Atil, Berk, Gupta, Vipul, Das, Sarkar Snigdha Sarathi, Passonneau, Rebecca J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Model Unlearning Objectives Vary for Distinct Language Functions
by: Atil, Berk, et al.
Published: (2026)
by: Atil, Berk, et al.
Published: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
by: Atil, Berk, et al.
Published: (2026)
by: Atil, Berk, et al.
Published: (2026)
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025)
by: Atil, Berk, et al.
Published: (2025)
VerAs: Verify then Assess STEM Lab Reports
by: Atil, Berk, et al.
Published: (2024)
by: Atil, Berk, et al.
Published: (2024)
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
by: Das, Sarkar Snigdha Sarathi, et al.
Published: (2024)
Verbosity $\neq$ Veracity: Demystify Verbosity Compensation Behavior of Large Language Models
by: Zhang, Yusen, et al.
Published: (2024)
by: Zhang, Yusen, et al.
Published: (2024)
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
by: Gupta, Rushil, et al.
Published: (2025)
by: Gupta, Rushil, et al.
Published: (2025)
Sociodemographic Bias in Language Models: A Survey and Forward Path
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
by: Gupta, Vipul, et al.
Published: (2023)
by: Gupta, Vipul, et al.
Published: (2023)
Can Editing LLMs Inject Harm?
by: Chen, Canyu, et al.
Published: (2024)
by: Chen, Canyu, et al.
Published: (2024)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
Joint Training for Selective Prediction
by: Li, Zhaohui, et al.
Published: (2024)
by: Li, Zhaohui, et al.
Published: (2024)
Still Not There: Can LLMs Outperform Smaller Task-Specific Seq2Seq Models on the Poetry-to-Prose Conversion Task?
by: Das, Kunal Kingkar, et al.
Published: (2025)
by: Das, Kunal Kingkar, et al.
Published: (2025)
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
by: Karizaki, Mahsa Sheikhi, et al.
Published: (2024)
by: Karizaki, Mahsa Sheikhi, et al.
Published: (2024)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Evaluating LLMs at Detecting Errors in LLM Responses
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
by: Yang, Haoyan, et al.
Published: (2024)
by: Yang, Haoyan, et al.
Published: (2024)
Can We Locate and Prevent Stereotypes in LLMs?
by: D'Souza, Alex
Published: (2026)
by: D'Souza, Alex
Published: (2026)
Efficient PRM Training Data Synthesis via Formal Verification
by: Kamoi, Ryo, et al.
Published: (2025)
by: Kamoi, Ryo, et al.
Published: (2025)
Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers
by: Qi, Zhenting, et al.
Published: (2024)
by: Qi, Zhenting, et al.
Published: (2024)
LLM-REVal: Can We Trust LLM Reviewers Yet?
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Can LLMs Correct Physicians, Yet? Investigating Effective Interaction Methods in the Medical Domain
by: Sayin, Burcu, et al.
Published: (2024)
by: Sayin, Burcu, et al.
Published: (2024)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
by: Neumann, Terrence, et al.
Published: (2024)
by: Neumann, Terrence, et al.
Published: (2024)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
Can We Edit LLMs for Long-Tail Biomedical Knowledge?
by: Yi, Xinhao, et al.
Published: (2025)
by: Yi, Xinhao, et al.
Published: (2025)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
by: Bianchi, Owen, et al.
Published: (2025)
by: Bianchi, Owen, et al.
Published: (2025)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
by: Wu, Taiqiang, et al.
Published: (2026)
by: Wu, Taiqiang, et al.
Published: (2026)
Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling
by: Bansal, Hritik, et al.
Published: (2024)
by: Bansal, Hritik, et al.
Published: (2024)
Catch Me If You Can? Not Yet: LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors
by: Wang, Zhengxiang, et al.
Published: (2025)
by: Wang, Zhengxiang, et al.
Published: (2025)
Can LLMs Generate High-Quality Task-Specific Conversations?
by: Li, Shengqi, et al.
Published: (2025)
by: Li, Shengqi, et al.
Published: (2025)
An Evaluation of LLMs for Detecting Harmful Computing Terms
by: Jacas, Joshua, et al.
Published: (2025)
by: Jacas, Joshua, et al.
Published: (2025)
Non-Determinism of "Deterministic" LLM Settings
by: Atil, Berk, et al.
Published: (2024)
by: Atil, Berk, et al.
Published: (2024)
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
by: Loweimi, Erfan, et al.
Published: (2026)
by: Loweimi, Erfan, et al.
Published: (2026)
Can LLMs Grasp Implicit Cultural Values? Benchmarking LLMs' Cultural Intelligence with CQ-Bench
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
Changing Answer Order Can Decrease MMLU Accuracy
by: Gupta, Vipul, et al.
Published: (2024)
by: Gupta, Vipul, et al.
Published: (2024)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
Can We Infer Confidential Properties of Training Data from LLMs?
by: Huang, Pengrun, et al.
Published: (2025)
by: Huang, Pengrun, et al.
Published: (2025)
Similar Items
-
Model Unlearning Objectives Vary for Distinct Language Functions
by: Atil, Berk, et al.
Published: (2026) -
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
by: Atil, Berk, et al.
Published: (2025) -
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
by: Atil, Berk, et al.
Published: (2026) -
Something Just Like TRuST : Toxicity Recognition of Span and Target
by: Atil, Berk, et al.
Published: (2025) -
VerAs: Verify then Assess STEM Lab Reports
by: Atil, Berk, et al.
Published: (2024)