Dynamic Evaluation for Oversensitivity in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Pu, Sophia Xiao, Cheng, Sitao, Wang, Xin Eric, Wang, William Yang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
di: Li, Xirui, et al.
Pubblicazione: (2024)
di: Li, Xirui, et al.
Pubblicazione: (2024)
Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
di: Cheng, Sitao, et al.
Pubblicazione: (2024)
di: Cheng, Sitao, et al.
Pubblicazione: (2024)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
di: Zhou, Ruiwen, et al.
Pubblicazione: (2024)
di: Zhou, Ruiwen, et al.
Pubblicazione: (2024)
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
di: Cai, Yunna, et al.
Pubblicazione: (2025)
di: Cai, Yunna, et al.
Pubblicazione: (2025)
Disentangling Memory and Reasoning Ability in Large Language Models
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
di: Jin, Mingyu, et al.
Pubblicazione: (2024)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
di: Pu, Xiao, et al.
Pubblicazione: (2025)
di: Pu, Xiao, et al.
Pubblicazione: (2025)
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
di: Ni, Jiliang, et al.
Pubblicazione: (2025)
di: Ni, Jiliang, et al.
Pubblicazione: (2025)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
di: Xiao, Yang, et al.
Pubblicazione: (2023)
di: Xiao, Yang, et al.
Pubblicazione: (2023)
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
di: Cheng, Sitao, et al.
Pubblicazione: (2025)
di: Cheng, Sitao, et al.
Pubblicazione: (2025)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
di: Liu, Zhiwei, et al.
Pubblicazione: (2025)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
di: Long, Lin, et al.
Pubblicazione: (2024)
di: Long, Lin, et al.
Pubblicazione: (2024)
Can LLMs Solve longer Math Word Problems Better?
di: Xu, Xin, et al.
Pubblicazione: (2024)
di: Xu, Xin, et al.
Pubblicazione: (2024)
Designing and Evaluating Dialogue LLMs for Co-Creative Improvised Theatre
di: Branch, Boyd, et al.
Pubblicazione: (2024)
di: Branch, Boyd, et al.
Pubblicazione: (2024)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
LEDOM: Reverse Language Model
di: Yin, Xunjian, et al.
Pubblicazione: (2025)
di: Yin, Xunjian, et al.
Pubblicazione: (2025)
Compass-v3: Scaling Domain-Specific LLMs for Multilingual E-Commerce in Southeast Asia
di: Maria, Sophia
Pubblicazione: (2025)
di: Maria, Sophia
Pubblicazione: (2025)
MCEval: A Dynamic Framework for Fair Multilingual Cultural Evaluation of LLMs
di: Huang, Shulin, et al.
Pubblicazione: (2025)
di: Huang, Shulin, et al.
Pubblicazione: (2025)
Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Bridging the Knowledge-Action Gap by Evaluating LLMs in Dynamic Dental Clinical Scenarios
di: Ma, Hongyang, et al.
Pubblicazione: (2026)
di: Ma, Hongyang, et al.
Pubblicazione: (2026)
LLMs vs. Chinese Anime Enthusiasts: A Comparative Study on Emotionally Supportive Role-Playing
di: Qiu, Lanlan, et al.
Pubblicazione: (2025)
di: Qiu, Lanlan, et al.
Pubblicazione: (2025)
TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
di: Huang, Xiang, et al.
Pubblicazione: (2024)
di: Huang, Xiang, et al.
Pubblicazione: (2024)
Semantics-Adaptive Activation Intervention for LLMs via Dynamic Steering Vectors
di: Wang, Weixuan, et al.
Pubblicazione: (2024)
di: Wang, Weixuan, et al.
Pubblicazione: (2024)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
di: Liu, Heyang, et al.
Pubblicazione: (2025)
di: Liu, Heyang, et al.
Pubblicazione: (2025)
PFID: Privacy First Inference Delegation Framework for LLMs
di: Yang, Haoyan, et al.
Pubblicazione: (2024)
di: Yang, Haoyan, et al.
Pubblicazione: (2024)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
di: Yu, Zeping, et al.
Pubblicazione: (2024)
di: Yu, Zeping, et al.
Pubblicazione: (2024)
DITING: A Multi-Agent Evaluation Framework for Benchmarking Web Novel Translation
di: Zhang, Enze, et al.
Pubblicazione: (2025)
di: Zhang, Enze, et al.
Pubblicazione: (2025)
Understanding the Effects of Domain Finetuning on LLMs
di: Tanwar, Eshaan, et al.
Pubblicazione: (2025)
di: Tanwar, Eshaan, et al.
Pubblicazione: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
di: Liu, Chengzhi, et al.
Pubblicazione: (2026)
di: Liu, Chengzhi, et al.
Pubblicazione: (2026)
Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
How Good are LLMs at Relation Extraction under Low-Resource Scenario? Comprehensive Evaluation
di: Jinensibieke, Dawulie, et al.
Pubblicazione: (2024)
di: Jinensibieke, Dawulie, et al.
Pubblicazione: (2024)
MolViBench: Evaluating LLMs on Molecular Vibe Coding
di: Li, Jiatong, et al.
Pubblicazione: (2026)
di: Li, Jiatong, et al.
Pubblicazione: (2026)
EvoWiki: Evaluating LLMs on Evolving Knowledge
di: Tang, Wei, et al.
Pubblicazione: (2024)
di: Tang, Wei, et al.
Pubblicazione: (2024)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
di: Brown, Oscar, et al.
Pubblicazione: (2024)
di: Brown, Oscar, et al.
Pubblicazione: (2024)
Guardians of Discourse: Evaluating LLMs on Multilingual Offensive Language Detection
di: He, Jianfei, et al.
Pubblicazione: (2024)
di: He, Jianfei, et al.
Pubblicazione: (2024)
Differentiable Evolutionary Reinforcement Learning
di: Cheng, Sitao, et al.
Pubblicazione: (2025)
di: Cheng, Sitao, et al.
Pubblicazione: (2025)
SciDA: Scientific Dynamic Assessor of LLMs
di: Zhou, Junting, et al.
Pubblicazione: (2025)
di: Zhou, Junting, et al.
Pubblicazione: (2025)
Evaluating the Effectiveness of Black-Box Prompt Optimization as the Scale of LLMs Continues to Grow
di: Zhou, Ziyu, et al.
Pubblicazione: (2025)
di: Zhou, Ziyu, et al.
Pubblicazione: (2025)
Can LLMs Track Their Output Length? A Dynamic Feedback Mechanism for Precise Length Regulation
di: Xiao, Meiman, et al.
Pubblicazione: (2026)
di: Xiao, Meiman, et al.
Pubblicazione: (2026)
Unveiling the Competitive Dynamics: A Comparative Evaluation of American and Chinese LLMs
di: Jiang, Zhenhui, et al.
Pubblicazione: (2024)
di: Jiang, Zhenhui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
di: Li, Xirui, et al.
Pubblicazione: (2024) -
Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
di: Cheng, Sitao, et al.
Pubblicazione: (2024) -
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
di: Zhou, Ruiwen, et al.
Pubblicazione: (2024) -
Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
di: Cai, Yunna, et al.
Pubblicazione: (2025) -
Disentangling Memory and Reasoning Ability in Large Language Models
di: Jin, Mingyu, et al.
Pubblicazione: (2024)