TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
Fuente:
arXiv
Salvato in:
| Autori principali: | Hui, Zheng, Dong, Yijiang River, Shareghi, Ehsan, Collier, Nigel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
di: Dong, Yijiang River, et al.
Pubblicazione: (2024)
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
di: Hui, Zheng, et al.
Pubblicazione: (2025)
di: Hui, Zheng, et al.
Pubblicazione: (2025)
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
di: Dong, Yijiang River, et al.
Pubblicazione: (2026)
di: Dong, Yijiang River, et al.
Pubblicazione: (2026)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
ReasonGraph: Visualisation of Reasoning Paths
di: Li, Zongqian, et al.
Pubblicazione: (2025)
di: Li, Zongqian, et al.
Pubblicazione: (2025)
Quantifying the Persona Effect in LLM Simulations
di: Hu, Tiancheng, et al.
Pubblicazione: (2024)
di: Hu, Tiancheng, et al.
Pubblicazione: (2024)
Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
di: Zhou, Han, et al.
Pubblicazione: (2024)
di: Zhou, Han, et al.
Pubblicazione: (2024)
One STEP at a time: Language Agents are Stepwise Planners
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
di: Nguyen, Minh, et al.
Pubblicazione: (2024)
Interpreting Latent Student Knowledge Representations in Programming Assignments
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
di: Liu, Yinhong, et al.
Pubblicazione: (2024)
PiVe: Prompting with Iterative Verification Improving Graph-based Generative Capability of LLMs
di: Han, Jiuzhou, et al.
Pubblicazione: (2023)
di: Han, Jiuzhou, et al.
Pubblicazione: (2023)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
di: Dong, Zhichen, et al.
Pubblicazione: (2024)
di: Dong, Zhichen, et al.
Pubblicazione: (2024)
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM Experts
di: Ghosh, Shaona, et al.
Pubblicazione: (2024)
di: Ghosh, Shaona, et al.
Pubblicazione: (2024)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
di: Ren, Richard, et al.
Pubblicazione: (2024)
di: Ren, Richard, et al.
Pubblicazione: (2024)
Value of Information: A Framework for Human-Agent Communication
di: Dong, Yijiang River, et al.
Pubblicazione: (2026)
di: Dong, Yijiang River, et al.
Pubblicazione: (2026)
iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
di: Hu, Tiancheng, et al.
Pubblicazione: (2025)
All Roads Lead to Rome: Graph-Based Confidence Estimation for Large Language Model Reasoning
di: Zhang, Caiqi, et al.
Pubblicazione: (2025)
di: Zhang, Caiqi, et al.
Pubblicazione: (2025)
Test Case-Informed Knowledge Tracing for Open-ended Coding Tasks
di: Duan, Zhangqi, et al.
Pubblicazione: (2024)
di: Duan, Zhangqi, et al.
Pubblicazione: (2024)
When Personalization Meets Reality: A Multi-Faceted Analysis of Personalized Preference Learning
di: Dong, Yijiang River, et al.
Pubblicazione: (2025)
di: Dong, Yijiang River, et al.
Pubblicazione: (2025)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
di: Scarlatos, Alexander, et al.
Pubblicazione: (2025)
di: Scarlatos, Alexander, et al.
Pubblicazione: (2025)
DiVERT: Distractor Generation with Variational Errors Represented as Text for Math Multiple-choice Questions
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance
di: Theuma, Adrian, et al.
Pubblicazione: (2024)
di: Theuma, Adrian, et al.
Pubblicazione: (2024)
AppellateGen: A Benchmark for Appellate Legal Judgment Generation
di: Yang, Hongkun, et al.
Pubblicazione: (2026)
di: Yang, Hongkun, et al.
Pubblicazione: (2026)
Identifying Climate Targets in National Laws and Policies using Machine Learning
di: Juhasz, Matyas, et al.
Pubblicazione: (2024)
di: Juhasz, Matyas, et al.
Pubblicazione: (2024)
SyllabusQA: A Course Logistics Question Answering Dataset
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
di: Fernandez, Nigel, et al.
Pubblicazione: (2024)
KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
di: Duan, Zhangqi, et al.
Pubblicazione: (2026)
di: Duan, Zhangqi, et al.
Pubblicazione: (2026)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
di: Tedeschi, Simone, et al.
Pubblicazione: (2024)
di: Tedeschi, Simone, et al.
Pubblicazione: (2024)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
di: Wang, Minghan, et al.
Pubblicazione: (2026)
di: Wang, Minghan, et al.
Pubblicazione: (2026)
Baichuan4-Finance Technical Report
di: Zhang, Hanyu, et al.
Pubblicazione: (2024)
di: Zhang, Hanyu, et al.
Pubblicazione: (2024)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
di: Dumitran, Adrian-Marius, et al.
Pubblicazione: (2025)
GG-BBQ: German Gender Bias Benchmark for Question Answering
di: Satheesh, Shalaka, et al.
Pubblicazione: (2025)
di: Satheesh, Shalaka, et al.
Pubblicazione: (2025)
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
Unintended Impacts of LLM Alignment on Global Representation
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
di: Ryan, Michael J., et al.
Pubblicazione: (2024)
DetoxLLM: A Framework for Detoxification with Explanations
di: Khondaker, Md Tawkat Islam, et al.
Pubblicazione: (2024)
di: Khondaker, Md Tawkat Islam, et al.
Pubblicazione: (2024)
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2023)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
di: Guldimann, Philipp, et al.
Pubblicazione: (2024)
di: Guldimann, Philipp, et al.
Pubblicazione: (2024)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
di: Majumdar, Ayan, et al.
Pubblicazione: (2025)
di: Majumdar, Ayan, et al.
Pubblicazione: (2025)
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
di: Peczuh, Marisa C., et al.
Pubblicazione: (2025)
Strategic Demonstration Selection for Improved Fairness in LLM In-Context Learning
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
di: Hu, Jingyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can LLM be a Personalized Judge?
di: Dong, Yijiang River, et al.
Pubblicazione: (2024) -
Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning
di: Hui, Zheng, et al.
Pubblicazione: (2025) -
Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
di: Dong, Yijiang River, et al.
Pubblicazione: (2026) -
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
di: Hu, Tiancheng, et al.
Pubblicazione: (2025) -
ReasonGraph: Visualisation of Reasoning Paths
di: Li, Zongqian, et al.
Pubblicazione: (2025)