Policy-Grounded Safety Evaluation of 20 Large Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Contreras, Juan Manuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
von: Palit, Sayon, et al.
Veröffentlicht: (2025)
Large Language Model Interface for Home Energy Management Systems
von: Michelon, François, et al.
Veröffentlicht: (2025)
von: Michelon, François, et al.
Veröffentlicht: (2025)
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
von: Kurian, Ashley, et al.
Veröffentlicht: (2025)
Automated Evaluation of Gender Bias Across 13 Large Multimodal Models
von: Contreras, Juan Manuel
Veröffentlicht: (2025)
von: Contreras, Juan Manuel
Veröffentlicht: (2025)
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
von: Li, Houyi, et al.
Veröffentlicht: (2025)
von: Li, Houyi, et al.
Veröffentlicht: (2025)
Towards Platonic Representation for Table Reasoning: A Foundation for Permutation-Invariant Retrieval
von: Tchuitcheu, Willy Carlos, et al.
Veröffentlicht: (2026)
von: Tchuitcheu, Willy Carlos, et al.
Veröffentlicht: (2026)
The Hidden Attention of Mamba Models
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
von: Ali, Ameen, et al.
Veröffentlicht: (2024)
Evaluating Model Robustness Using Adaptive Sparse L0 Regularization
von: Liu, Weiyou, et al.
Veröffentlicht: (2024)
von: Liu, Weiyou, et al.
Veröffentlicht: (2024)
PromptSAM+: Malware Detection based on Prompt Segment Anything Model
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
von: Wei, Xingyuan, et al.
Veröffentlicht: (2024)
Circularity and Symmetries of $p$ and $p^{2}$-polygons
von: Haag, Rolf
Veröffentlicht: (2025)
von: Haag, Rolf
Veröffentlicht: (2025)
REMoH: A Reflective Evolution of Multi-objective Heuristics approach via Large Language Models
von: Forniés-Tabuenca, Diego, et al.
Veröffentlicht: (2025)
von: Forniés-Tabuenca, Diego, et al.
Veröffentlicht: (2025)
Deep Learning for Human Locomotion Analysis in Lower-Limb Exoskeletons: A Comparative Study
von: Coser, Omar, et al.
Veröffentlicht: (2025)
von: Coser, Omar, et al.
Veröffentlicht: (2025)
Mobile Phone Sensor-based Nigerian Driving Dataset to Detect Alcohol-influenced Behaviours
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
von: Thompson, Iniakpokeikiye Peter, et al.
Veröffentlicht: (2025)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
von: S, Remya Ajai A, et al.
Veröffentlicht: (2024)
Algorithmic Trading Strategy Development and Optimisation
von: Yuan, Owen Nyo Wei, et al.
Veröffentlicht: (2026)
von: Yuan, Owen Nyo Wei, et al.
Veröffentlicht: (2026)
From Pixels to Privacy: Temporally Consistent Video Anonymization via Token Pruning for Privacy Preserving Action Recognition
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
von: Aslam, Nazia, et al.
Veröffentlicht: (2026)
LAraBench: Benchmarking Arabic AI with Large Language Models
von: Abdelali, Ahmed, et al.
Veröffentlicht: (2023)
von: Abdelali, Ahmed, et al.
Veröffentlicht: (2023)
EnergyMamba: An Uncertainty-Aware Graph-Enhanced Selective State Space Model for Energy Consumption Prediction
von: Yu, Dahai, et al.
Veröffentlicht: (2026)
von: Yu, Dahai, et al.
Veröffentlicht: (2026)
Software Implementation of Digital Filtering via Tustin's Bilinear Transform
von: Herron, Connor W.
Veröffentlicht: (2024)
von: Herron, Connor W.
Veröffentlicht: (2024)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
MicroWorld: Empowering Multimodal Large Language Models to Bridge the Microscopic Domain Gap with Multimodal Attribute Graph
von: Li, Manyu, et al.
Veröffentlicht: (2026)
von: Li, Manyu, et al.
Veröffentlicht: (2026)
WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
von: Sun, Zihao, et al.
Veröffentlicht: (2025)
von: Sun, Zihao, et al.
Veröffentlicht: (2025)
Guardians of the Web: The Evolution and Future of Website Information Security
von: Islam, Md Saiful, et al.
Veröffentlicht: (2025)
von: Islam, Md Saiful, et al.
Veröffentlicht: (2025)
Continual Learning, Not Training: Online Adaptation For Agents
von: Jaglan, Aman, et al.
Veröffentlicht: (2025)
von: Jaglan, Aman, et al.
Veröffentlicht: (2025)
Mapping the Urban Mobility Intelligence Frontier: A Scientometric Analysis of Data-Driven Pedestrian Trajectory Prediction and Simulation
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
von: Xu, Junhao, et al.
Veröffentlicht: (2025)
Uncovering Bias Paths with LLM-guided Causal Discovery: An Active Learning and Dynamic Scoring Approach
von: Zanna, Khadija, et al.
Veröffentlicht: (2025)
von: Zanna, Khadija, et al.
Veröffentlicht: (2025)
The impact of postediting on AI generative translation in Yemeni context: Translating literary prose by ChatGPT
von: Al-wagieh, Nasim, et al.
Veröffentlicht: (2026)
von: Al-wagieh, Nasim, et al.
Veröffentlicht: (2026)
RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation
von: Luo, Qinyu, et al.
Veröffentlicht: (2024)
von: Luo, Qinyu, et al.
Veröffentlicht: (2024)
Smart Data-Driven GRU Predictor for SnO$_2$ Thin films Characteristics
von: Bouamra, Faiza, et al.
Veröffentlicht: (2024)
von: Bouamra, Faiza, et al.
Veröffentlicht: (2024)
Distribution Consistency based Self-Training for Graph Neural Networks with Sparse Labels
von: Wang, Fali, et al.
Veröffentlicht: (2024)
von: Wang, Fali, et al.
Veröffentlicht: (2024)
Biometrics Employing Neural Network
von: Bhuiyan, Sajjad
Veröffentlicht: (2024)
von: Bhuiyan, Sajjad
Veröffentlicht: (2024)
HySem: A context length optimized LLM pipeline for unstructured tabular extraction
von: PP, Narayanan, et al.
Veröffentlicht: (2024)
von: PP, Narayanan, et al.
Veröffentlicht: (2024)
OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
von: Alam, Firoj, et al.
Veröffentlicht: (2025)
R-Genie: Reasoning-Guided Generative Image Editing
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
von: Zhang, Dong, et al.
Veröffentlicht: (2025)
Robust Reward Modeling for Large Language Models via Causal Decomposition
von: Lu, Yunsheng, et al.
Veröffentlicht: (2026)
von: Lu, Yunsheng, et al.
Veröffentlicht: (2026)
Factual Dialogue Summarization via Learning from Large Language Models
von: Zhu, Rongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Rongxin, et al.
Veröffentlicht: (2024)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
Evaluating the Limitations of Local LLMs in Solving Complex Programming Challenges
von: Matotek, Kadin, et al.
Veröffentlicht: (2025)
von: Matotek, Kadin, et al.
Veröffentlicht: (2025)
LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
von: Sikdar, Prateek Kumar
Veröffentlicht: (2026)
von: Sikdar, Prateek Kumar
Veröffentlicht: (2026)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
von: Pratyush, Spandan
Veröffentlicht: (2026)
von: Pratyush, Spandan
Veröffentlicht: (2026)
Ähnliche Einträge
-
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
von: Palit, Sayon, et al.
Veröffentlicht: (2025) -
Large Language Model Interface for Home Energy Management Systems
von: Michelon, François, et al.
Veröffentlicht: (2025) -
Train to Defend: First Defense Against Cryptanalytic Neural Network Parameter Extraction Attacks
von: Kurian, Ashley, et al.
Veröffentlicht: (2025) -
Automated Evaluation of Gender Bias Across 13 Large Multimodal Models
von: Contreras, Juan Manuel
Veröffentlicht: (2025) -
Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
von: Li, Houyi, et al.
Veröffentlicht: (2025)