LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lucas, Tom, Buscemi, Alessio, Capozucca, Alfredo, Castignani, German, Delacroix, Barbara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
von: Huang, Xuanxiang, et al.
Veröffentlicht: (2025)
von: Huang, Xuanxiang, et al.
Veröffentlicht: (2025)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
von: Kelly, Matthew
Veröffentlicht: (2025)
von: Kelly, Matthew
Veröffentlicht: (2025)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026)
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
von: von Cossel, Oskar
Veröffentlicht: (2026)
von: von Cossel, Oskar
Veröffentlicht: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
von: Beltoft, Stine, et al.
Veröffentlicht: (2025)
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
von: Cohen, Liran, et al.
Veröffentlicht: (2025)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
von: Leonesi, Matteo, et al.
Veröffentlicht: (2026)
Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
OntoLogX: Ontology-Guided Knowledge Graph Extraction from Cybersecurity Logs with Large Language Models
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
von: Cotti, Luca, et al.
Veröffentlicht: (2025)
HIP Network: Historical Information Passing Network for Extrapolation Reasoning on Temporal Knowledge Graph
von: He, Yongquan, et al.
Veröffentlicht: (2024)
von: He, Yongquan, et al.
Veröffentlicht: (2024)
AgentLeak: A Full-Stack Benchmark for Privacy Leakage in Multi-Agent LLM Systems
von: Yagoubi, Faouzi El, et al.
Veröffentlicht: (2026)
von: Yagoubi, Faouzi El, et al.
Veröffentlicht: (2026)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
von: Naik, Akshat, et al.
Veröffentlicht: (2025)
von: Naik, Akshat, et al.
Veröffentlicht: (2025)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
von: Heydari, Sina, et al.
Veröffentlicht: (2026)
The Company You Keep: How LLMs Respond to Dark Triad Traits
von: Lu, Zeyi, et al.
Veröffentlicht: (2026)
von: Lu, Zeyi, et al.
Veröffentlicht: (2026)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2024)
Automated Circuit Interpretation via Probe Prompting
von: Birardi, Giuseppe
Veröffentlicht: (2025)
von: Birardi, Giuseppe
Veröffentlicht: (2025)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
von: Bercovich, Ivan, et al.
Veröffentlicht: (2026)
Robustness of Spatio-temporal Graph Neural Networks for Fault Location in Partially Observable Distribution Grids
von: Karabulut, Burak, et al.
Veröffentlicht: (2026)
von: Karabulut, Burak, et al.
Veröffentlicht: (2026)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
von: DeLeeuw, Caleb
Veröffentlicht: (2026)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
von: Yeste, Víctor, et al.
Veröffentlicht: (2026)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
von: Li, Shenghao
Veröffentlicht: (2025)
von: Li, Shenghao
Veröffentlicht: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
von: Wang, Xinyue, et al.
Veröffentlicht: (2026)
von: Wang, Xinyue, et al.
Veröffentlicht: (2026)
MeMo: Towards Language Models with Associative Memory Mechanisms
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
von: Zanzotto, Fabio Massimo, et al.
Veröffentlicht: (2025)
Integration of Contextual Descriptors in Ontology Alignment for Enrichment of Semantic Correspondence
von: Manziuk, Eduard, et al.
Veröffentlicht: (2024)
von: Manziuk, Eduard, et al.
Veröffentlicht: (2024)
Privacy as Commodity: MFG-RegretNet for Large-Scale Privacy Trading in Federated Learning
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
von: Sun, Kangkang, et al.
Veröffentlicht: (2026)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
von: Raman, Vishal, et al.
Veröffentlicht: (2025)
On Fact and Frequency: LLM Responses to Misinformation Expressed with Uncertainty
von: van de Sande, Yana, et al.
Veröffentlicht: (2025)
von: van de Sande, Yana, et al.
Veröffentlicht: (2025)
Comparing Fairness of Generative Mobility Models
von: Wang, Daniel, et al.
Veröffentlicht: (2024)
von: Wang, Daniel, et al.
Veröffentlicht: (2024)
The Limits of Obliviate: Evaluating Unlearning in LLMs via Stimulus-Knowledge Entanglement-Behavior Framework
von: Shah, Aakriti, et al.
Veröffentlicht: (2025)
von: Shah, Aakriti, et al.
Veröffentlicht: (2025)
Computable Gap Assessment of Artificial Intelligence Governance in Children's Centres: Evidence-Mechanism-Governance-Indicator Modelling of UNICEF's Guidance on AI and Children 3.0 Based on the Graph-GAP Framework
von: Meng, Wei
Veröffentlicht: (2025)
von: Meng, Wei
Veröffentlicht: (2025)
Logic interpretations of ANN partition cells
von: Schmitt, Ingo
Veröffentlicht: (2024)
von: Schmitt, Ingo
Veröffentlicht: (2024)
On the Power and Limitations of Examples for Description Logic Concepts
von: Cate, Balder ten, et al.
Veröffentlicht: (2024)
von: Cate, Balder ten, et al.
Veröffentlicht: (2024)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
von: Blanco-Justicia, Alberto, et al.
Veröffentlicht: (2024)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
von: Tiwari, Rishabh, et al.
Veröffentlicht: (2026)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
von: Chua, Jaymari, et al.
Veröffentlicht: (2025)
Enhancing Large Language Models through Neuro-Symbolic Integration and Ontological Reasoning
von: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Veröffentlicht: (2025)
von: Vsevolodovna, Ruslan Idelfonso Magana, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uncovering Bugs in Formal Explainers: A Case Study with PyXAI
von: Huang, Xuanxiang, et al.
Veröffentlicht: (2025) -
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
von: Kelly, Matthew
Veröffentlicht: (2025) -
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
von: Rosenblatt, Lucas, et al.
Veröffentlicht: (2026) -
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
von: Aharon, Eliya Naomi, et al.
Veröffentlicht: (2026) -
Beyond Imperfect Alternatives with Rulemapping: A Neuro-Symbolic Case Study on Online Hate Speech
von: von Cossel, Oskar
Veröffentlicht: (2026)