Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Beltoft, Stine, Galke, Lukas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
by: Yeste, Víctor, et al.
Published: (2026)
by: Yeste, Víctor, et al.
Published: (2026)
The AI Fiction Paradox
by: Elkins, Katherine
Published: (2026)
by: Elkins, Katherine
Published: (2026)
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
by: Lucas, Tom, et al.
Published: (2026)
by: Lucas, Tom, et al.
Published: (2026)
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
by: Leonesi, Matteo, et al.
Published: (2026)
by: Leonesi, Matteo, et al.
Published: (2026)
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
by: Cohen, Liran, et al.
Published: (2025)
by: Cohen, Liran, et al.
Published: (2025)
The Epistemic Suite: A Post-Foundational Diagnostic Methodology for Assessing AI Knowledge Claims
by: Kelly, Matthew
Published: (2025)
by: Kelly, Matthew
Published: (2025)
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
by: Naik, Akshat, et al.
Published: (2025)
by: Naik, Akshat, et al.
Published: (2025)
LFC-DA: Logical Formula-Controlled Data Augmentation for Enhanced Logical Reasoning
by: Li, Shenghao
Published: (2025)
by: Li, Shenghao
Published: (2025)
Controlled Territory and Conflict Tracking (CONTACT): (Geo-)Mapping Occupied Territory from Open Source Intelligence
by: Mandal, Paul K., et al.
Published: (2025)
by: Mandal, Paul K., et al.
Published: (2025)
BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders
by: DeLeeuw, Caleb
Published: (2026)
by: DeLeeuw, Caleb
Published: (2026)
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories
by: Bercovich, Ivan, et al.
Published: (2026)
by: Bercovich, Ivan, et al.
Published: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning
by: Kim, Kyuhee, et al.
Published: (2026)
by: Kim, Kyuhee, et al.
Published: (2026)
Privately Fine-Tuned LLMs Preserve Temporal Dynamics in Tabular Data
by: Rosenblatt, Lucas, et al.
Published: (2026)
by: Rosenblatt, Lucas, et al.
Published: (2026)
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
by: Blanco-Justicia, Alberto, et al.
Published: (2024)
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
by: Madan, Kapil
Published: (2025)
by: Madan, Kapil
Published: (2025)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation
by: Zhu, Pengyun, et al.
Published: (2026)
by: Zhu, Pengyun, et al.
Published: (2026)
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
by: Webster, Kevin T
Published: (2025)
by: Webster, Kevin T
Published: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
by: Boumber, Dainis, et al.
Published: (2024)
by: Boumber, Dainis, et al.
Published: (2024)
Training Language Models to Use Prolog as a Tool
by: Mellgren, Niklas, et al.
Published: (2025)
by: Mellgren, Niklas, et al.
Published: (2025)
Qwerty AI: Explainable Automated Age Rating and Content Safety Assessment for Russian-Language Screenplays
by: Zmanovskii, Nikita
Published: (2025)
by: Zmanovskii, Nikita
Published: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Extreme Self-Preference in Language Models
by: Lehr, Steven A., et al.
Published: (2025)
by: Lehr, Steven A., et al.
Published: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
Exploring and Mitigating Gender Bias in Encoder-Based Transformer Models
by: Hossain, Ariyan, et al.
Published: (2025)
by: Hossain, Ariyan, et al.
Published: (2025)
Measuring the Authority Stack of AI Systems: Empirical Analysis of 366,120 Forced-Choice Responses Across 8 AI Models
by: Lee, Seulki
Published: (2026)
by: Lee, Seulki
Published: (2026)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
by: Basu, Abhinaba, et al.
Published: (2026)
by: Basu, Abhinaba, et al.
Published: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedback
by: Du, Yishan, et al.
Published: (2025)
by: Du, Yishan, et al.
Published: (2025)
MeMo: Towards Language Models with Associative Memory Mechanisms
by: Zanzotto, Fabio Massimo, et al.
Published: (2025)
by: Zanzotto, Fabio Massimo, et al.
Published: (2025)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
by: Heydari, Sina, et al.
Published: (2026)
by: Heydari, Sina, et al.
Published: (2026)
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
by: Kryshtal, Andrii
Published: (2026)
by: Kryshtal, Andrii
Published: (2026)
Revealing Hidden Bias in AI: Lessons from Large Language Models
by: Beatty, Django, et al.
Published: (2024)
by: Beatty, Django, et al.
Published: (2024)
Replicating TEMPEST at Scale: Multi-Turn Adversarial Attacks Against Trillion-Parameter Frontier Models
by: Young, Richard
Published: (2025)
by: Young, Richard
Published: (2025)
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation
by: Hartmann, David, et al.
Published: (2026)
by: Hartmann, David, et al.
Published: (2026)
Similar Items
-
Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
by: Yeste, Víctor, et al.
Published: (2026) -
More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
by: Yeste, Víctor, et al.
Published: (2026) -
Do Schwartz Higher-Order Values Help Sentence-Level Human Value Detection? A Study of Hierarchical Gating and Calibration
by: Yeste, Víctor, et al.
Published: (2026) -
The AI Fiction Paradox
by: Elkins, Katherine
Published: (2026) -
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
by: Lucas, Tom, et al.
Published: (2026)