Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisions
Fuente:
arXiv
Salvato in:
| Autori principali: | Elkins, Katherine, Chun, Jon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The AI Fiction Paradox
di: Elkins, Katherine
Pubblicazione: (2026)
di: Elkins, Katherine
Pubblicazione: (2026)
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
di: Chun, Jon, et al.
Pubblicazione: (2024)
di: Chun, Jon, et al.
Pubblicazione: (2024)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
di: Tikhonov, Alexey, et al.
Pubblicazione: (2026)
di: Tikhonov, Alexey, et al.
Pubblicazione: (2026)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
di: Sela, Omer
Pubblicazione: (2026)
di: Sela, Omer
Pubblicazione: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
di: Hari, Vishnu, et al.
Pubblicazione: (2025)
di: Hari, Vishnu, et al.
Pubblicazione: (2025)
ARF-RLHF: Adaptive Reward-Following for RLHF through Emotion-Driven Self-Supervision and Trace-Biased Dynamic Optimization
di: Zhang, YuXuan
Pubblicazione: (2025)
di: Zhang, YuXuan
Pubblicazione: (2025)
Post-Training Probability Manifold Correction via Structured SVD Pruning and Self-Referential Distillation
di: Flouro, Aaron R., et al.
Pubblicazione: (2026)
di: Flouro, Aaron R., et al.
Pubblicazione: (2026)
An Automatic Text Classification Method Based on Hierarchical Taxonomies, Neural Networks and Document Embedding: The NETHIC Tool
di: Lomasto, Luigi, et al.
Pubblicazione: (2026)
di: Lomasto, Luigi, et al.
Pubblicazione: (2026)
RL-LLM-DT: An Automatic Decision Tree Generation Method Based on RL Evaluation and LLM Enhancement
di: Lin, Junjie, et al.
Pubblicazione: (2024)
di: Lin, Junjie, et al.
Pubblicazione: (2024)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
Task Memory Engine (TME): Enhancing State Awareness for Multi-Step LLM Agent Tasks
di: Ye, Ye
Pubblicazione: (2025)
di: Ye, Ye
Pubblicazione: (2025)
Monotonicity as an Architectural Bias for Robust Language Models
di: Cooper, Patrick, et al.
Pubblicazione: (2026)
di: Cooper, Patrick, et al.
Pubblicazione: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
di: Pizzo, David Alejandro Trejo
Pubblicazione: (2026)
di: Pizzo, David Alejandro Trejo
Pubblicazione: (2026)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
di: Basu, Abhinaba, et al.
Pubblicazione: (2026)
di: Basu, Abhinaba, et al.
Pubblicazione: (2026)
The Mirror Loop: Recursive Non-Convergence in Generative Reasoning Systems
di: DeVilling, Bentley
Pubblicazione: (2025)
di: DeVilling, Bentley
Pubblicazione: (2025)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
di: Mutlu, Abdulvahap, et al.
Pubblicazione: (2026)
di: Mutlu, Abdulvahap, et al.
Pubblicazione: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
di: Chen, Yihong, et al.
Pubblicazione: (2022)
di: Chen, Yihong, et al.
Pubblicazione: (2022)
Dynamic Dual-Granularity Skill Bank for Agentic RL
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
di: Basu, Abhinaba, et al.
Pubblicazione: (2026)
di: Basu, Abhinaba, et al.
Pubblicazione: (2026)
BitSkip: An Empirical Analysis of Quantization and Early Exit Composition in Transformers
di: Bhuvaneswaran, Ramshankar, et al.
Pubblicazione: (2025)
di: Bhuvaneswaran, Ramshankar, et al.
Pubblicazione: (2025)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
The Efficiency Attenuation Phenomenon: A Computational Challenge to the Language of Thought Hypothesis
di: Zhang, Di
Pubblicazione: (2026)
di: Zhang, Di
Pubblicazione: (2026)
Measuring Intent Comprehension in LLMs
di: Kunievsky, Nadav, et al.
Pubblicazione: (2025)
di: Kunievsky, Nadav, et al.
Pubblicazione: (2025)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
di: Jia, Xiao
Pubblicazione: (2026)
di: Jia, Xiao
Pubblicazione: (2026)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
di: Kaiser, Daniel, et al.
Pubblicazione: (2026)
di: Kaiser, Daniel, et al.
Pubblicazione: (2026)
KDSTM: Neural Semi-supervised Topic Modeling with Knowledge Distillation
di: Xu, Weijie, et al.
Pubblicazione: (2023)
di: Xu, Weijie, et al.
Pubblicazione: (2023)
When Your Own Output Becomes Your Training Data: Noise-to-Meaning Loops and a Formal RSI Trigger
di: Ando, Rintaro
Pubblicazione: (2025)
di: Ando, Rintaro
Pubblicazione: (2025)
Chronicals: A High-Performance Framework for LLM Fine-Tuning with 3.51x Speedup over Unsloth
di: Nair, Arjun S.
Pubblicazione: (2026)
di: Nair, Arjun S.
Pubblicazione: (2026)
BreakFun: Jailbreaking LLMs via Schema Exploitation
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
di: Oskooei, Amirkia Rafiei, et al.
Pubblicazione: (2025)
Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping
di: Lai, Guannan, et al.
Pubblicazione: (2025)
di: Lai, Guannan, et al.
Pubblicazione: (2025)
Harnessing non-adversarial robustness in large language models
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
di: Zhou, Qinghua, et al.
Pubblicazione: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
di: Hill, Brennen
Pubblicazione: (2025)
di: Hill, Brennen
Pubblicazione: (2025)
Autonomous Multi-Robot Infrastructure for AI-Enabled Healthcare Delivery and Diagnostics
di: Kalaivanan, Nakhul, et al.
Pubblicazione: (2025)
di: Kalaivanan, Nakhul, et al.
Pubblicazione: (2025)
It's 2025 -- Narrative Learning is the new baseline to beat for explainable machine learning
di: Baker, Gregory D.
Pubblicazione: (2025)
di: Baker, Gregory D.
Pubblicazione: (2025)
Pre-train to Gain: Robust Learning Without Clean Labels
di: Szczecina, David, et al.
Pubblicazione: (2025)
di: Szczecina, David, et al.
Pubblicazione: (2025)
LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
di: Weigang, Li, et al.
Pubblicazione: (2025)
di: Weigang, Li, et al.
Pubblicazione: (2025)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
di: Keeman, Michael
Pubblicazione: (2026)
di: Keeman, Michael
Pubblicazione: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
di: Mathew, Aby Mammen
Pubblicazione: (2026)
di: Mathew, Aby Mammen
Pubblicazione: (2026)
Documenti analoghi
-
The AI Fiction Paradox
di: Elkins, Katherine
Pubblicazione: (2026) -
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values
di: Chun, Jon, et al.
Pubblicazione: (2024) -
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025) -
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
di: Tikhonov, Alexey, et al.
Pubblicazione: (2026) -
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
di: Sela, Omer
Pubblicazione: (2026)