Measuring Moral Inconsistencies in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Bonagiri, Vamshi Krishna, Vennam, Sreeram, Gaur, Manas, Kumaraguru, Ponnurangam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SaGE: Evaluating Moral Consistency in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
LLM Vocabulary Compression for Low-Compute Environments
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
di: Govil, Priyanshul, et al.
Pubblicazione: (2024)
di: Govil, Priyanshul, et al.
Pubblicazione: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
Higher Order Structures For Graph Explanations
di: Sinha, Akshit, et al.
Pubblicazione: (2024)
di: Sinha, Akshit, et al.
Pubblicazione: (2024)
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
di: Mishra, Debangan, et al.
Pubblicazione: (2025)
di: Mishra, Debangan, et al.
Pubblicazione: (2025)
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025)
di: Roy, Amartya, et al.
Pubblicazione: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
Representation Surgery: Theory and Practice of Affine Steering
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
di: Garg, Rahul, et al.
Pubblicazione: (2024)
di: Garg, Rahul, et al.
Pubblicazione: (2024)
Neurosymbolic Retrievers for Retrieval-augmented Generation
di: Saxena, Yash, et al.
Pubblicazione: (2026)
di: Saxena, Yash, et al.
Pubblicazione: (2026)
Enhancing AI Safety Through the Fusion of Low Rank Adapters
di: Gudipudi, Satya Swaroop, et al.
Pubblicazione: (2024)
di: Gudipudi, Satya Swaroop, et al.
Pubblicazione: (2024)
On the Convergence of Moral Self-Correction in Large Language Models
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
di: Liu, Guangliang, et al.
Pubblicazione: (2025)
OncoReason: Structuring Clinical Reasoning in LLMs for Robust and Interpretable Survival Prediction
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
di: Hemadri, Raghu Vamshi, et al.
Pubblicazione: (2025)
Labels Generated by Large Language Models Help Measure People's Empathy in Vitro
di: Hasan, Md Rakibul, et al.
Pubblicazione: (2025)
di: Hasan, Md Rakibul, et al.
Pubblicazione: (2025)
Great Models Think Alike and this Undermines AI Oversight
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
di: Goel, Shashwat, et al.
Pubblicazione: (2025)
IoT-Based Preventive Mental Health Using Knowledge Graphs and Standards for Better Well-Being
di: Gyrard, Amelie, et al.
Pubblicazione: (2024)
di: Gyrard, Amelie, et al.
Pubblicazione: (2024)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
The Moral Gap of Large Language Models
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
di: Skorski, Maciej, et al.
Pubblicazione: (2025)
Long-context Non-factoid Question Answering in Indic Languages
di: Mishra, Ritwik, et al.
Pubblicazione: (2025)
di: Mishra, Ritwik, et al.
Pubblicazione: (2025)
Semantic Sensitivities and Inconsistent Predictions: Measuring the Fragility of NLI Models
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
di: Arakelyan, Erik, et al.
Pubblicazione: (2024)
Multimodal Language Models Cannot Spot Spatial Inconsistencies
di: Khangaonkar, Om, et al.
Pubblicazione: (2026)
di: Khangaonkar, Om, et al.
Pubblicazione: (2026)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
di: Gambardella, Andrew, et al.
Pubblicazione: (2025)
Improving Automatic VQA Evaluation Using Large Language Models
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
di: Mañas, Oscar, et al.
Pubblicazione: (2023)
HLDC: Hindi Legal Documents Corpus
di: Kapoor, Arnav, et al.
Pubblicazione: (2022)
di: Kapoor, Arnav, et al.
Pubblicazione: (2022)
SPIRIT: Short-term Prediction of solar IRradIance for zero-shot Transfer learning using Foundation Models
di: Mishra, Aditya, et al.
Pubblicazione: (2025)
di: Mishra, Aditya, et al.
Pubblicazione: (2025)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
di: Morlat, Geoffroy, et al.
Pubblicazione: (2025)
di: Morlat, Geoffroy, et al.
Pubblicazione: (2025)
Bias after Prompting: Persistent Discrimination in Large Language Models
di: Sivakumar, Nivedha, et al.
Pubblicazione: (2025)
di: Sivakumar, Nivedha, et al.
Pubblicazione: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
di: Aerni, Michael, et al.
Pubblicazione: (2024)
di: Aerni, Michael, et al.
Pubblicazione: (2024)
Human-Interpretable Adversarial Prompt Attack on Large Language Models with Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
A Moral Imperative: The Need for Continual Superalignment of Large Language Models
di: Puthumanaillam, Gokul, et al.
Pubblicazione: (2024)
di: Puthumanaillam, Gokul, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SaGE: Evaluating Moral Consistency in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024) -
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025) -
LLM Vocabulary Compression for Low-Compute Environments
di: Vennam, Sreeram, et al.
Pubblicazione: (2024) -
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
di: Vennam, Sreeram, et al.
Pubblicazione: (2024) -
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)