SaGE: Evaluating Moral Consistency in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Bonagiri, Vamshi Krishna, Vennam, Sreeram, Govil, Priyanshul, Kumaraguru, Ponnurangam, Gaur, Manas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Measuring Moral Inconsistencies in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
di: Govil, Priyanshul, et al.
Pubblicazione: (2024)
di: Govil, Priyanshul, et al.
Pubblicazione: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
di: Aneja, Krishak, et al.
Pubblicazione: (2026)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
di: Kodali, Prashant, et al.
Pubblicazione: (2024)
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
KnowledgePrompts: Exploring the Abilities of Large Language Models to Solve Proportional Analogies via Knowledge-Enhanced Prompting
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
di: Wijesiriwardene, Thilini, et al.
Pubblicazione: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
di: Joishy, Anish R, et al.
Pubblicazione: (2025)
LLM Vocabulary Compression for Low-Compute Environments
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
di: Vennam, Sreeram, et al.
Pubblicazione: (2024)
K-PERM: Personalized Response Generation Using Dynamic Knowledge Retrieval and Persona-Adaptive Queries
di: Raj, Kanak, et al.
Pubblicazione: (2023)
di: Raj, Kanak, et al.
Pubblicazione: (2023)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
di: Vedula, Bhaskara Hanuma, et al.
Pubblicazione: (2026)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2025)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2026)
Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2025)
Multilingual Coreference Resolution in Low-resource South Asian Languages
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
di: Vamshi, Bodla Krishna, et al.
Pubblicazione: (2026)
InSaAF: Incorporating Safety through Accuracy and Fairness | Are LLMs ready for the Indian Legal Domain?
di: Tripathi, Yogesh, et al.
Pubblicazione: (2024)
di: Tripathi, Yogesh, et al.
Pubblicazione: (2024)
Structured Definitions and Segmentations for Legal Reasoning in LLMs: A Study on Indian Legal Data
di: Khatri, Mann, et al.
Pubblicazione: (2025)
di: Khatri, Mann, et al.
Pubblicazione: (2025)
The Moral Consistency Pipeline: Continuous Ethical Evaluation for Large Language Models
di: Jamshidi, Saeid, et al.
Pubblicazione: (2025)
di: Jamshidi, Saeid, et al.
Pubblicazione: (2025)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
di: Mohammadi, Seyedali, et al.
Pubblicazione: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
di: Das, Nilanjana, et al.
Pubblicazione: (2026)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
SEMMA: A Semantic Aware Knowledge Graph Foundation Model
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
di: Arun, Arvindh, et al.
Pubblicazione: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
di: Sinha, Shiven, et al.
Pubblicazione: (2024)
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
di: Lamba, Naveen, et al.
Pubblicazione: (2025)
Higher Order Structures For Graph Explanations
di: Sinha, Akshit, et al.
Pubblicazione: (2024)
di: Sinha, Akshit, et al.
Pubblicazione: (2024)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
di: Saxena, Yash, et al.
Pubblicazione: (2024)
di: Saxena, Yash, et al.
Pubblicazione: (2024)
Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs
di: P, Dinesh Srivasthav, et al.
Pubblicazione: (2026)
di: P, Dinesh Srivasthav, et al.
Pubblicazione: (2026)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
di: P, Vedanta S, et al.
Pubblicazione: (2026)
di: P, Vedanta S, et al.
Pubblicazione: (2026)
Just KIDDIN: Knowledge Infusion and Distillation for Detection of INdecent Memes
di: Garg, Rahul, et al.
Pubblicazione: (2024)
di: Garg, Rahul, et al.
Pubblicazione: (2024)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
di: Huang, Allison, et al.
Pubblicazione: (2024)
di: Huang, Allison, et al.
Pubblicazione: (2024)
DCR-Consistency: Divide-Conquer-Reasoning for Consistency Evaluation and Improvement of Large Language Models
di: Cui, Wendi, et al.
Pubblicazione: (2024)
di: Cui, Wendi, et al.
Pubblicazione: (2024)
Neurosymbolic Retrievers for Retrieval-augmented Generation
di: Saxena, Yash, et al.
Pubblicazione: (2026)
di: Saxena, Yash, et al.
Pubblicazione: (2026)
A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models
di: Petkar, Soham, et al.
Pubblicazione: (2025)
di: Petkar, Soham, et al.
Pubblicazione: (2025)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
di: Das, Nilanjana, et al.
Pubblicazione: (2024)
CMoralEval: A Moral Evaluation Benchmark for Chinese Large Language Models
di: Yu, Linhao, et al.
Pubblicazione: (2024)
di: Yu, Linhao, et al.
Pubblicazione: (2024)
Cross-Examiner: Evaluating Consistency of Large Language Model-Generated Explanations
di: Villa, Danielle, et al.
Pubblicazione: (2025)
di: Villa, Danielle, et al.
Pubblicazione: (2025)
Firm or Fickle? Evaluating Large Language Models Consistency in Sequential Interactions
di: Li, Yubo, et al.
Pubblicazione: (2025)
di: Li, Yubo, et al.
Pubblicazione: (2025)
Tracing Moral Foundations in Large Language Models
di: Yu, Chenxiao, et al.
Pubblicazione: (2026)
di: Yu, Chenxiao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Measuring Moral Inconsistencies in Large Language Models
di: Bonagiri, Vamshi Krishna, et al.
Pubblicazione: (2024) -
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
di: Govil, Priyanshul, et al.
Pubblicazione: (2024) -
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
di: Aneja, Krishak, et al.
Pubblicazione: (2026) -
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
di: Kodali, Prashant, et al.
Pubblicazione: (2024) -
Multilingual Non-Factoid Question Answering with Answer Paragraph Selection
di: Mishra, Ritwik, et al.
Pubblicazione: (2024)