Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sahoo, Subramanyam, Jain, Vinija, Chadha, Aman, Chaudhary, Divya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Catch Me If You Can: How Smaller Reasoning Models Pretend to Reason with Mathematical Fidelity
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Policy myopia as a mechanism of gradual disempowerment in Post-AGI governance, Circa 2049
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
von: Mahalingam, Aakash, et al.
Veröffentlicht: (2024)
von: Mahalingam, Aakash, et al.
Veröffentlicht: (2024)
From Prejudice to Parity: A New Approach to Debiasing Large Language Model Word Embeddings
von: Rakshit, Aishik, et al.
Veröffentlicht: (2024)
von: Rakshit, Aishik, et al.
Veröffentlicht: (2024)
Moral Sensitivity in LLMs: A Tiered Evaluation of Contextual Bias via Behavioral Profiling and Mechanistic Interpretability
von: Aggarwal, Yash, et al.
Veröffentlicht: (2026)
von: Aggarwal, Yash, et al.
Veröffentlicht: (2026)
The Controllability Trap: A Governance Framework for Military AI Agents
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
The Last Vote: A Multi-Stakeholder Framework for Language Model Governance
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Boardwalk Empire: How Generative AI is Revolutionizing Economic Paradigms
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2024)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2024)
A Comprehensive Survey of Hallucination in Large Language, Image, Video and Audio Foundation Models
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
von: Sahoo, Pranab, et al.
Veröffentlicht: (2024)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
von: Kumar, Tanay, et al.
Veröffentlicht: (2026)
von: Kumar, Tanay, et al.
Veröffentlicht: (2026)
AI based Content Creation and Product Recommendation Applications in E-commerce: An Ethical overview
von: Jain, Aditi Madhusudan, et al.
Veröffentlicht: (2025)
von: Jain, Aditi Madhusudan, et al.
Veröffentlicht: (2025)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
Decoding the Diversity: A Review of the Indic AI Research Landscape
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
von: Haider, Batool, et al.
Veröffentlicht: (2025)
von: Haider, Batool, et al.
Veröffentlicht: (2025)
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
Dial-a-ride problem with modular platooning and en-route transfers
von: Fu, Zhexi, et al.
Veröffentlicht: (2022)
von: Fu, Zhexi, et al.
Veröffentlicht: (2022)
Analysis of Indian Agricultural Ecosystem using Knowledge-based Tantra Framework
von: Prabhu, Shreekanth M, et al.
Veröffentlicht: (2021)
von: Prabhu, Shreekanth M, et al.
Veröffentlicht: (2021)
LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
von: Sidibomma, Rushendra, et al.
Veröffentlicht: (2024)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
MAAT: Multi-phase Adapter-Aware Targeted Unlearning
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
von: Yagnik, Suryash, et al.
Veröffentlicht: (2026)
Identity, Crimes, and Law Enforcement in the Metaverse
von: Qin, Hua Xuan, et al.
Veröffentlicht: (2022)
von: Qin, Hua Xuan, et al.
Veröffentlicht: (2022)
Parameter Efficient Fine Tuning: A Comprehensive Analysis Across Applications
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
von: Balne, Charith Chandra Sai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026) -
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)