Gespeichert in:
| Hauptverfasser: | Das, Amit, Rahgouy, Mostafa, Feng, Dongji, Zhang, Zheng, Bhattacharya, Tathagata, Raychawdhary, Nilanjana, Jamshidi, Fatemeh, Jain, Vinija, Chadha, Aman, Sandage, Mary, Pope, Lauramarie, Dozier, Gerry, Seals, Cheryl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.02472 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024)
von: Das, Amit, et al.
Veröffentlicht: (2024)
Investigating Hallucination in Conversations for Low Resource Languages
von: Das, Amit, et al.
Veröffentlicht: (2025)
von: Das, Amit, et al.
Veröffentlicht: (2025)
Towards Effective Authorship Attribution: Integrating Class-Incremental Learning
von: Rahgouy, Mostafa, et al.
Veröffentlicht: (2024)
von: Rahgouy, Mostafa, et al.
Veröffentlicht: (2024)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026)
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
Are Small Language Models Ready to Compete with Large Language Models for Practical Applications?
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
von: Sinha, Neelabh, et al.
Veröffentlicht: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Reasoning or Rhetoric? An Empirical Analysis of Moral Reasoning Explanations in Large Language Models
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
von: Kasat, Aryan, et al.
Veröffentlicht: (2026)
Dial E for Ethical Enforcement: institutional VETO power as a governance primitive
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Born With a Silver Spoon? Investigating Socioeconomic Bias in Large Language Models
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
von: Singh, Smriti, et al.
Veröffentlicht: (2024)
AlignGuard-LoRA: Alignment-Preserving Fine-Tuning via Fisher-Guided Decomposition and Riemannian-Geodesic Collision Regularization
von: Das, Amitava, et al.
Veröffentlicht: (2025)
von: Das, Amitava, et al.
Veröffentlicht: (2025)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2025)
Exploring the Impact of Large Language Models on Recommender Systems: An Extensive Review
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
von: Vats, Arpita, et al.
Veröffentlicht: (2024)
The Civilising Offensive
Veröffentlicht: (2022)
Veröffentlicht: (2022)
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems
von: Liu, Zishuo, et al.
Veröffentlicht: (2025)
von: Liu, Zishuo, et al.
Veröffentlicht: (2025)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
Offensive Robot Cybersecurity
von: Mayoral-Vilches, Víctor
Veröffentlicht: (2025)
von: Mayoral-Vilches, Víctor
Veröffentlicht: (2025)
SOMALIA: Puntland Offensive
Veröffentlicht: (2025)
Veröffentlicht: (2025)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
von: Das, Arion, et al.
Veröffentlicht: (2026)
von: Das, Arion, et al.
Veröffentlicht: (2026)
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
von: Ghosh, Akash, et al.
Veröffentlicht: (2024)
D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
von: Raina, Samarth, et al.
Veröffentlicht: (2025)
AlignMerge - Alignment-Preserving Large Language Model Merging via Fisher-Guided Geometric Constraints
von: Roy, Aniruddha, et al.
Veröffentlicht: (2025)
von: Roy, Aniruddha, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Accelerated Generation Techniques in Large Language Models
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
von: Khoshnoodi, Mahsa, et al.
Veröffentlicht: (2024)
Neural FOXP2 -- Language Specific Neuron Steering for Targeted Language Improvement in LLMs
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
von: Saha, Anusa, et al.
Veröffentlicht: (2026)
ECLIPTICA -- A Framework for Switchable LLM Alignment via CITA - Contrastive Instruction-Tuned Alignment
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
von: Wanaskar, Kapil, et al.
Veröffentlicht: (2026)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
von: Kumar, Tanay, et al.
Veröffentlicht: (2026)
von: Kumar, Tanay, et al.
Veröffentlicht: (2026)
Multilingual State Space Models for Structured Question Answering in Indic Languages
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
MOD-X: A Modular Open Decentralized eXchange Framework proposal for Heterogeneous Interoperable Artificial Intelligence Agents
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
Decoding the Diversity: A Review of the Indic AI Research Landscape
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
von: KJ, Sankalp, et al.
Veröffentlicht: (2024)
Refining Text-to-Image Generation: Towards Accurate Training-Free Glyph-Enhanced Image Generation
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
von: Lakhanpal, Sanyam, et al.
Veröffentlicht: (2024)
Offensive Lineup Analysis in Basketball with Clustering Players Based on Shooting Style and Offensive Role
von: Yamada, Kazuhiro, et al.
Veröffentlicht: (2024)
von: Yamada, Kazuhiro, et al.
Veröffentlicht: (2024)
What is the AGI in Offensive Security ?
von: Cho, Youngwoong
Veröffentlicht: (2026)
von: Cho, Youngwoong
Veröffentlicht: (2026)
Ähnliche Einträge
-
Investigating Annotator Bias in Large Language Models for Hate Speech Detection
von: Das, Amit, et al.
Veröffentlicht: (2024) -
Investigating Hallucination in Conversations for Low Resource Languages
von: Das, Amit, et al.
Veröffentlicht: (2025) -
Towards Effective Authorship Attribution: Integrating Class-Incremental Learning
von: Rahgouy, Mostafa, et al.
Veröffentlicht: (2024) -
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
von: Krishnappa, Pushwitha, et al.
Veröffentlicht: (2026) -
TRACEALIGN -- Tracing the Drift: Attributing Alignment Failures to Training-Time Belief Sources in LLMs
von: Das, Amitava, et al.
Veröffentlicht: (2025)