InSaAF: Incorporating Safety through Accuracy and Fairness | Are LLMs ready for the Indian Legal Domain?
Fuente:
arXiv
Saved in:
| Main Authors: | Tripathi, Yogesh, Donakanti, Raghav, Girhepuje, Sahil, Kavathekar, Ishan, Vedula, Bhaskara Hanuma, Krishnan, Gokul S, Goyal, Shreya, Goel, Anmol, Ravindran, Balaraman, Kumaraguru, Ponnurangam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Models Trained on Indian Legal Data Fair?
by: Girhepuje, Sahil, et al.
Published: (2023)
by: Girhepuje, Sahil, et al.
Published: (2023)
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
by: Kavathekar, Ishan, et al.
Published: (2025)
by: Kavathekar, Ishan, et al.
Published: (2025)
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
MACA: A Framework for Distilling Trustworthy LLMs into Efficient Retrievers
by: Gudipudi, Satya Swaroop, et al.
Published: (2026)
by: Gudipudi, Satya Swaroop, et al.
Published: (2026)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
by: Kavathekar, Ishan, et al.
Published: (2025)
by: Kavathekar, Ishan, et al.
Published: (2025)
LExT: Towards Evaluating Trustworthiness of Natural Language Explanations
by: Shailya, Krithi, et al.
Published: (2025)
by: Shailya, Krithi, et al.
Published: (2025)
Counter Turing Test ($CT^2$): Investigating AI-Generated Text Detection for Hindi -- Ranking LLMs based on Hindi AI Detectability Index ($ADI_{hi}$)
by: Kavathekar, Ishan, et al.
Published: (2024)
by: Kavathekar, Ishan, et al.
Published: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
Where Should I Study? Biased Language Models Decide! Evaluating Fairness in LMs for Academic Recommendations
by: Shailya, Krithi, et al.
Published: (2025)
by: Shailya, Krithi, et al.
Published: (2025)
IndiCASA: A Dataset and Bias Evaluation Framework in LLMs Using Contrastive Embedding Similarity in the Indian Context
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Participatory Approaches in AI Development and Governance: Case Studies
by: Parthasarathy, Ambreesh, et al.
Published: (2024)
by: Parthasarathy, Ambreesh, et al.
Published: (2024)
mFARM: Towards Multi-Faceted Fairness Assessment based on HARMs in Clinical Decision Support
by: Adappanavar, Shreyash, et al.
Published: (2025)
by: Adappanavar, Shreyash, et al.
Published: (2025)
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Participatory Approaches in AI Development and Governance: A Principled Approach
by: Parthasarathy, Ambreesh, et al.
Published: (2024)
by: Parthasarathy, Ambreesh, et al.
Published: (2024)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
by: P, Vedanta S, et al.
Published: (2026)
by: P, Vedanta S, et al.
Published: (2026)
A Survey on Offensive AI Within Cybersecurity
by: Girhepuje, Sahil, et al.
Published: (2024)
by: Girhepuje, Sahil, et al.
Published: (2024)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
by: Mishra, Debangan, et al.
Published: (2025)
by: Mishra, Debangan, et al.
Published: (2025)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
Unifying Model-Free Efficiency and Model-Based Representations via Latent Dynamics
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
by: Acharjee, Jashaswimalya, et al.
Published: (2026)
Learning Interpretable Models Using Uncertainty Oracles
by: Ghose, Abhishek, et al.
Published: (2019)
by: Ghose, Abhishek, et al.
Published: (2019)
Enhancing AI Safety Through the Fusion of Low Rank Adapters
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
Reimagining Self-Adaptation in the Age of Large Language Models
by: Donakanti, Raghav, et al.
Published: (2024)
by: Donakanti, Raghav, et al.
Published: (2024)
Personal Narratives Empower Politically Disinclined Individuals to Engage in Political Discussions
by: Chebrolu, Tejasvi, et al.
Published: (2025)
by: Chebrolu, Tejasvi, et al.
Published: (2025)
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025)
by: Sinha, Shiven, et al.
Published: (2025)
Television Discourse Decoded: Comprehensive Multimodal Analytics at Scale
by: Agarwal, Anmol, et al.
Published: (2024)
by: Agarwal, Anmol, et al.
Published: (2024)
X-posing Free Speech: Examining the Impact of Moderation Relaxation on Online Social Networks
by: Arun, Arvindh, et al.
Published: (2024)
by: Arun, Arvindh, et al.
Published: (2024)
Analyzing Patterns and Influence of Advertising in Print Newspapers
by: Vardhan, N Harsha, et al.
Published: (2025)
by: Vardhan, N Harsha, et al.
Published: (2025)
Long-context Non-factoid Question Answering in Indic Languages
by: Mishra, Ritwik, et al.
Published: (2025)
by: Mishra, Ritwik, et al.
Published: (2025)
HLDC: Hindi Legal Documents Corpus
by: Kapoor, Arnav, et al.
Published: (2022)
by: Kapoor, Arnav, et al.
Published: (2022)
Nature vs. Supra-Nature: How Montessori Children Conceptualize the Real and Imaginary Through Visual Art - DataSet
by: Donakanti, Vinay Shyam
Published: (2026)
by: Donakanti, Vinay Shyam
Published: (2026)
SWAN: Sparse Winnowed Attention for Reduced Inference Memory via Decompression-Free KV-Cache Compression
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
by: S, Santhosh G, et al.
Published: (2025)
by: S, Santhosh G, et al.
Published: (2025)
Ketto and the Science of Giving: A Data-Driven Investigation of Crowdfunding for India
by: Chandra, Karuna, et al.
Published: (2025)
by: Chandra, Karuna, et al.
Published: (2025)
MetaGMT: Improving Actionable Interpretability of Graph Multilinear Networks via Meta-Learning Filtration
by: Bhattacharya, Rishabh, et al.
Published: (2025)
by: Bhattacharya, Rishabh, et al.
Published: (2025)
Higher Order Structures For Graph Explanations
by: Sinha, Akshit, et al.
Published: (2024)
by: Sinha, Akshit, et al.
Published: (2024)
Similar Items
-
Are Models Trained on Indian Legal Data Fair?
by: Girhepuje, Sahil, et al.
Published: (2023) -
Small Models, Big Tasks: An Exploratory Empirical Study on Small Language Models for Function Calling
by: Kavathekar, Ishan, et al.
Published: (2025) -
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
by: Vedula, Bhaskara Hanuma, et al.
Published: (2026) -
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025) -
MACA: A Framework for Distilling Trustworthy LLMs into Efficient Retrievers
by: Gudipudi, Satya Swaroop, et al.
Published: (2026)