DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
Fuente:
arXiv
Saved in:
| Main Authors: | Seth, Agrima, Ahuja, Sanchit, Bali, Kalika, Sitaram, Sunayana |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024)
by: Hada, Rishav, et al.
Published: (2024)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024)
by: Jain, Prachi, et al.
Published: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
by: Ahuja, Sanchit, et al.
Published: (2023)
by: Ahuja, Sanchit, et al.
Published: (2023)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2023)
by: Kumar, Somnath, et al.
Published: (2023)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
by: Mukherjee, Sagnik, et al.
Published: (2024)
by: Mukherjee, Sagnik, et al.
Published: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
by: Hamna, Hamna, et al.
Published: (2025)
by: Hamna, Hamna, et al.
Published: (2025)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
by: Ochieng, Millicent, et al.
Published: (2024)
by: Ochieng, Millicent, et al.
Published: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
by: Hada, Rishav, et al.
Published: (2023)
by: Hada, Rishav, et al.
Published: (2023)
KAHANI: Culturally-Nuanced Visual Storytelling Tool for Non-Western Cultures
by: Hamna, et al.
Published: (2024)
by: Hamna, et al.
Published: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
by: Chitale, Pranjal A., et al.
Published: (2025)
by: Chitale, Pranjal A., et al.
Published: (2025)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
by: Schneider, Florian, et al.
Published: (2024)
by: Schneider, Florian, et al.
Published: (2024)
Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
by: Hada, Rishav, et al.
Published: (2024)
by: Hada, Rishav, et al.
Published: (2024)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models
by: Ahuja, Sanchit, et al.
Published: (2026)
by: Ahuja, Sanchit, et al.
Published: (2026)
Uncovering inequalities in new knowledge learning by large language models across different languages
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
by: Yadav, Hemant, et al.
Published: (2024)
by: Yadav, Hemant, et al.
Published: (2024)
Improving Self Consistency in LLMs through Probabilistic Tokenization
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
The Human Flourishing Geographic Index: A County-Level Dataset for the United States, 2013--2023
by: Iacus, Stefano M., et al.
Published: (2025)
by: Iacus, Stefano M., et al.
Published: (2025)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
by: Khandelwal, Khyati, et al.
Published: (2023)
by: Khandelwal, Khyati, et al.
Published: (2023)
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
by: Gumma, Varun, et al.
Published: (2024)
by: Gumma, Varun, et al.
Published: (2024)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Through the Prism of Culture: Evaluating LLMs' Understanding of Indian Subcultures and Traditions
by: Chhikara, Garima, et al.
Published: (2025)
by: Chhikara, Garima, et al.
Published: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
by: Li, Cheng, et al.
Published: (2024)
by: Li, Cheng, et al.
Published: (2024)
Tackling Social Bias against the Poor: A Dataset and Taxonomy on Aporophobia
by: Curto, Georgina, et al.
Published: (2025)
by: Curto, Georgina, et al.
Published: (2025)
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
by: Ahuja, Sanchit, et al.
Published: (2025)
by: Ahuja, Sanchit, et al.
Published: (2025)
IndRegBias: A Dataset for Studying Indian Regional Biases in English and Code-Mixed Social Media Comments
by: Panda, Debasmita, et al.
Published: (2026)
by: Panda, Debasmita, et al.
Published: (2026)
What's Not on the Plate? Rethinking Food Computing through Indigenous Indian Datasets
by: Gogoi, Pamir, et al.
Published: (2025)
by: Gogoi, Pamir, et al.
Published: (2025)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
by: Ahuja, Sanchit, et al.
Published: (2024)
by: Ahuja, Sanchit, et al.
Published: (2024)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
by: Aggarwal, Divyanshu, et al.
Published: (2024)
by: Aggarwal, Divyanshu, et al.
Published: (2024)
Mapping Violence: Developing an Extensive Framework to Build a Bangla Sectarian Expression Dataset from Social Media Interactions
by: Tasnim, Nazia, et al.
Published: (2024)
by: Tasnim, Nazia, et al.
Published: (2024)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
by: Nadeem, Afrozah, et al.
Published: (2026)
by: Nadeem, Afrozah, et al.
Published: (2026)
Towards High-Fidelity Synthetic Multi-platform Social Media Datasets via Large Language Models
by: Tari, Henry, et al.
Published: (2025)
by: Tari, Henry, et al.
Published: (2025)
Similar Items
-
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025) -
Contamination Report for Multilingual Benchmarks
by: Ahuja, Sanchit, et al.
Published: (2024) -
METAL: Towards Multilingual Meta-Evaluation
by: Hada, Rishav, et al.
Published: (2024) -
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
by: Jain, Prachi, et al.
Published: (2024) -
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)