Can Large Language Models Make Everyone Happy?
Fuente:
arXiv
Salvato in:
| Autori principali: | Naseem, Usman, Kashyap, Gautam Siddharth, Shabbir, Ebad, Ray, Sushant Kumar, Mohammad, Abdullah, Ali, Rafiq |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Do Large Language Models Reflect Demographic Pluralism in Safety?
di: Naseem, Usman, et al.
Pubblicazione: (2026)
di: Naseem, Usman, et al.
Pubblicazione: (2026)
Are Aligned Large Language Models Still Misaligned?
di: Naseem, Usman, et al.
Pubblicazione: (2026)
di: Naseem, Usman, et al.
Pubblicazione: (2026)
Are Large Language Models Economically Viable for Industry Deployment?
di: Mohammad, Abdullah, et al.
Pubblicazione: (2026)
di: Mohammad, Abdullah, et al.
Pubblicazione: (2026)
Can Argus Judge Them All? Comparing VLMs Across Domains
di: Joshi, Harsh, et al.
Pubblicazione: (2025)
di: Joshi, Harsh, et al.
Pubblicazione: (2025)
Truth, Trust, and Trouble: Medical AI on the Edge
di: Azeez, Mohammad Anas, et al.
Pubblicazione: (2025)
di: Azeez, Mohammad Anas, et al.
Pubblicazione: (2025)
LLMs on a Budget? Say HOLA
di: Siddiqui, Zohaib Hasan, et al.
Pubblicazione: (2025)
di: Siddiqui, Zohaib Hasan, et al.
Pubblicazione: (2025)
AlignCultura: Towards Culturally Aligned Large Language Models?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
di: Ray, Sushant Kumar, et al.
Pubblicazione: (2026)
di: Ray, Sushant Kumar, et al.
Pubblicazione: (2026)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2026)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
ChildGuard: A Specialized Dataset for Combatting Child-Targeted Hate Speech
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
di: Naseem, Usman
Pubblicazione: (2026)
di: Naseem, Usman
Pubblicazione: (2026)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
di: Tripathi, Sahil, et al.
Pubblicazione: (2026)
di: Tripathi, Sahil, et al.
Pubblicazione: (2026)
How Can Multimodal Remote Sensing Datasets Transform Classification via SpatialNet-ViT?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
di: Maskey, Utsav, et al.
Pubblicazione: (2025)
From Text to Transformation: A Comprehensive Review of Large Language Models' Versatility
di: Kaur, Pravneet, et al.
Pubblicazione: (2024)
di: Kaur, Pravneet, et al.
Pubblicazione: (2024)
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
di: Afzoon, Saleh, et al.
Pubblicazione: (2024)
di: Afzoon, Saleh, et al.
Pubblicazione: (2024)
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs
di: Gautam, Sushant
Pubblicazione: (2024)
di: Gautam, Sushant
Pubblicazione: (2024)
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content
di: Abishethvarman, Vadivel, et al.
Pubblicazione: (2025)
di: Abishethvarman, Vadivel, et al.
Pubblicazione: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
di: Ren, Kaixuan, et al.
Pubblicazione: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
di: Bhandari, Pranav, et al.
Pubblicazione: (2026)
Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
di: Almutairi, Ali, et al.
Pubblicazione: (2025)
di: Almutairi, Ali, et al.
Pubblicazione: (2025)
Can Reasoning LLMs Enhance Clinical Document Classification?
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
di: Mustafa, Akram, et al.
Pubblicazione: (2025)
Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings
di: Dey, Krishno, et al.
Pubblicazione: (2024)
di: Dey, Krishno, et al.
Pubblicazione: (2024)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
di: Alawwad, Hessa A., et al.
Pubblicazione: (2025)
di: Alawwad, Hessa A., et al.
Pubblicazione: (2025)
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
di: Yuan, Jiahao, et al.
Pubblicazione: (2024)
di: Yuan, Jiahao, et al.
Pubblicazione: (2024)
LLM for Everyone: Representing the Underrepresented in Large Language Models
di: Cahyawijaya, Samuel
Pubblicazione: (2024)
di: Cahyawijaya, Samuel
Pubblicazione: (2024)
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models
di: Veeramani, Hariram, et al.
Pubblicazione: (2024)
di: Veeramani, Hariram, et al.
Pubblicazione: (2024)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
di: Nadeem, Afrozah, et al.
Pubblicazione: (2025)
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
di: Govindarajan, Vijay, et al.
Pubblicazione: (2025)
di: Govindarajan, Vijay, et al.
Pubblicazione: (2025)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
di: Afzoon, Saleh, et al.
Pubblicazione: (2026)
di: Afzoon, Saleh, et al.
Pubblicazione: (2026)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
di: Maskey, Utsav, et al.
Pubblicazione: (2026)
di: Maskey, Utsav, et al.
Pubblicazione: (2026)
Is Safety Standard Same for Everyone? User-Specific Safety Evaluation of Large Language Models
di: In, Yeonjun, et al.
Pubblicazione: (2025)
di: In, Yeonjun, et al.
Pubblicazione: (2025)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
di: Ren, Juan, et al.
Pubblicazione: (2025)
di: Ren, Juan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Do Large Language Models Reflect Demographic Pluralism in Safety?
di: Naseem, Usman, et al.
Pubblicazione: (2026) -
Are Aligned Large Language Models Still Misaligned?
di: Naseem, Usman, et al.
Pubblicazione: (2026) -
Are Large Language Models Economically Viable for Industry Deployment?
di: Mohammad, Abdullah, et al.
Pubblicazione: (2026) -
Can Argus Judge Them All? Comparing VLMs Across Domains
di: Joshi, Harsh, et al.
Pubblicazione: (2025) -
Truth, Trust, and Trouble: Medical AI on the Edge
di: Azeez, Mohammad Anas, et al.
Pubblicazione: (2025)