Gespeichert in:
| Hauptverfasser: | Abishethvarman, Vadivel, Chandna, Bhavik, Jalan, Pratik, Naseem, Usman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.00973 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025)
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
von: Naseem, Usman
Veröffentlicht: (2026)
von: Naseem, Usman
Veröffentlicht: (2026)
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
von: Afzoon, Saleh, et al.
Veröffentlicht: (2024)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2024)
From Native Memes to Global Moderation: Cross-Cultural Evaluation of Vision-Language Models for Hateful Meme Detection
von: Wang, Mo, et al.
Veröffentlicht: (2026)
von: Wang, Mo, et al.
Veröffentlicht: (2026)
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
Do Large Language Models Reflect Demographic Pluralism in Safety?
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism
von: Chandna, Bhavik, et al.
Veröffentlicht: (2026)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2026)
AlignCultura: Towards Culturally Aligned Large Language Models?
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
Should LLM Safety Be More Than Refusing Harmful Instructions?
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
von: Maskey, Utsav, et al.
Veröffentlicht: (2025)
CogMem: A Cognitive Memory Architecture for Sustained Multi-Turn Reasoning in Large Language Models
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
von: Zhang, Yiran, et al.
Veröffentlicht: (2025)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2026)
Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings
von: Dey, Krishno, et al.
Veröffentlicht: (2024)
von: Dey, Krishno, et al.
Veröffentlicht: (2024)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2026)
Evaluating Personality Traits in Large Language Models: Insights from Psychological Questionnaires
von: Bhandari, Pranav, et al.
Veröffentlicht: (2025)
von: Bhandari, Pranav, et al.
Veröffentlicht: (2025)
Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
von: Alawwad, Hessa A., et al.
Veröffentlicht: (2025)
von: Alawwad, Hessa A., et al.
Veröffentlicht: (2025)
A Counterfactual Explanation Framework for Retrieval Models
von: Chandna, Bhavik, et al.
Veröffentlicht: (2024)
von: Chandna, Bhavik, et al.
Veröffentlicht: (2024)
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Reversal of Thought: Enhancing Large Language Models with Preference-Guided Reverse Reasoning Warm-up
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2024)
Can Large Language Models Make Everyone Happy?
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
Framing Political Bias in Multilingual LLMs Across Pakistani Languages
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Enhancing ESG Impact Type Identification through Early Fusion and Multilingual Models
von: Veeramani, Hariram, et al.
Veröffentlicht: (2024)
von: Veeramani, Hariram, et al.
Veröffentlicht: (2024)
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare
von: Shetty, Anudeex, et al.
Veröffentlicht: (2025)
von: Shetty, Anudeex, et al.
Veröffentlicht: (2025)
Are Aligned Large Language Models Still Misaligned?
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
von: Naseem, Usman, et al.
Veröffentlicht: (2026)
Evaluating Hierarchical Clinical Document Classification Using Reasoning-Based LLMs
von: Mustafa, Akram, et al.
Veröffentlicht: (2025)
von: Mustafa, Akram, et al.
Veröffentlicht: (2025)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
von: Hasan, Md. Arid, et al.
Veröffentlicht: (2024)
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
von: Maskey, Utsav, et al.
Veröffentlicht: (2026)
Are Large Language Models Economically Viable for Industry Deployment?
von: Mohammad, Abdullah, et al.
Veröffentlicht: (2026)
von: Mohammad, Abdullah, et al.
Veröffentlicht: (2026)
Bias Beyond Borders: Political Ideology Evaluation and Steering in Multilingual LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2026)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2026)
ChineseSafe: A Chinese Benchmark for Evaluating Safety in Large Language Models
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
von: Zhang, Hengxiang, et al.
Veröffentlicht: (2024)
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
von: Ren, Juan, et al.
Veröffentlicht: (2025)
von: Ren, Juan, et al.
Veröffentlicht: (2025)
PersoPilot: An Adaptive AI-Copilot for Transparent Contextualized Persona Classification and Personalized Response Generation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
von: Yadav, Sumit, et al.
Veröffentlicht: (2025)
PersoDPO: Scalable Preference Optimization for Instruction-Adherent, Persona-Grounded Dialogue via Multi-LLM Evaluation
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
von: Afzoon, Saleh, et al.
Veröffentlicht: (2026)
Steering Towards Fairness: Mitigating Political Bias in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
von: Kashyap, Gautam Siddharth, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ExtremeAIGC: Benchmarking LMM Vulnerability to AI-Generated Extremist Content
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025) -
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
von: Chandna, Bhavik, et al.
Veröffentlicht: (2025) -
Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
von: Naseem, Usman
Veröffentlicht: (2026) -
Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities
von: Maskey, Utsav, et al.
Veröffentlicht: (2025) -
PersoBench: Benchmarking Personalized Response Generation in Large Language Models
von: Afzoon, Saleh, et al.
Veröffentlicht: (2024)