ACCESS DENIED INC: The First Benchmark Environment for Sensitivity Awareness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fazlija, Dren, Orlov, Arkadij, Sikdar, Sandipan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Sensitivity-Aware Language Models
von: Fazlija, Dren, et al.
Veröffentlicht: (2026)
von: Fazlija, Dren, et al.
Veröffentlicht: (2026)
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2026)
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2026)
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025)
SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples
von: Fazlija, Dren, et al.
Veröffentlicht: (2025)
von: Fazlija, Dren, et al.
Veröffentlicht: (2025)
How Real Is Real? A Human Evaluation Framework for Unrestricted Adversarial Examples
von: Fazlija, Dren, et al.
Veröffentlicht: (2024)
von: Fazlija, Dren, et al.
Veröffentlicht: (2024)
TIGQA:An Expert Annotated Question Answering Dataset in Tigrinya
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2024)
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2024)
ACCESS DENIED? : CULTURAL CAPITAL AND DIGITAL ACCESS
von: Shraddha Kumbhojkar
Veröffentlicht: (2019)
von: Shraddha Kumbhojkar
Veröffentlicht: (2019)
The Impact of Synthetic Data on Object Detection Model Performance: A Comparative Analysis with Real-World Data
von: Bay, Muammer, et al.
Veröffentlicht: (2025)
von: Bay, Muammer, et al.
Veröffentlicht: (2025)
Operating Within the Operational Design Domain: Zero-Shot Perception with Vision-Language Models
von: Ünal, Berkehan, et al.
Veröffentlicht: (2026)
von: Ünal, Berkehan, et al.
Veröffentlicht: (2026)
Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
von: Majhi, Sandipan, et al.
Veröffentlicht: (2025)
Litmus (Re)Agent: A Benchmark and Agentic System for Predictive Evaluation of Multilingual Models
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
C-QUERI: Congressional Questions, Exchanges, and Responses in Institutions Dataset
von: Rudra, Manjari, et al.
Veröffentlicht: (2025)
von: Rudra, Manjari, et al.
Veröffentlicht: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
von: Kumar, Shanu, et al.
Veröffentlicht: (2024)
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
von: Seo, Gyuhyeon, et al.
Veröffentlicht: (2025)
von: Seo, Gyuhyeon, et al.
Veröffentlicht: (2025)
How Sensitive Are Safety Benchmarks to Judge Configuration Choices?
von: Zhang, Xinran
Veröffentlicht: (2026)
von: Zhang, Xinran
Veröffentlicht: (2026)
LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Models
von: Sikdar, Prateek Kumar
Veröffentlicht: (2026)
von: Sikdar, Prateek Kumar
Veröffentlicht: (2026)
Benchmarking Machine Translation with Cultural Awareness
von: Yao, Binwei, et al.
Veröffentlicht: (2023)
von: Yao, Binwei, et al.
Veröffentlicht: (2023)
VM14K: First Vietnamese Medical Benchmark
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
von: Nguyen, Thong, et al.
Veröffentlicht: (2025)
Properties of Group Fairness Metrics for Rankings
von: Schumacher, Tobias, et al.
Veröffentlicht: (2022)
von: Schumacher, Tobias, et al.
Veröffentlicht: (2022)
Exploring Disparity-Accuracy Trade-offs in Face Recognition Systems: The Role of Datasets, Architectures, and Loss Functions
von: Jaiswal, Siddharth D, et al.
Veröffentlicht: (2025)
von: Jaiswal, Siddharth D, et al.
Veröffentlicht: (2025)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
von: Bellibatlu, Rohith Reddy, et al.
Veröffentlicht: (2026)
MMTutorBench: The First Multimodal Benchmark for AI Math Tutoring
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
von: Yang, Tengchao, et al.
Veröffentlicht: (2025)
Culturally-Aware Conversations: A Framework & Benchmark for LLMs
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
von: Havaldar, Shreya, et al.
Veröffentlicht: (2025)
Are Arabic Benchmarks Reliable? QIMMA's Quality-First Approach to LLM Evaluation
von: AlQadi, Leen, et al.
Veröffentlicht: (2026)
von: AlQadi, Leen, et al.
Veröffentlicht: (2026)
VISLA Benchmark: Evaluating Embedding Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Factual and Edit-Sensitive Graph-to-Sequence Generation via Graph-Aware Adaptive Noising
von: Shahane, Aditya Hemant, et al.
Veröffentlicht: (2026)
von: Shahane, Aditya Hemant, et al.
Veröffentlicht: (2026)
Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness
von: Li, Bryan, et al.
Veröffentlicht: (2024)
von: Li, Bryan, et al.
Veröffentlicht: (2024)
CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation
von: Villa-Cueva, Emilio, et al.
Veröffentlicht: (2025)
von: Villa-Cueva, Emilio, et al.
Veröffentlicht: (2025)
Benchmarking Prompt Sensitivity in Large Language Models
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
von: Razavi, Amirhossein, et al.
Veröffentlicht: (2025)
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
von: Morabito, Robert, et al.
Veröffentlicht: (2024)
I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench
von: Li, Yuan, et al.
Veröffentlicht: (2024)
von: Li, Yuan, et al.
Veröffentlicht: (2024)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
Rethinking Tokenization for Rich Morphology: The Dominance of Unigram over BPE and Morphological Alignment
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
von: Vemula, Saketh Reddy, et al.
Veröffentlicht: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
von: Liu, Zefang, et al.
Veröffentlicht: (2025)
Culture-Aware Machine Translation in Large Language Models: Benchmarking and Investigation
von: Yuan, Zekun, et al.
Veröffentlicht: (2026)
von: Yuan, Zekun, et al.
Veröffentlicht: (2026)
Environment-Aware Code Generation: How far are We?
von: Wu, Tongtong, et al.
Veröffentlicht: (2026)
von: Wu, Tongtong, et al.
Veröffentlicht: (2026)
PalmX 2025: The First Shared Task on Benchmarking LLMs on Arabic and Islamic Culture
von: Alwajih, Fakhraddin, et al.
Veröffentlicht: (2025)
von: Alwajih, Fakhraddin, et al.
Veröffentlicht: (2025)
Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework
von: Joarder, Shourov, et al.
Veröffentlicht: (2026)
von: Joarder, Shourov, et al.
Veröffentlicht: (2026)
RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
von: Xu, Zukang, et al.
Veröffentlicht: (2025)
UAL-Bench: The First Comprehensive Unusual Activity Localization Benchmark
von: Abdullah, Hasnat Md, et al.
Veröffentlicht: (2024)
von: Abdullah, Hasnat Md, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Towards Sensitivity-Aware Language Models
von: Fazlija, Dren, et al.
Veröffentlicht: (2026) -
LGSE: Lexically Grounded Subword Embedding Initialization for Low-Resource Language Adaptation
von: Teklehaymanot, Hailay, et al.
Veröffentlicht: (2026) -
MoVoC: Morphology-Aware Subword Construction for Geez Script Languages
von: Teklehaymanot, Hailay Kidu, et al.
Veröffentlicht: (2025) -
SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples
von: Fazlija, Dren, et al.
Veröffentlicht: (2025) -
How Real Is Real? A Human Evaluation Framework for Unrestricted Adversarial Examples
von: Fazlija, Dren, et al.
Veröffentlicht: (2024)