SDUs DAISY: A Benchmark for Danish Culture
Fuente:
arXiv
Salvato in:
| Autori principali: | Nielsen, Jacob, Beltoft, Stine L., Schneider-Kamp, Peter, Poech, Lukas Galke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
di: Brach, William, et al.
Pubblicazione: (2026)
di: Brach, William, et al.
Pubblicazione: (2026)
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
di: Beltoft, Stine Lyngsø, et al.
Pubblicazione: (2026)
di: Beltoft, Stine Lyngsø, et al.
Pubblicazione: (2026)
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
di: Beltoft, Stine, et al.
Pubblicazione: (2025)
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
di: Torrielli, Federico, et al.
Pubblicazione: (2026)
di: Torrielli, Federico, et al.
Pubblicazione: (2026)
Chain of Summaries: Summarization Through Iterative Questioning
di: Brach, William, et al.
Pubblicazione: (2025)
di: Brach, William, et al.
Pubblicazione: (2025)
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
di: From, Mogens Henrik, et al.
Pubblicazione: (2025)
di: From, Mogens Henrik, et al.
Pubblicazione: (2025)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
di: Barmina, Gianluca, et al.
Pubblicazione: (2025)
di: Barmina, Gianluca, et al.
Pubblicazione: (2025)
When are 1.58 bits enough? A Bottom-up Exploration of BitNet Quantization
di: Nielsen, Jacob, et al.
Pubblicazione: (2024)
di: Nielsen, Jacob, et al.
Pubblicazione: (2024)
Continual Quantization-Aware Pre-Training: When to transition from 16-bit to 1.58-bit pre-training for BitNet language models?
di: Nielsen, Jacob, et al.
Pubblicazione: (2025)
di: Nielsen, Jacob, et al.
Pubblicazione: (2025)
SommBench: Assessing Sommelier Expertise of Language Models
di: Brach, William, et al.
Pubblicazione: (2026)
di: Brach, William, et al.
Pubblicazione: (2026)
Isolating Culture Neurons in Multilingual Large Language Models
di: Namazifard, Danial, et al.
Pubblicazione: (2025)
di: Namazifard, Danial, et al.
Pubblicazione: (2025)
Training Language Models to Use Prolog as a Tool
di: Mellgren, Niklas, et al.
Pubblicazione: (2025)
di: Mellgren, Niklas, et al.
Pubblicazione: (2025)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
di: Nielsen, Jacob, et al.
Pubblicazione: (2024)
di: Nielsen, Jacob, et al.
Pubblicazione: (2024)
Dynaword: From One-shot to Continuously Developed Datasets
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2025)
di: Enevoldsen, Kenneth, et al.
Pubblicazione: (2025)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
The Differential Meaning of Models: A Framework for Analyzing the Structural Consequences of Semantic Modeling Decisions
di: Stine, Zachary K., et al.
Pubblicazione: (2025)
di: Stine, Zachary K., et al.
Pubblicazione: (2025)
Natural Language Processing for Electronic Health Records in Scandinavian Languages: Norwegian, Swedish, and Danish
di: Woldaregay, Ashenafi Zebene, et al.
Pubblicazione: (2025)
di: Woldaregay, Ashenafi Zebene, et al.
Pubblicazione: (2025)
FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models
di: Pirchert, Annemette Brok, et al.
Pubblicazione: (2026)
di: Pirchert, Annemette Brok, et al.
Pubblicazione: (2026)
ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
di: Ahmed, Md Shamim, et al.
Pubblicazione: (2026)
From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation
di: Han, Bangju, et al.
Pubblicazione: (2026)
di: Han, Bangju, et al.
Pubblicazione: (2026)
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
di: Choi, Dasol, et al.
Pubblicazione: (2026)
di: Choi, Dasol, et al.
Pubblicazione: (2026)
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia
di: Ayash, Lama, et al.
Pubblicazione: (2025)
di: Ayash, Lama, et al.
Pubblicazione: (2025)
TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
di: Akinode, Victor, et al.
Pubblicazione: (2026)
di: Akinode, Victor, et al.
Pubblicazione: (2026)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
di: Kabir, Daeen, et al.
Pubblicazione: (2025)
di: Kabir, Daeen, et al.
Pubblicazione: (2025)
Learning from Sufficient Rationales: Analysing the Relationship Between Explanation Faithfulness and Token-level Regularisation Strategies
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement
di: Kamp, Jonathan, et al.
Pubblicazione: (2024)
di: Kamp, Jonathan, et al.
Pubblicazione: (2024)
Explanation Bias is a Product: Revealing the Hidden Lexical and Position Preferences in Post-Hoc Feature Attribution
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
di: Kamp, Jonathan, et al.
Pubblicazione: (2025)
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
di: Jin, Jiho, et al.
Pubblicazione: (2026)
di: Jin, Jiho, et al.
Pubblicazione: (2026)
CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
di: Chiu, Yu Ying, et al.
Pubblicazione: (2024)
CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks
di: Lin, Peiqin, et al.
Pubblicazione: (2026)
di: Lin, Peiqin, et al.
Pubblicazione: (2026)
CamelEval: Advancing Culturally Aligned Arabic Language Models and Benchmarks
di: Qian, Zhaozhi, et al.
Pubblicazione: (2024)
di: Qian, Zhaozhi, et al.
Pubblicazione: (2024)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
di: Chang, Chen-Chi, et al.
Pubblicazione: (2024)
di: Chang, Chen-Chi, et al.
Pubblicazione: (2024)
MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks
di: Veuthey, Jaime Raldua, et al.
Pubblicazione: (2025)
di: Veuthey, Jaime Raldua, et al.
Pubblicazione: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2026)
IndoBias: A Dual Track Culturally Grounded Benchmark for LLMs Bias Evaluation in Indonesian Languages
di: Hanif, Ikhlasul Akmal, et al.
Pubblicazione: (2026)
di: Hanif, Ikhlasul Akmal, et al.
Pubblicazione: (2026)
LLMs as annotators of credibility assessment in Danish asylum decisions: evaluating classification performance and errors beyond aggregated metrics
di: Humblot-Renaux, Galadrielle, et al.
Pubblicazione: (2026)
di: Humblot-Renaux, Galadrielle, et al.
Pubblicazione: (2026)
Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports
di: Singh, Punit Kumar, et al.
Pubblicazione: (2025)
di: Singh, Punit Kumar, et al.
Pubblicazione: (2025)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
IRLBench: A Multi-modal, Culturally Grounded, Parallel Irish-English Benchmark for Open-Ended LLM Reasoning Evaluation
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
di: Tran, Khanh-Tung, et al.
Pubblicazione: (2025)
LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
di: Gao, Zihan, et al.
Pubblicazione: (2025)
di: Gao, Zihan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
di: Brach, William, et al.
Pubblicazione: (2026) -
Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
di: Beltoft, Stine Lyngsø, et al.
Pubblicazione: (2026) -
Not Everything That Counts Can Be Counted: A Case for Safe Qualitative AI
di: Beltoft, Stine, et al.
Pubblicazione: (2025) -
Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
di: Torrielli, Federico, et al.
Pubblicazione: (2026) -
Chain of Summaries: Summarization Through Iterative Questioning
di: Brach, William, et al.
Pubblicazione: (2025)