Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jensen, Benjamin, Reynolds, Ian, Atalan, Yasir, Garcia, Michael, Woo, Austin, Chen, Anthony, Howarth, Trevor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
von: Dinh, Tu Anh, et al.
Veröffentlicht: (2024)
von: Dinh, Tu Anh, et al.
Veröffentlicht: (2024)
Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
von: Schuster, Jakob, et al.
Veröffentlicht: (2026)
von: Schuster, Jakob, et al.
Veröffentlicht: (2026)
Cross-lingual Human-Preference Alignment for Neural Machine Translation with Direct Quality Optimization
von: Uhlig, Kaden, et al.
Veröffentlicht: (2024)
von: Uhlig, Kaden, et al.
Veröffentlicht: (2024)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
RMGAP: Benchmarking the Generalization of Reward Models across Diverse Preferences
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
von: Zhou, Yangyang, et al.
Veröffentlicht: (2026)
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
von: Bose, Joy
Veröffentlicht: (2026)
von: Bose, Joy
Veröffentlicht: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
von: Oketunji, Abiodun Finbarrs, et al.
Veröffentlicht: (2023)
Engineering A Large Language Model From Scratch
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
von: Neto, Pedro Barbosa de Carvalho
Veröffentlicht: (2026)
PL-Guard: Benchmarking Language Model Safety for Polish
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
von: Krasnodębska, Aleksandra, et al.
Veröffentlicht: (2025)
A Benchmark of French ASR Systems Based on Error Severity
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
von: Tholly, Antoine, et al.
Veröffentlicht: (2025)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
von: Bouchekif, Abdessalam, et al.
Veröffentlicht: (2026)
LCFO: Long Context and Long Form Output Dataset and Benchmarking
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
Improving Retrospective Language Agents via Joint Policy Gradient Optimization
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
von: Feng, Xueyang, et al.
Veröffentlicht: (2025)
Policy-driven Knowledge Selection and Response Generation for Document-grounded Dialogue
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
von: Ma, Longxuan, et al.
Veröffentlicht: (2024)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
von: Muller, Sacha, et al.
Veröffentlicht: (2024)
SLAP: Stratified Loss-based Pruning for On-Policy Data-Efficient Instruction Tuning
von: Zou, Run, et al.
Veröffentlicht: (2026)
von: Zou, Run, et al.
Veröffentlicht: (2026)
LLM-GLOBE: A Benchmark Evaluating the Cultural Values Embedded in LLM Output
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
von: Karinshak, Elise, et al.
Veröffentlicht: (2024)
BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation
von: The Omnilingual MT Team, et al.
Veröffentlicht: (2025)
von: The Omnilingual MT Team, et al.
Veröffentlicht: (2025)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
von: Gaim, Fitsum, et al.
Veröffentlicht: (2025)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
von: Wang, Xintao, et al.
Veröffentlicht: (2026)
RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors
von: Dugan, Liam, et al.
Veröffentlicht: (2024)
von: Dugan, Liam, et al.
Veröffentlicht: (2024)
LaTIM: Measuring Latent Token-to-Token Interactions in Mamba Models
von: Pitorro, Hugo, et al.
Veröffentlicht: (2025)
von: Pitorro, Hugo, et al.
Veröffentlicht: (2025)
EmoS: A High-Fidelity Multimodal Benchmark for Fine-grained Streaming Emotional Understanding
von: Guo, Pengze, et al.
Veröffentlicht: (2026)
von: Guo, Pengze, et al.
Veröffentlicht: (2026)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
von: Ahmad, Sarfraz, et al.
Veröffentlicht: (2025)
von: Ahmad, Sarfraz, et al.
Veröffentlicht: (2025)
EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generation
von: Yang, Pei, et al.
Veröffentlicht: (2026)
von: Yang, Pei, et al.
Veröffentlicht: (2026)
Surprisingly Fragile: Assessing and Addressing Prompt Instability in Multimodal Foundation Models
von: Stewart, Ian, et al.
Veröffentlicht: (2024)
von: Stewart, Ian, et al.
Veröffentlicht: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
von: Wang, Yuxia, et al.
Veröffentlicht: (2024)
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
von: Lu, Haolang, et al.
Veröffentlicht: (2025)
Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
von: Hinterleitner, Lukas, et al.
Veröffentlicht: (2026)
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
von: Jiang, Yilin, et al.
Veröffentlicht: (2025)
Towards Fundamental Language Models: Does Linguistic Competence Scale with Model Size?
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
von: Collado-Montañez, Jaime, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026) -
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
von: Gupta, Abhay, et al.
Veröffentlicht: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025) -
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
von: Dinh, Tu Anh, et al.
Veröffentlicht: (2024)