Salvato in:
| Autori principali: | Xie, Huiyuan, Steffek, Felix, de Faria, Joana Ribeiro, Carter, Christine, Rutherford, Jonathan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2409.08098 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Automatic Information Extraction From Employment Tribunal Judgements Using Large Language Models
di: de Faria, Joana Ribeiro, et al.
Pubblicazione: (2024)
di: de Faria, Joana Ribeiro, et al.
Pubblicazione: (2024)
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
di: Sargeant, Holli, et al.
Pubblicazione: (2024)
di: Sargeant, Holli, et al.
Pubblicazione: (2024)
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
di: Wang, Xing, et al.
Pubblicazione: (2025)
di: Wang, Xing, et al.
Pubblicazione: (2025)
LLM vs. Lawyers: Identifying a Subset of Summary Judgments in a Large UK Case Law Dataset
di: Izzidien, Ahmed, et al.
Pubblicazione: (2024)
di: Izzidien, Ahmed, et al.
Pubblicazione: (2024)
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
di: Kim, Dahyun, et al.
Pubblicazione: (2024)
TACLer: Tailored Curriculum Reinforcement Learning for Efficient Reasoning
di: Lai, Huiyuan, et al.
Pubblicazione: (2026)
di: Lai, Huiyuan, et al.
Pubblicazione: (2026)
The Cambridge Law Corpus: A Dataset for Legal AI Research
di: Östling, Andreas, et al.
Pubblicazione: (2023)
di: Östling, Andreas, et al.
Pubblicazione: (2023)
CaseReportBench: An LLM Benchmark Dataset for Dense Information Extraction in Clinical Case Reports
di: Zhang, Xiao Yu Cindy, et al.
Pubblicazione: (2025)
di: Zhang, Xiao Yu Cindy, et al.
Pubblicazione: (2025)
AyutthayaAlpha: A Thai-Latin Script Transliteration Transformer
di: Lauc, Davor, et al.
Pubblicazione: (2024)
di: Lauc, Davor, et al.
Pubblicazione: (2024)
Evaluating the Quality of Benchmark Datasets for Low-Resource Languages: A Case Study on Turkish
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
di: Cengiz, Ayşe Aysu, et al.
Pubblicazione: (2025)
CliniBench: A Clinical Outcome Prediction Benchmark for Generative and Encoder-Based Language Models
di: Grundmann, Paul, et al.
Pubblicazione: (2025)
di: Grundmann, Paul, et al.
Pubblicazione: (2025)
PhayaThaiBERT: Enhancing a Pretrained Thai Language Model with Unassimilated Loanwords
di: Sriwirote, Panyut, et al.
Pubblicazione: (2023)
di: Sriwirote, Panyut, et al.
Pubblicazione: (2023)
MASSW: A New Dataset and Benchmark Tasks for AI-Assisted Scientific Workflows
di: Zhang, Xingjian, et al.
Pubblicazione: (2024)
di: Zhang, Xingjian, et al.
Pubblicazione: (2024)
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
di: Li, Keyu, et al.
Pubblicazione: (2025)
di: Li, Keyu, et al.
Pubblicazione: (2025)
OpenJAI-v1.0: An Open Thai Large Language Model
di: Trakuekul, Pontakorn, et al.
Pubblicazione: (2025)
di: Trakuekul, Pontakorn, et al.
Pubblicazione: (2025)
Topic-Conversation Relevance (TCR) Dataset and Benchmarks
di: Fan, Yaran, et al.
Pubblicazione: (2024)
di: Fan, Yaran, et al.
Pubblicazione: (2024)
C2RUST-BENCH: A Minimized, Representative Dataset for C-to-Rust Transpilation Evaluation
di: Sirlanci, Melih, et al.
Pubblicazione: (2025)
di: Sirlanci, Melih, et al.
Pubblicazione: (2025)
GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation
di: Sorodoc, Ionut-Teodor, et al.
Pubblicazione: (2025)
di: Sorodoc, Ionut-Teodor, et al.
Pubblicazione: (2025)
Who Benchmarks the Benchmarks? A Case Study of LLM Evaluation in Icelandic
di: Ingimundarson, Finnur Ágúst, et al.
Pubblicazione: (2026)
di: Ingimundarson, Finnur Ágúst, et al.
Pubblicazione: (2026)
Towards Explainability in Legal Outcome Prediction Models
di: Valvoda, Josef, et al.
Pubblicazione: (2024)
di: Valvoda, Josef, et al.
Pubblicazione: (2024)
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
di: Ye, Fangda, et al.
Pubblicazione: (2026)
di: Ye, Fangda, et al.
Pubblicazione: (2026)
Jawaher: A Multidialectal Dataset of Arabic Proverbs for LLM Benchmarking
di: Magdy, Samar M., et al.
Pubblicazione: (2025)
di: Magdy, Samar M., et al.
Pubblicazione: (2025)
SwaQuAD-24: QA Benchmark Dataset in Swahili
di: Kondoro, Alfred Malengo
Pubblicazione: (2024)
di: Kondoro, Alfred Malengo
Pubblicazione: (2024)
BR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law
di: Júnior, Juvenal Domingos, et al.
Pubblicazione: (2025)
di: Júnior, Juvenal Domingos, et al.
Pubblicazione: (2025)
Exposing Assumptions in AI Benchmarks through Cognitive Modelling
di: Rystrøm, Jonathan H., et al.
Pubblicazione: (2024)
di: Rystrøm, Jonathan H., et al.
Pubblicazione: (2024)
TEG-DB: A Comprehensive Dataset and Benchmark of Textual-Edge Graphs
di: Li, Zhuofeng, et al.
Pubblicazione: (2024)
di: Li, Zhuofeng, et al.
Pubblicazione: (2024)
PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning
di: Pham, Hung Manh, et al.
Pubblicazione: (2026)
di: Pham, Hung Manh, et al.
Pubblicazione: (2026)
Emotion and Intent Joint Understanding in Multimodal Conversation: A Benchmarking Dataset
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing
di: Fein, Daniel, et al.
Pubblicazione: (2025)
di: Fein, Daniel, et al.
Pubblicazione: (2025)
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection
di: Harris, Sheetal, et al.
Pubblicazione: (2024)
di: Harris, Sheetal, et al.
Pubblicazione: (2024)
BLUCK: A Benchmark Dataset for Bengali Linguistic Understanding and Cultural Knowledge
di: Kabir, Daeen, et al.
Pubblicazione: (2025)
di: Kabir, Daeen, et al.
Pubblicazione: (2025)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
di: Zakizadeh, Mahdi, et al.
Pubblicazione: (2025)
di: Zakizadeh, Mahdi, et al.
Pubblicazione: (2025)
Breaking the Silence: A Dataset and Benchmark for Bangla Text-to-Gloss Translation
di: Abdullah, Sharif Mohammad, et al.
Pubblicazione: (2025)
di: Abdullah, Sharif Mohammad, et al.
Pubblicazione: (2025)
Order Matters in Hallucination: Reasoning Order as Benchmark and Reflexive Prompting for Large-Language-Models
di: Xie, Zikai
Pubblicazione: (2024)
di: Xie, Zikai
Pubblicazione: (2024)
Can Large Language Models Predict the Outcome of Judicial Decisions?
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2025)
di: Kmainasi, Mohamed Bayan, et al.
Pubblicazione: (2025)
HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
di: Abdaljalil, Samir, et al.
Pubblicazione: (2025)
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs
di: Ma, Yunsheng, et al.
Pubblicazione: (2023)
di: Ma, Yunsheng, et al.
Pubblicazione: (2023)
MCFEND: A Multi-source Benchmark Dataset for Chinese Fake News Detection
di: Li, Yupeng, et al.
Pubblicazione: (2024)
di: Li, Yupeng, et al.
Pubblicazione: (2024)
LLM-Generated Negative News Headlines Dataset: Creation and Benchmarking Against Real Journalism
di: Babalola, Olusola, et al.
Pubblicazione: (2025)
di: Babalola, Olusola, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Automatic Information Extraction From Employment Tribunal Judgements Using Large Language Models
di: de Faria, Joana Ribeiro, et al.
Pubblicazione: (2024) -
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
di: Sargeant, Holli, et al.
Pubblicazione: (2024) -
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
di: Wang, Xing, et al.
Pubblicazione: (2025) -
LLM vs. Lawyers: Identifying a Subset of Summary Judgments in a Large UK Case Law Dataset
di: Izzidien, Ahmed, et al.
Pubblicazione: (2024) -
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)