Guardado en:
| Autor principal: | Gupta, Kshitij |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2502.07747 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
por: Wang, Chonghua, et al.
Publicado: (2024)
por: Wang, Chonghua, et al.
Publicado: (2024)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
por: Pandey, Atharva, et al.
Publicado: (2025)
por: Pandey, Atharva, et al.
Publicado: (2025)
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
por: Mishra, Kshitij, et al.
Publicado: (2026)
por: Mishra, Kshitij, et al.
Publicado: (2026)
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
por: Garg, Raghav, et al.
Publicado: (2025)
por: Garg, Raghav, et al.
Publicado: (2025)
Suvach -- Generated Hindi QA benchmark
por: Narayanan, Vaishak, et al.
Publicado: (2024)
por: Narayanan, Vaishak, et al.
Publicado: (2024)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
por: Xu, Xinnuo, et al.
Publicado: (2025)
por: Xu, Xinnuo, et al.
Publicado: (2025)
LongStory: Coherent, Complete and Length Controlled Long story Generation
por: Park, Kyeongman, et al.
Publicado: (2023)
por: Park, Kyeongman, et al.
Publicado: (2023)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
por: Katsis, Yannis, et al.
Publicado: (2025)
por: Katsis, Yannis, et al.
Publicado: (2025)
Dynamic benchmarking framework for LLM-based conversational data capture
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
por: Aluffi, Pietro Alessandro, et al.
Publicado: (2025)
LLMzSzŁ: a comprehensive LLM benchmark for Polish
por: Jassem, Krzysztof, et al.
Publicado: (2025)
por: Jassem, Krzysztof, et al.
Publicado: (2025)
Batayan: A Filipino NLP benchmark for evaluating Large Language Models
por: Montalan, Jann Railey, et al.
Publicado: (2025)
por: Montalan, Jann Railey, et al.
Publicado: (2025)
LongTail-Swap: benchmarking language models' abilities on rare words
por: Algayres, Robin, et al.
Publicado: (2025)
por: Algayres, Robin, et al.
Publicado: (2025)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
por: Gatti, Prajwal, et al.
Publicado: (2025)
por: Gatti, Prajwal, et al.
Publicado: (2025)
Polish-English medical knowledge transfer: A new benchmark and results
por: Grzybowski, Łukasz, et al.
Publicado: (2024)
por: Grzybowski, Łukasz, et al.
Publicado: (2024)
Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
por: Abdaljalil, Samir, et al.
Publicado: (2026)
por: Abdaljalil, Samir, et al.
Publicado: (2026)
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
por: Ichmoukhamedov, Timour, et al.
Publicado: (2024)
Simple and Scalable Strategies to Continually Pre-train Large Language Models
por: Ibrahim, Adam, et al.
Publicado: (2024)
por: Ibrahim, Adam, et al.
Publicado: (2024)
MinorBench: A hand-built benchmark for content-based risks for children
por: Khoo, Shaun, et al.
Publicado: (2025)
por: Khoo, Shaun, et al.
Publicado: (2025)
TelcoLM: collecting data, adapting, and benchmarking language models for the telecommunication domain
por: Barboule, Camille, et al.
Publicado: (2024)
por: Barboule, Camille, et al.
Publicado: (2024)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
por: Gupta, Lavanya, et al.
Publicado: (2024)
por: Gupta, Lavanya, et al.
Publicado: (2024)
Digital Twin Ecosystem for Oncology Clinical Operations
por: Pandey, Himanshu, et al.
Publicado: (2024)
por: Pandey, Himanshu, et al.
Publicado: (2024)
Multilingual LLMs Are Not Multilingual Thinkers: Evidence from Hindi Analogy Evaluation
por: Gupta, Ashray, et al.
Publicado: (2025)
por: Gupta, Ashray, et al.
Publicado: (2025)
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
por: Panagoulias, Dimitrios P., et al.
Publicado: (2024)
The Russian-focused embedders' exploration: ruMTEB benchmark and Russian embedding model design
por: Snegirev, Artem, et al.
Publicado: (2024)
por: Snegirev, Artem, et al.
Publicado: (2024)
A benchmark for joint dialogue satisfaction, emotion recognition, and emotion state transition prediction
por: Bian, Jing, et al.
Publicado: (2026)
por: Bian, Jing, et al.
Publicado: (2026)
Multilingual Controlled Generation And Gold-Standard-Agnostic Evaluation of Code-Mixed Sentences
por: Gupta, Ayushman, et al.
Publicado: (2024)
por: Gupta, Ayushman, et al.
Publicado: (2024)
A thorough benchmark of automatic text classification: From traditional approaches to large language models
por: Cunha, Washington, et al.
Publicado: (2025)
por: Cunha, Washington, et al.
Publicado: (2025)
QuanTemp: A real-world open-domain benchmark for fact-checking numerical claims
por: V, Venktesh, et al.
Publicado: (2024)
por: V, Venktesh, et al.
Publicado: (2024)
Leaving the barn door open for Clever Hans: Simple features predict LLM benchmark answers
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2024)
Towards Understanding the Robustness of LLM-based Evaluations under Perturbations
por: Chaudhary, Manav, et al.
Publicado: (2024)
por: Chaudhary, Manav, et al.
Publicado: (2024)
BEARCUBS: A benchmark for computer-using web agents
por: Song, Yixiao, et al.
Publicado: (2025)
por: Song, Yixiao, et al.
Publicado: (2025)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
por: Surikuchi, Aditya K, et al.
Publicado: (2024)
por: Surikuchi, Aditya K, et al.
Publicado: (2024)
Creativity Benchmark: A benchmark for marketing creativity for large language models
por: Bhat, Ninad, et al.
Publicado: (2025)
por: Bhat, Ninad, et al.
Publicado: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
por: Chen, Zixun, et al.
Publicado: (2025)
por: Chen, Zixun, et al.
Publicado: (2025)
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
por: Narsupalli, Yaswanth, et al.
Publicado: (2024)
por: Narsupalli, Yaswanth, et al.
Publicado: (2024)
Evaluating Concurrent Robustness of Language Models Across Diverse Challenge Sets
por: Gupta, Vatsal, et al.
Publicado: (2023)
por: Gupta, Vatsal, et al.
Publicado: (2023)
A New HOPE: Domain-agnostic Automatic Evaluation of Text Chunking
por: Brådland, Henrik, et al.
Publicado: (2025)
por: Brådland, Henrik, et al.
Publicado: (2025)
Evaluating Large Language Models on Rare Disease Diagnosis: A Case Study using House M.D
por: Gupta, Arsh, et al.
Publicado: (2025)
por: Gupta, Arsh, et al.
Publicado: (2025)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
Is 'Hope' a person or an idea? A pilot benchmark for NER: comparing traditional NLP tools and large language models on ambiguous entities
por: Latifi, Payam
Publicado: (2025)
por: Latifi, Payam
Publicado: (2025)
Ejemplares similares
-
Ada-LEval: Evaluating long-context LLMs with length-adaptable benchmarks
por: Wang, Chonghua, et al.
Publicado: (2024) -
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
por: Pandey, Atharva, et al.
Publicado: (2025) -
SD-E$^2$: Semantic Exploration for Reasoning Under Token Budgets
por: Mishra, Kshitij, et al.
Publicado: (2026) -
CXMArena: Unified Dataset to benchmark performance in realistic CXM Scenarios
por: Garg, Raghav, et al.
Publicado: (2025) -
Suvach -- Generated Hindi QA benchmark
por: Narayanan, Vaishak, et al.
Publicado: (2024)