DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Hasan, Md Hasebul, Charu, Krity Haque, Sridhar, Eshwara Prasad, Deb, Shuchisnigdha, Islam, Mohammad A. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
por: Hasan, Md Hasebul, et al.
Publicado: (2025)
por: Hasan, Md Hasebul, et al.
Publicado: (2025)
A Multi-Modal Deep Learning Based Approach for House Price Prediction
por: Hasan, Md Hasebul, et al.
Publicado: (2024)
por: Hasan, Md Hasebul, et al.
Publicado: (2024)
EasyMath: A 0-shot Math Benchmark for SLMs
por: Karki, Drishya, et al.
Publicado: (2025)
por: Karki, Drishya, et al.
Publicado: (2025)
Ensemble Language Models for Multilingual Sentiment Analysis
por: Hasan, Md Arid
Publicado: (2024)
por: Hasan, Md Arid
Publicado: (2024)
Depthwise Separable Convolutions with Deep Residual Convolutions
por: Hasan, Md Arid, et al.
Publicado: (2024)
por: Hasan, Md Arid, et al.
Publicado: (2024)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
por: Yuan, Weikang, et al.
Publicado: (2025)
por: Yuan, Weikang, et al.
Publicado: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
por: Haque, Md. Asraful, et al.
Publicado: (2026)
por: Haque, Md. Asraful, et al.
Publicado: (2026)
Grokking in the Wild: Data Augmentation for Real-World Multi-Hop Reasoning with Transformers
por: Abramov, Roman, et al.
Publicado: (2025)
por: Abramov, Roman, et al.
Publicado: (2025)
Evaluating LLM Metrics Through Real-World Capabilities
por: Miller, Justin K, et al.
Publicado: (2025)
por: Miller, Justin K, et al.
Publicado: (2025)
On Preserving the Knowledge of Long Clinical Texts
por: Hasan, Mohammad Junayed, et al.
Publicado: (2023)
por: Hasan, Mohammad Junayed, et al.
Publicado: (2023)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
por: Schopf, Tim, et al.
Publicado: (2026)
por: Schopf, Tim, et al.
Publicado: (2026)
Guardians of the Web: The Evolution and Future of Website Information Security
por: Islam, Md Saiful, et al.
Publicado: (2025)
por: Islam, Md Saiful, et al.
Publicado: (2025)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
por: Arief, Hasan
Publicado: (2026)
por: Arief, Hasan
Publicado: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
por: Wang, Yuxia, et al.
Publicado: (2024)
por: Wang, Yuxia, et al.
Publicado: (2024)
DESS: DeBERTa Enhanced Syntactic-Semantic Aspect Sentiment Triplet Extraction
por: Thenuwara, Vishal, et al.
Publicado: (2025)
por: Thenuwara, Vishal, et al.
Publicado: (2025)
SpokenNativQA: Multilingual Everyday Spoken Queries for LLMs
por: Alam, Firoj, et al.
Publicado: (2025)
por: Alam, Firoj, et al.
Publicado: (2025)
Detecting Prompt Injection Attacks Against Application Using Classifiers
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
UrduFactCheck: An Agentic Fact-Checking Framework for Urdu with Evidence Boosting and Benchmarking
por: Ahmad, Sarfraz, et al.
Publicado: (2025)
por: Ahmad, Sarfraz, et al.
Publicado: (2025)
CSTRL: Context-Driven Sequential Transfer Learning for Abstractive Radiology Report Summarization
por: Naznin, Mst. Fahmida Sultana, et al.
Publicado: (2025)
por: Naznin, Mst. Fahmida Sultana, et al.
Publicado: (2025)
Enhancing experiential learning through virtual reality: System design and a case study in additive manufacturing
por: Rafia Rahman Rafa, et al.
Publicado: (2024)
por: Rafia Rahman Rafa, et al.
Publicado: (2024)
BLP-2023 Task 2: Sentiment Analysis
por: Hasan, Md. Arid, et al.
Publicado: (2023)
por: Hasan, Md. Arid, et al.
Publicado: (2023)
Multilingual De-Duplication Strategies: Applying scalable similarity search with monolingual & multilingual embedding models
por: Pasch, Stefan, et al.
Publicado: (2024)
por: Pasch, Stefan, et al.
Publicado: (2024)
TituLLMs: A Family of Bangla LLMs with Comprehensive Benchmarking
por: Nahin, Shahriar Kabir, et al.
Publicado: (2025)
por: Nahin, Shahriar Kabir, et al.
Publicado: (2025)
From Facts to Folklore: Evaluating Large Language Models on Bengali Cultural Knowledge
por: Chowdhury, Nafis, et al.
Publicado: (2025)
por: Chowdhury, Nafis, et al.
Publicado: (2025)
Benchmarking Retrieval Strategies for Biomedical Retrieval-Augmented Generation: A Controlled Empirical Study
por: Bal, Devi Prasad, et al.
Publicado: (2026)
por: Bal, Devi Prasad, et al.
Publicado: (2026)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
por: Xu, Qiyuan, et al.
Publicado: (2026)
por: Xu, Qiyuan, et al.
Publicado: (2026)
Evaluating the Environmental Impact of using SLMs and Prompt Engineering for Code Generation
por: Mamun, Md Afif Al, et al.
Publicado: (2026)
por: Mamun, Md Afif Al, et al.
Publicado: (2026)
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
por: Mousi, Basel, et al.
Publicado: (2024)
por: Mousi, Basel, et al.
Publicado: (2024)
Automated Bug Triaging using Instruction-Tuned Large Language Models
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
por: Kiashemshaki, Kiana, et al.
Publicado: (2025)
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings
por: Hasan, Md. Arid, et al.
Publicado: (2024)
por: Hasan, Md. Arid, et al.
Publicado: (2024)
Physical oceanography during Walther Herwig I cruise WH27
por: Anonymous
Publicado: (2011)
por: Anonymous
Publicado: (2011)
Enhancing Wide-Angle Image Using Narrow-Angle View of the Same Scene
por: Safwan, Hussain Md., et al.
Publicado: (2025)
por: Safwan, Hussain Md., et al.
Publicado: (2025)
Primary production within the euphotic layer of the Norwegian Sea in July 1997: measurements by different methods
por: Sapozhnikov, Victor V, et al.
Publicado: (2000)
por: Sapozhnikov, Victor V, et al.
Publicado: (2000)
LLM-Based Multi-Task Bangla Hate Speech Detection: Type, Severity, and Target
por: Hasan, Md Arid, et al.
Publicado: (2025)
por: Hasan, Md Arid, et al.
Publicado: (2025)
Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs
por: Paulsen, Norman
Publicado: (2025)
por: Paulsen, Norman
Publicado: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
por: Yao, Ben, et al.
Publicado: (2025)
por: Yao, Ben, et al.
Publicado: (2025)
Enhancing Mental Health Counseling Support in Bangladesh using Culturally-Grounded Knowledge
por: Hasan, Md Arid, et al.
Publicado: (2026)
por: Hasan, Md Arid, et al.
Publicado: (2026)
Enhancing Automated Essay Scoring with Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training
por: Choi, Hongseok, et al.
Publicado: (2026)
por: Choi, Hongseok, et al.
Publicado: (2026)
Ejemplares similares
-
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
por: Hasan, Md Hasebul, et al.
Publicado: (2025) -
A Multi-Modal Deep Learning Based Approach for House Price Prediction
por: Hasan, Md Hasebul, et al.
Publicado: (2024) -
EasyMath: A 0-shot Math Benchmark for SLMs
por: Karki, Drishya, et al.
Publicado: (2025) -
Ensemble Language Models for Multilingual Sentiment Analysis
por: Hasan, Md Arid
Publicado: (2024) -
Depthwise Separable Convolutions with Deep Residual Convolutions
por: Hasan, Md Arid, et al.
Publicado: (2024)