Unified Deployment-Aware Evaluation of Open Reasoning Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Manik, Md Motaleb Hossen, Wang, Ge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emergent decentralized regulation in a purely synthetic society
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
Question-Answering System for Bangla: Fine-tuning BERT-Bangla for a Closed Domain
di: Roy, Subal Chandra, et al.
Pubblicazione: (2024)
di: Roy, Subal Chandra, et al.
Pubblicazione: (2024)
OpenClaw Agents on Moltbook: Risky Instruction Sharing and Norm Enforcement in an Agent-Only Social Network
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time Interaction
di: Islam, Md Zabirul, et al.
Pubblicazione: (2025)
di: Islam, Md Zabirul, et al.
Pubblicazione: (2025)
SlideChain: Semantic Provenance for Lecture Understanding via Blockchain Registration
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2025)
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2025)
ADAPT: AI-Driven Decentralized Adaptive Publishing Testbed
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026)
ChatGPT vs. DeepSeek: A Comparative Study on AI-Based Code Generation
di: Manik, Md Motaleb Hossen
Pubblicazione: (2025)
di: Manik, Md Motaleb Hossen
Pubblicazione: (2025)
N-ReLU: Zero-Mean Stochastic Extension of ReLU
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2025)
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2025)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
di: Hua, Tim Tian, et al.
Pubblicazione: (2025)
Pitfalls of Evaluating Language Models with Open Benchmarks
di: Hasan, Md. Najib, et al.
Pubblicazione: (2025)
di: Hasan, Md. Najib, et al.
Pubblicazione: (2025)
Social Media Sentiments Analysis on the July Revolution in Bangladesh: A Hybrid Transformer Based Machine Learning Approach
di: Hossen, Md. Sabbir, et al.
Pubblicazione: (2025)
di: Hossen, Md. Sabbir, et al.
Pubblicazione: (2025)
Can't say cant? Measuring and Reasoning of Dark Jargons in Large Language Models
di: Ji, Xu, et al.
Pubblicazione: (2024)
di: Ji, Xu, et al.
Pubblicazione: (2024)
CodeMixBench: Evaluating Large Language Models on Code Generation with Code-Mixed Prompts
di: Sheokand, Manik, et al.
Pubblicazione: (2025)
di: Sheokand, Manik, et al.
Pubblicazione: (2025)
Risks, Causes, and Mitigations of Widespread Deployments of Large Language Models (LLMs): A Survey
di: Sakib, Md Nazmus, et al.
Pubblicazione: (2024)
di: Sakib, Md Nazmus, et al.
Pubblicazione: (2024)
Language of Persuasion and Misrepresentation in Business Communication: A Textual Detection Approach
di: Hossen, Sayem, et al.
Pubblicazione: (2025)
di: Hossen, Sayem, et al.
Pubblicazione: (2025)
EffiReason-Bench: A Unified Benchmark for Evaluating and Advancing Efficient Reasoning in Large Language Models
di: Huang, Junquan, et al.
Pubblicazione: (2025)
di: Huang, Junquan, et al.
Pubblicazione: (2025)
Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models
di: Islam, Shayekh Bin, et al.
Pubblicazione: (2024)
di: Islam, Shayekh Bin, et al.
Pubblicazione: (2024)
CMULAB: An Open-Source Framework for Training and Deployment of Natural Language Processing Models
di: Sheikh, Zaid, et al.
Pubblicazione: (2024)
di: Sheikh, Zaid, et al.
Pubblicazione: (2024)
Secure On-Premise Deployment of Open-Weights Large Language Models in Radiology: An Isolation-First Architecture with Prospective Pilot Evaluation
di: Nowak, Sebastian, et al.
Pubblicazione: (2026)
di: Nowak, Sebastian, et al.
Pubblicazione: (2026)
Knowledge Reasoning Language Model: Unifying Knowledge and Language for Inductive Knowledge Graph Reasoning
di: Zhuo, Xingrui, et al.
Pubblicazione: (2025)
di: Zhuo, Xingrui, et al.
Pubblicazione: (2025)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
di: Mekky, Ali, et al.
Pubblicazione: (2025)
di: Mekky, Ali, et al.
Pubblicazione: (2025)
A CNN-Based Malaria Diagnosis from Blood Cell Images with SHAP and LIME Explainability
di: Abir, Md. Ismiel Hossen, et al.
Pubblicazione: (2025)
di: Abir, Md. Ismiel Hossen, et al.
Pubblicazione: (2025)
Evaluating OpenAI GPT Models for Translation of Endangered Uralic Languages: A Comparison of Reasoning and Non-Reasoning Architectures
di: Tereshchenko, Yehor, et al.
Pubblicazione: (2025)
di: Tereshchenko, Yehor, et al.
Pubblicazione: (2025)
MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models
di: Dihan, Mahir Labib, et al.
Pubblicazione: (2024)
di: Dihan, Mahir Labib, et al.
Pubblicazione: (2024)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
di: Amjad, Husnain, et al.
Pubblicazione: (2026)
di: Amjad, Husnain, et al.
Pubblicazione: (2026)
Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
di: Zhu, Junda, et al.
Pubblicazione: (2025)
di: Zhu, Junda, et al.
Pubblicazione: (2025)
OpenLearnLM Benchmark: A Unified Framework for Evaluating Knowledge, Skill, and Attitude in Educational Large Language Models
di: Lee, Unggi, et al.
Pubblicazione: (2026)
di: Lee, Unggi, et al.
Pubblicazione: (2026)
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
di: Yang, Ge, et al.
Pubblicazione: (2024)
di: Yang, Ge, et al.
Pubblicazione: (2024)
Evaluating Accounting Reasoning Capabilities of Large Language Models
di: Zhou, Jie, et al.
Pubblicazione: (2026)
di: Zhou, Jie, et al.
Pubblicazione: (2026)
Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language Models
di: Liu, Runxuan, et al.
Pubblicazione: (2026)
di: Liu, Runxuan, et al.
Pubblicazione: (2026)
Transformer-Driven Triple Fusion Framework for Enhanced Multimodal Author Intent Classification in Low-Resource Bangla
di: Islam, Ariful, et al.
Pubblicazione: (2025)
di: Islam, Ariful, et al.
Pubblicazione: (2025)
BanglaSentNet: An Explainable Hybrid Deep Learning Framework for Multi-Aspect Sentiment Analysis with Cross-Domain Transfer Learning
di: Islam, Ariful, et al.
Pubblicazione: (2025)
di: Islam, Ariful, et al.
Pubblicazione: (2025)
ThinkDial: An Open Recipe for Controlling Reasoning Effort in Large Language Models
di: He, Qianyu, et al.
Pubblicazione: (2025)
di: He, Qianyu, et al.
Pubblicazione: (2025)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
di: Wang, Junlin, et al.
Pubblicazione: (2024)
di: Wang, Junlin, et al.
Pubblicazione: (2024)
Analyzing the Dynamics of COVID-19 Lockdown Success: Insights from Regional Data and Public Health Measures
di: Manik, Md. Motaleb Hossen, et al.
Pubblicazione: (2024)
di: Manik, Md. Motaleb Hossen, et al.
Pubblicazione: (2024)
The Hawthorne Effect in Reasoning Models: Evaluating and Steering Test Awareness
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2025)
di: Abdelnabi, Sahar, et al.
Pubblicazione: (2025)
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
di: Wang, Jun, et al.
Pubblicazione: (2024)
di: Wang, Jun, et al.
Pubblicazione: (2024)
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective
di: Yu, Yiyao, et al.
Pubblicazione: (2025)
di: Yu, Yiyao, et al.
Pubblicazione: (2025)
FERA: Uncertainty-Aware Federated Reasoning for Large Language Models
di: Wang, Ruhan, et al.
Pubblicazione: (2026)
di: Wang, Ruhan, et al.
Pubblicazione: (2026)
Small Language Models: Architectures, Techniques, Evaluation, Problems and Future Adaptation
di: Sakib, Tanjil Hasan, et al.
Pubblicazione: (2025)
di: Sakib, Tanjil Hasan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Emergent decentralized regulation in a purely synthetic society
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026) -
Question-Answering System for Bangla: Fine-tuning BERT-Bangla for a Closed Domain
di: Roy, Subal Chandra, et al.
Pubblicazione: (2024) -
OpenClaw Agents on Moltbook: Risky Instruction Sharing and Norm Enforcement in an Agent-Only Social Network
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2026) -
ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time Interaction
di: Islam, Md Zabirul, et al.
Pubblicazione: (2025) -
SlideChain: Semantic Provenance for Lecture Understanding via Blockchain Registration
di: Manik, Md Motaleb Hossen, et al.
Pubblicazione: (2025)