Subjective Question Generation and Answer Evaluation using NLP
Fuente:
arXiv
Guardado en:
| Autores principales: | Islam, G. M. Refatul, Shaheer, Safwan, Nur, Yaseen, Hamid, Mohammad Rafid |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Reviewriter: AI-Generated Instructions For Peer Review Writing
por: Su, Xiaotian, et al.
Publicado: (2025)
por: Su, Xiaotian, et al.
Publicado: (2025)
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
por: Nguyen, Minh Hoang, et al.
Publicado: (2026)
por: Nguyen, Minh Hoang, et al.
Publicado: (2026)
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
por: Bayram, M. Ali, et al.
Publicado: (2024)
por: Bayram, M. Ali, et al.
Publicado: (2024)
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
por: Alhamzeh, Alaa, et al.
Publicado: (2025)
por: Alhamzeh, Alaa, et al.
Publicado: (2025)
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
por: Brant, Thiago, et al.
Publicado: (2026)
por: Brant, Thiago, et al.
Publicado: (2026)
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
Detecting Prompt Injection Attacks Against Application Using Classifiers
por: Shaheer, Safwan, et al.
Publicado: (2025)
por: Shaheer, Safwan, et al.
Publicado: (2025)
Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts
por: Seo, Yeongbin, et al.
Publicado: (2025)
por: Seo, Yeongbin, et al.
Publicado: (2025)
Towards Probabilistic Question Answering Over Tabular Data
por: Shen, Chen, et al.
Publicado: (2025)
por: Shen, Chen, et al.
Publicado: (2025)
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
por: Bayram, M. Ali, et al.
Publicado: (2025)
por: Bayram, M. Ali, et al.
Publicado: (2025)
The Syntactic Acceptability Dataset (Preview): A Resource for Machine Learning and Linguistic Analysis of English
por: Juzek, Tom S
Publicado: (2025)
por: Juzek, Tom S
Publicado: (2025)
Tokens with Meaning: A Hybrid Tokenization Approach for Turkish
por: Bayram, M. Ali, et al.
Publicado: (2025)
por: Bayram, M. Ali, et al.
Publicado: (2025)
Evaluating Input Feature Explanations through a Unified Diagnostic Evaluation Framework
por: Sun, Jingyi, et al.
Publicado: (2024)
por: Sun, Jingyi, et al.
Publicado: (2024)
Exploring State Tracking Capabilities of Large Language Models
por: Rezaee, Kiamehr, et al.
Publicado: (2025)
por: Rezaee, Kiamehr, et al.
Publicado: (2025)
Evaluating Pixel Language Models on Non-Standardized Languages
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2024)
FairLangProc: A Python package for fairness in NLP
por: Pérez-Peralta, Arturo, et al.
Publicado: (2025)
por: Pérez-Peralta, Arturo, et al.
Publicado: (2025)
Experimentation in Content Moderation using RWKV
por: Yildirim, Umut, et al.
Publicado: (2024)
por: Yildirim, Umut, et al.
Publicado: (2024)
The Unlikely Duel: Evaluating Creative Writing in LLMs through a Unique Scenario
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
por: Gómez-Rodríguez, Carlos, et al.
Publicado: (2024)
Low-Resource English-Tigrinya MT: Leveraging Multilingual Models, Custom Tokenizers, and Clean Evaluation Benchmarks
por: Teklehaymanot, Hailay Kidu, et al.
Publicado: (2025)
por: Teklehaymanot, Hailay Kidu, et al.
Publicado: (2025)
ScoreRAG: A Retrieval-Augmented Generation Framework with Consistency-Relevance Scoring and Structured Summarization for News Generation
por: Lin, Pei-Yun, et al.
Publicado: (2025)
por: Lin, Pei-Yun, et al.
Publicado: (2025)
How GenAI Mentor Configurations Shape Early Collaborative Dynamics: A Classroom Comparison of Individual and Shared Agents
por: Zha, Siyu, et al.
Publicado: (2026)
por: Zha, Siyu, et al.
Publicado: (2026)
A Survey on Natural Language Counterfactual Generation
por: Wang, Yongjie, et al.
Publicado: (2024)
por: Wang, Yongjie, et al.
Publicado: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
por: Pan, Leyi, et al.
Publicado: (2025)
por: Pan, Leyi, et al.
Publicado: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
por: Weigang, Li, et al.
Publicado: (2025)
por: Weigang, Li, et al.
Publicado: (2025)
NurValues: Real-World Nursing Values Evaluation for Large Language Models in Clinical Context
por: Yao, Ben, et al.
Publicado: (2025)
por: Yao, Ben, et al.
Publicado: (2025)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2023)
por: Muñoz-Ortiz, Alberto, et al.
Publicado: (2023)
Adaptive Steering and Remasking for Safe Generation in Diffusion Language Models
por: Lee, Yejin, et al.
Publicado: (2026)
por: Lee, Yejin, et al.
Publicado: (2026)
When Retrieval Succeeds and Fails: Rethinking Retrieval-Augmented Generation for LLMs
por: Wang, Yongjie, et al.
Publicado: (2025)
por: Wang, Yongjie, et al.
Publicado: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
por: Tu, Songjun, et al.
Publicado: (2026)
por: Tu, Songjun, et al.
Publicado: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
por: Nguyen, Minh Hoang, et al.
Publicado: (2025)
Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese
por: Pranida, Salsabila Zahirah, et al.
Publicado: (2025)
por: Pranida, Salsabila Zahirah, et al.
Publicado: (2025)
Large Language Models(LLMs) on Tabular Data: Prediction, Generation, and Understanding -- A Survey
por: Fang, Xi, et al.
Publicado: (2024)
por: Fang, Xi, et al.
Publicado: (2024)
Measuring text summarization factuality using atomic facts entailment metrics in the context of retrieval augmented generation
por: Kriman, N. E.
Publicado: (2024)
por: Kriman, N. E.
Publicado: (2024)
LegalGuardian: A Privacy-Preserving Framework for Secure Integration of Large Language Models in Legal Practice
por: Demir, M. Mikail, et al.
Publicado: (2025)
por: Demir, M. Mikail, et al.
Publicado: (2025)
Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"
por: Dingfelder, Philipp, et al.
Publicado: (2025)
por: Dingfelder, Philipp, et al.
Publicado: (2025)
Multiplication in Multimodal LLMs: Computation with Text, Image, and Audio Inputs
por: Balter, Samuel G., et al.
Publicado: (2026)
por: Balter, Samuel G., et al.
Publicado: (2026)
Generating Natural-Language Surgical Feedback: From Structured Representation to Domain-Grounded Evaluation
por: Nasriddinov, Firdavs, et al.
Publicado: (2025)
por: Nasriddinov, Firdavs, et al.
Publicado: (2025)
Evaluation of Table Representations to Answer Questions from Tables in Documents : A Case Study using 3GPP Specifications
por: Roychowdhury, Sujoy, et al.
Publicado: (2024)
por: Roychowdhury, Sujoy, et al.
Publicado: (2024)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
por: Soman, Sumit, et al.
Publicado: (2025)
por: Soman, Sumit, et al.
Publicado: (2025)
Ejemplares similares
-
Reviewriter: AI-Generated Instructions For Peer Review Writing
por: Su, Xiaotian, et al.
Publicado: (2025) -
From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoring
por: Nguyen, Minh Hoang, et al.
Publicado: (2026) -
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
por: Bayram, M. Ali, et al.
Publicado: (2024) -
Argument Quality Annotation and Gender Bias Detection in Financial Communication through Large Language Models
por: Alhamzeh, Alaa, et al.
Publicado: (2025) -
Estimating Exam Item Difficulty with LLMs: A Benchmark on Brazil's ENEM Corpus
por: Brant, Thiago, et al.
Publicado: (2026)