Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Madusanka, Tharindu, Pratt-Hartmann, Ian, Batista-Navarro, Riza |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
by: Wu, Yulong, et al.
Published: (2025)
by: Wu, Yulong, et al.
Published: (2025)
A Note on the Complexity of the Satisfiability Problem for Graded Modal Logics
by: Kazakov, Yevgeny, et al.
Published: (2009)
by: Kazakov, Yevgeny, et al.
Published: (2009)
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024)
by: Feher, Darius, et al.
Published: (2024)
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)
by: Kumarage, Tharindu, et al.
Published: (2024)
Exploring the Performance of Large Language Models on Subjective Span Identification Tasks
by: Dmonte, Alphaeus, et al.
Published: (2026)
by: Dmonte, Alphaeus, et al.
Published: (2026)
CANTONMT: Investigating Back-Translation and Model-Switch Mechanisms for Cantonese-English Neural Machine Translation
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
CantonMT: Cantonese to English NMT Platform with Fine-Tuned Models Using Synthetic Back-Translation Data
by: Hong, Kung Yin, et al.
Published: (2024)
by: Hong, Kung Yin, et al.
Published: (2024)
Can Language Models Solve Graph Problems in Natural Language?
by: Wang, Heng, et al.
Published: (2023)
by: Wang, Heng, et al.
Published: (2023)
Towards Generalized Offensive Language Identification
by: Dmonte, Alphaeus, et al.
Published: (2024)
by: Dmonte, Alphaeus, et al.
Published: (2024)
Advancements in Natural Language Processing: Exploring Transformer-Based Architectures for Text Understanding
by: Wu, Tianhao, et al.
Published: (2025)
by: Wu, Tianhao, et al.
Published: (2025)
Combining Transformers with Natural Language Explanations
by: Ruggeri, Federico, et al.
Published: (2021)
by: Ruggeri, Federico, et al.
Published: (2021)
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
by: Hettiarachchi, Hansi, et al.
Published: (2024)
Defining and Evaluating Decision and Composite Risk in Language Models Applied to Natural Language Inference
by: Shen, Ke, et al.
Published: (2024)
by: Shen, Ke, et al.
Published: (2024)
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
by: Yuksekgonul, Mert, et al.
Published: (2023)
by: Yuksekgonul, Mert, et al.
Published: (2023)
GOLD: Geometry Problem Solver with Natural Language Description
by: Zhang, Jiaxin, et al.
Published: (2024)
by: Zhang, Jiaxin, et al.
Published: (2024)
WRDScore: New Metric for Evaluation of Natural Language Generation Models
by: Mussabayev, Ravil
Published: (2024)
by: Mussabayev, Ravil
Published: (2024)
Learning Tractable Distributions Of Language Model Continuations
by: Yidou-Weng, Gwen, et al.
Published: (2025)
by: Yidou-Weng, Gwen, et al.
Published: (2025)
Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
by: Li, Hang, et al.
Published: (2025)
by: Li, Hang, et al.
Published: (2025)
Exploring the Role of Reasoning Structures for Constructing Proofs in Multi-Step Natural Language Reasoning with Large Language Models
by: Zheng, Zi'ou, et al.
Published: (2024)
by: Zheng, Zi'ou, et al.
Published: (2024)
Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks
by: Anand, Avinash, et al.
Published: (2024)
by: Anand, Avinash, et al.
Published: (2024)
DataAgent: Evaluating Large Language Models' Ability to Answer Zero-Shot, Natural Language Queries
by: Mishra, Manit, et al.
Published: (2024)
by: Mishra, Manit, et al.
Published: (2024)
Transformer-based Causal Language Models Perform Clustering
by: Wu, Xinbo, et al.
Published: (2024)
by: Wu, Xinbo, et al.
Published: (2024)
Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM
by: Shetty, Samay U., et al.
Published: (2026)
by: Shetty, Samay U., et al.
Published: (2026)
Aspect-based Sentiment Evaluation of Chess Moves (ASSESS): an NLP-based Method for Evaluating Chess Strategies from Textbooks
by: Alrdahi, Haifa, et al.
Published: (2024)
by: Alrdahi, Haifa, et al.
Published: (2024)
GraphArena: Evaluating and Exploring Large Language Models on Graph Computation
by: Tang, Jianheng, et al.
Published: (2024)
by: Tang, Jianheng, et al.
Published: (2024)
On the Importance and Evaluation of Narrativity in Natural Language AI Explanations
by: Cedro, Mateusz, et al.
Published: (2026)
by: Cedro, Mateusz, et al.
Published: (2026)
Transformer-based Language Models for Reasoning in the Description Logic ALCQ
by: Poulis, Angelos, et al.
Published: (2024)
by: Poulis, Angelos, et al.
Published: (2024)
Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
by: Jia, Boyu, et al.
Published: (2025)
by: Jia, Boyu, et al.
Published: (2025)
Report Cards: Qualitative Evaluation of Language Models Using Natural Language Summaries
by: Yang, Blair, et al.
Published: (2024)
by: Yang, Blair, et al.
Published: (2024)
Quantization-Aware and Tensor-Compressed Training of Transformers for Natural Language Understanding
by: Yang, Zi, et al.
Published: (2023)
by: Yang, Zi, et al.
Published: (2023)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
A Survey of Automatic Hallucination Evaluation on Natural Language Generation
by: Qi, Siya, et al.
Published: (2024)
by: Qi, Siya, et al.
Published: (2024)
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
by: Saad-Falcon, Jon, et al.
Published: (2024)
by: Saad-Falcon, Jon, et al.
Published: (2024)
Genshin: General Shield for Natural Language Processing with Large Language Models
by: Peng, Xiao, et al.
Published: (2024)
by: Peng, Xiao, et al.
Published: (2024)
Querying Structured Data Through Natural Language Using Language Models
by: Valentin-Micu, Hontan, et al.
Published: (2026)
by: Valentin-Micu, Hontan, et al.
Published: (2026)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
by: Giorgi, Flavio, et al.
Published: (2024)
by: Giorgi, Flavio, et al.
Published: (2024)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
by: Zheng, Xin, et al.
Published: (2024)
by: Zheng, Xin, et al.
Published: (2024)
Progtuning: Progressive Fine-tuning Framework for Transformer-based Language Models
by: Ji, Xiaoshuang, et al.
Published: (2025)
by: Ji, Xiaoshuang, et al.
Published: (2025)
Similar Items
-
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
by: Wu, Yulong, et al.
Published: (2025) -
Pay Attention to Real World Perturbations! Natural Robustness Evaluation in Machine Reading Comprehension
by: Wu, Yulong, et al.
Published: (2025) -
A Note on the Complexity of the Satisfiability Problem for Graded Modal Logics
by: Kazakov, Yevgeny, et al.
Published: (2009) -
Learning to Generate and Evaluate Fact-checking Explanations with Transformers
by: Feher, Darius, et al.
Published: (2024) -
Harnessing Artificial Intelligence to Combat Online Hate: Exploring the Challenges and Opportunities of Large Language Models in Hate Speech Detection
by: Kumarage, Tharindu, et al.
Published: (2024)