PragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational Dynamics
Fuente:
arXiv
Saved in:
| Main Authors: | Vashistha, Sachin, Bibhuti, Aryan, Naik, Atharva, Tutek, Martin, Aditya, Somak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023)
by: Rao, Abhinav, et al.
Published: (2023)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
by: Puerto, Haritz, et al.
Published: (2024)
by: Puerto, Haritz, et al.
Published: (2024)
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
by: Pandey, Saurabh Kumar, et al.
Published: (2025)
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024)
by: Rao, Abhinav, et al.
Published: (2024)
Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
by: Dukić, David, et al.
Published: (2025)
by: Dukić, David, et al.
Published: (2025)
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
by: Simhi, Adi, et al.
Published: (2026)
by: Simhi, Adi, et al.
Published: (2026)
MultiPragEval: Multilingual Pragmatic Evaluation of Large Language Models
by: Park, Dojun, et al.
Published: (2024)
by: Park, Dojun, et al.
Published: (2024)
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
by: Hong, Pengfei, et al.
Published: (2024)
by: Hong, Pengfei, et al.
Published: (2024)
MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
by: Dutta, Aritra, et al.
Published: (2026)
by: Dutta, Aritra, et al.
Published: (2026)
EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
by: Ray, Sourjyadip, et al.
Published: (2025)
by: Ray, Sourjyadip, et al.
Published: (2025)
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
by: Liu, Emmy, et al.
Published: (2026)
by: Liu, Emmy, et al.
Published: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
by: Ashuach, Tomer, et al.
Published: (2024)
by: Ashuach, Tomer, et al.
Published: (2024)
MATHSENSEI: A Tool-Augmented Large Language Model for Mathematical Reasoning
by: Das, Debrup, et al.
Published: (2024)
by: Das, Debrup, et al.
Published: (2024)
CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language Models
by: Oba, Miyu, et al.
Published: (2026)
by: Oba, Miyu, et al.
Published: (2026)
ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
by: Kapadnis, Manav Nitin, et al.
Published: (2026)
TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text
by: Adak, Sayantan, et al.
Published: (2024)
by: Adak, Sayantan, et al.
Published: (2024)
Context Parametrization with Compositional Adapters
by: Jukić, Josip, et al.
Published: (2025)
by: Jukić, Josip, et al.
Published: (2025)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
by: Sachdeva, Rachneet, et al.
Published: (2023)
by: Sachdeva, Rachneet, et al.
Published: (2023)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
by: Seth, Agrima, et al.
Published: (2025)
by: Seth, Agrima, et al.
Published: (2025)
Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages
by: Shahid, Farhana, et al.
Published: (2025)
by: Shahid, Farhana, et al.
Published: (2025)
Multi-Domain ABSA Conversation Dataset Generation via LLMs for Real-World Evaluation and Model Comparison
by: Pandit, Tejul, et al.
Published: (2025)
by: Pandit, Tejul, et al.
Published: (2025)
Can LLMs Infer Personality from Real World Conversations?
by: Zhu, Jianfeng, et al.
Published: (2025)
by: Zhu, Jianfeng, et al.
Published: (2025)
CRScore: Grounding Automated Evaluation of Code Review Comments in Code Claims and Smells
by: Naik, Atharva, et al.
Published: (2024)
by: Naik, Atharva, et al.
Published: (2024)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024)
by: Ray, Sourjyadip, et al.
Published: (2024)
DaLA: Danish Linguistic Acceptability Evaluation Guided by Real World Errors
by: Barmina, Gianluca, et al.
Published: (2025)
by: Barmina, Gianluca, et al.
Published: (2025)
The Battle of LLMs: A Comparative Study in Conversational QA Tasks
by: Rangapur, Aryan, et al.
Published: (2024)
by: Rangapur, Aryan, et al.
Published: (2024)
IOLBENCH: Benchmarking LLMs on Linguistic Reasoning
by: Goyal, Satyam, et al.
Published: (2025)
by: Goyal, Satyam, et al.
Published: (2025)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
by: Taktasheva, Ekaterina, et al.
Published: (2024)
by: Taktasheva, Ekaterina, et al.
Published: (2024)
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
by: Roy, Aniruddha, et al.
Published: (2025)
by: Roy, Aniruddha, et al.
Published: (2025)
Can LLMs Learn to Map the World from Local Descriptions?
by: Xia, Sirui, et al.
Published: (2025)
by: Xia, Sirui, et al.
Published: (2025)
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025)
by: Başar, Ezgi, et al.
Published: (2025)
Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models
by: Zhou, Xinyu, et al.
Published: (2024)
by: Zhou, Xinyu, et al.
Published: (2024)
NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
by: Dutta, Aritra, et al.
Published: (2025)
by: Dutta, Aritra, et al.
Published: (2025)
MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World Conversation
by: Xue, Haochen, et al.
Published: (2025)
by: Xue, Haochen, et al.
Published: (2025)
TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice
by: Hu, Gang, et al.
Published: (2026)
by: Hu, Gang, et al.
Published: (2026)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
by: Luo, Man, et al.
Published: (2023)
by: Luo, Man, et al.
Published: (2023)
Programming by Examples Meets Historical Linguistics: A Large Language Model Based Approach to Sound Law Induction
by: Naik, Atharva, et al.
Published: (2025)
by: Naik, Atharva, et al.
Published: (2025)
Similar Items
-
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
by: Rao, Abhinav, et al.
Published: (2023) -
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
by: Puerto, Haritz, et al.
Published: (2024) -
SMAB: MAB based word Sensitivity Estimation Framework and its Applications in Adversarial Text Generation
by: Pandey, Saurabh Kumar, et al.
Published: (2025) -
[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs
by: Rao, Abhinav, et al.
Published: (2024) -
Characterizing Linguistic Shifts in Croatian News via Diachronic Word Embeddings
by: Dukić, David, et al.
Published: (2025)