PLUGH: A Benchmark for Spatial Understanding and Reasoning in Large Language Models
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Tikhonov, Alexey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CAPE: Corrective Actions from Precondition Errors using Large Language Models
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
von: Radosky, Lukas, et al.
Veröffentlicht: (2026)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
von: Badshah, Sher, et al.
Veröffentlicht: (2024)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
von: Tikhonov, Alexey, et al.
Veröffentlicht: (2026)
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
von: Fu, Tianyu, et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
von: Basit, Abdul, et al.
Veröffentlicht: (2025)
Branching Narratives: Character Decision Points Detection
von: Tikhonov, Alexey
Veröffentlicht: (2024)
von: Tikhonov, Alexey
Veröffentlicht: (2024)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
von: Łajewska, Weronika, et al.
Veröffentlicht: (2025)
von: Łajewska, Weronika, et al.
Veröffentlicht: (2025)
LAraBench: Benchmarking Arabic AI with Large Language Models
von: Abdelali, Ahmed, et al.
Veröffentlicht: (2023)
von: Abdelali, Ahmed, et al.
Veröffentlicht: (2023)
Murphys Laws of AI Alignment: Why the Gap Always Wins
von: Gaikwad, Madhava
Veröffentlicht: (2025)
von: Gaikwad, Madhava
Veröffentlicht: (2025)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
InterFeat: A Pipeline for Finding Interesting Scientific Features
von: Ofer, Dan, et al.
Veröffentlicht: (2025)
von: Ofer, Dan, et al.
Veröffentlicht: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
von: Berg, Cameron, et al.
Veröffentlicht: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Large Language Models for Propaganda Span Annotation
von: Hasanain, Maram, et al.
Veröffentlicht: (2023)
von: Hasanain, Maram, et al.
Veröffentlicht: (2023)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
von: Günther, Michael, et al.
Veröffentlicht: (2024)
von: Günther, Michael, et al.
Veröffentlicht: (2024)
Query-Aware Flow Diffusion for Graph-Based RAG with Retrieval Guarantees
von: Zhou, Zhuoping, et al.
Veröffentlicht: (2026)
von: Zhou, Zhuoping, et al.
Veröffentlicht: (2026)
A Study into Investigating Temporal Robustness of LLMs
von: Wallat, Jonas, et al.
Veröffentlicht: (2025)
von: Wallat, Jonas, et al.
Veröffentlicht: (2025)
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
von: Butt, Umer, et al.
Veröffentlicht: (2024)
von: Butt, Umer, et al.
Veröffentlicht: (2024)
skLEP: A Slovak General Language Understanding Benchmark
von: Šuppa, Marek, et al.
Veröffentlicht: (2025)
von: Šuppa, Marek, et al.
Veröffentlicht: (2025)
IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text
von: Pall, Rajveer Singh
Veröffentlicht: (2026)
von: Pall, Rajveer Singh
Veröffentlicht: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
Efficient Code Embeddings from Code Generation Models
von: Kryvosheieva, Daria, et al.
Veröffentlicht: (2025)
von: Kryvosheieva, Daria, et al.
Veröffentlicht: (2025)
A Multiple-Fill-in-the-Blank Exam Approach for Enhancing Zero-Resource Hallucination Detection in Large Language Models
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
von: Munakata, Satoshi, et al.
Veröffentlicht: (2024)
Kodezi Chronos: A Debugging-First Language Model for Repository-Scale Code Understanding
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
von: Khan, Ishraq, et al.
Veröffentlicht: (2025)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
von: Abhishek, Alok, et al.
Veröffentlicht: (2025)
Extracting Sentence Embeddings from Pretrained Transformer Models
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness Calibration
von: Amanlou, Mohammad, et al.
Veröffentlicht: (2026)
von: Amanlou, Mohammad, et al.
Veröffentlicht: (2026)
Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
von: Alam, Firoj, et al.
Veröffentlicht: (2026)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
von: Li, Yangyang
Veröffentlicht: (2025)
von: Li, Yangyang
Veröffentlicht: (2025)
Multi-Task Contrastive Learning for 8192-Token Bilingual Text Embeddings
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
von: Mohr, Isabelle, et al.
Veröffentlicht: (2024)
Automatic Cardiac Risk Management Classification using large-context Electronic Patients Health Records
von: Vitale, Jacopo, et al.
Veröffentlicht: (2026)
von: Vitale, Jacopo, et al.
Veröffentlicht: (2026)
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
von: Günther, Michael, et al.
Veröffentlicht: (2025)
von: Günther, Michael, et al.
Veröffentlicht: (2025)
jina-reranker-v3: Last but Not Late Interaction for Listwise Document Reranking
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
von: Jha, Rohan, et al.
Veröffentlicht: (2024)
von: Jha, Rohan, et al.
Veröffentlicht: (2024)
jina-embeddings-v3: Multilingual Embeddings With Task LoRA
von: Sturua, Saba, et al.
Veröffentlicht: (2024)
von: Sturua, Saba, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
CAPE: Corrective Actions from Precondition Errors using Large Language Models
von: Raman, Shreyas Sundara, et al.
Veröffentlicht: (2022) -
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025) -
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
von: Kaiser, Daniel, et al.
Veröffentlicht: (2025) -
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
von: Radosky, Lukas, et al.
Veröffentlicht: (2026) -
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
von: Badshah, Sher, et al.
Veröffentlicht: (2024)