Distortions in Judged Spatial Relations in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fulman, Nir, Memduhoğlu, Abdulkadir, Zipf, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
von: Razumenko, Itay, et al.
Veröffentlicht: (2026)
von: Razumenko, Itay, et al.
Veröffentlicht: (2026)
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
von: Abdulkadir, Perry
Veröffentlicht: (2025)
von: Abdulkadir, Perry
Veröffentlicht: (2025)
Deep Learning Enhanced Road Traffic Analysis: Scalable Vehicle Detection and Velocity Estimation Using PlanetScope Imagery
von: Adamiak, Maciej, et al.
Veröffentlicht: (2024)
von: Adamiak, Maciej, et al.
Veröffentlicht: (2024)
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
FedJudge: Federated Legal Large Language Model
von: Yue, Linan, et al.
Veröffentlicht: (2023)
von: Yue, Linan, et al.
Veröffentlicht: (2023)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
Multi-Bit Distortion-Free Watermarking for Large Language Models
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
von: Boroujeny, Massieh Kordi, et al.
Veröffentlicht: (2024)
Radio: Rate-Distortion Optimization for Large Language Model Compression
von: Young, Sean I.
Veröffentlicht: (2025)
von: Young, Sean I.
Veröffentlicht: (2025)
Evaluation of Geographical Distortions in Language Models
von: Decoupes, Rémy, et al.
Veröffentlicht: (2024)
von: Decoupes, Rémy, et al.
Veröffentlicht: (2024)
Do Large Language Models Judge Error Severity Like Humans?
von: Sun, Diege, et al.
Veröffentlicht: (2025)
von: Sun, Diege, et al.
Veröffentlicht: (2025)
When Large Language Models are Reliable for Judging Empathic Communication
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
von: Kumar, Aakriti, et al.
Veröffentlicht: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Grounding Spatial Relations in Text-Only Language Models
von: Azkune, Gorka, et al.
Veröffentlicht: (2024)
von: Azkune, Gorka, et al.
Veröffentlicht: (2024)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
The Birth of Knowledge: Emergent Features across Time, Space, and Scale in Large Language Models
von: Sawmya, Shashata, et al.
Veröffentlicht: (2025)
von: Sawmya, Shashata, et al.
Veröffentlicht: (2025)
BenchOverflow: Measuring Overflow in Large Language Models via Plain-Text Prompts
von: Feiglin, Erin, et al.
Veröffentlicht: (2026)
von: Feiglin, Erin, et al.
Veröffentlicht: (2026)
Meta-Judging with Large Language Models: Concepts, Methods, and Challenges
von: Silva, Hugo, et al.
Veröffentlicht: (2026)
von: Silva, Hugo, et al.
Veröffentlicht: (2026)
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
von: Molenda, Piotr, et al.
Veröffentlicht: (2024)
Toward Trustworthy Difficulty Assessments: Large Language Models as Judges in Programming and Synthetic Tasks
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2025)
von: Tabib, H. M. Shadman, et al.
Veröffentlicht: (2025)
Potential and Perils of Large Language Models as Judges of Unstructured Textual Data
von: Bedemariam, Rewina, et al.
Veröffentlicht: (2025)
von: Bedemariam, Rewina, et al.
Veröffentlicht: (2025)
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
von: Niu, Tong, et al.
Veröffentlicht: (2024)
von: Niu, Tong, et al.
Veröffentlicht: (2024)
Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
von: Adib, Shefayat E Shams, et al.
Veröffentlicht: (2026)
On Relation-Specific Neurons in Large Language Models
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
von: Liu, Yihong, et al.
Veröffentlicht: (2025)
Judging It, Washing It: Scoring and Greenwashing Corporate Climate Disclosures using Large Language Models
von: Chuang, Marianne, et al.
Veröffentlicht: (2025)
von: Chuang, Marianne, et al.
Veröffentlicht: (2025)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
Exploratory Study into Relations between Cognitive Distortions and Emotional Appraisals
von: Agarwal, Navneet, et al.
Veröffentlicht: (2025)
von: Agarwal, Navneet, et al.
Veröffentlicht: (2025)
Exploring Spatial Schema Intuitions in Large Language and Vision Models
von: Wicke, Philipp, et al.
Veröffentlicht: (2024)
von: Wicke, Philipp, et al.
Veröffentlicht: (2024)
Local Normalization Distortion and the Thermodynamic Formalism of Decoding Strategies for Large Language Models
von: Kempton, Tom, et al.
Veröffentlicht: (2025)
von: Kempton, Tom, et al.
Veröffentlicht: (2025)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
Audio-Aware Large Language Models as Judges for Speaking Styles
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
von: Chiang, Cheng-Han, et al.
Veröffentlicht: (2025)
Locating and Extracting Relational Concepts in Large Language Models
von: Wang, Zijian, et al.
Veröffentlicht: (2024)
von: Wang, Zijian, et al.
Veröffentlicht: (2024)
Tracing Relational Knowledge Recall in Large Language Models
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
von: Popovič, Nicholas, et al.
Veröffentlicht: (2026)
Revisiting Relation Extraction in the era of Large Language Models
von: Wadhwa, Somin, et al.
Veröffentlicht: (2023)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2023)
Evaluating Spatial Understanding of Large Language Models
von: Yamada, Yutaro, et al.
Veröffentlicht: (2023)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2023)
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Relational Database Augmented Large Language Model
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
von: Qin, Zongyue, et al.
Veröffentlicht: (2024)
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector
von: Yang, Haoyan, et al.
Veröffentlicht: (2025)
von: Yang, Haoyan, et al.
Veröffentlicht: (2025)
SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?
von: Lee, Jonggeun, et al.
Veröffentlicht: (2026)
von: Lee, Jonggeun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew
von: Razumenko, Itay, et al.
Veröffentlicht: (2026) -
Auto-Tuning Safety Guardrails for Black-Box Large Language Models
von: Abdulkadir, Perry
Veröffentlicht: (2025) -
Deep Learning Enhanced Road Traffic Analysis: Scalable Vehicle Detection and Velocity Estimation Using PlanetScope Imagery
von: Adamiak, Maciej, et al.
Veröffentlicht: (2024) -
Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025) -
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)