Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
Fuente:
arXiv
Saved in:
| Main Authors: | Aksoy, Sinan G., Sabrio, Alexandra A., VonKaenel, Erik, Burke, Lee |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
by: Ferdousi, Rahatara, et al.
Published: (2025)
by: Ferdousi, Rahatara, et al.
Published: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025)
by: Štefánik, Michal, et al.
Published: (2025)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024)
by: Kornilov, Albert, et al.
Published: (2024)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026)
by: Yu, Zhiqi, et al.
Published: (2026)
Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
by: Souverain, Thomas
Published: (2025)
by: Souverain, Thomas
Published: (2025)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
by: Chen, Jiaju, et al.
Published: (2025)
by: Chen, Jiaju, et al.
Published: (2025)
Survey of Swarm Intelligence Approaches to Search Documents Based On Semantic Similarity
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
by: Shin, Suyeon, et al.
Published: (2024)
by: Shin, Suyeon, et al.
Published: (2024)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
Low-Resource Neural Machine Translation Using Recurrent Neural Networks and Transfer Learning: A Case Study on English-to-Igbo
by: Ekle, Ocheme Anthony, et al.
Published: (2025)
by: Ekle, Ocheme Anthony, et al.
Published: (2025)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
ERASMO: Leveraging Large Language Models for Enhanced Clustering Segmentation
by: Silva, Fillipe dos Santos, et al.
Published: (2024)
by: Silva, Fillipe dos Santos, et al.
Published: (2024)
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight
by: Bensch, Oliver, et al.
Published: (2024)
by: Bensch, Oliver, et al.
Published: (2024)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
by: Otal, Hakan T., et al.
Published: (2024)
by: Otal, Hakan T., et al.
Published: (2024)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
by: Abhishek, Alok, et al.
Published: (2025)
by: Abhishek, Alok, et al.
Published: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
by: Abhishek, Alok, et al.
Published: (2026)
by: Abhishek, Alok, et al.
Published: (2026)
TSDS: Data Selection for Task-Specific Model Finetuning
by: Liu, Zifan, et al.
Published: (2024)
by: Liu, Zifan, et al.
Published: (2024)
Reasoning Promotes Robustness in Theory of Mind Tasks
by: de Haan, Ian B., et al.
Published: (2026)
by: de Haan, Ian B., et al.
Published: (2026)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
by: Yoo, Seunghyun
Published: (2025)
by: Yoo, Seunghyun
Published: (2025)
AI Assistants for Spaceflight Procedures: Combining Generative Pre-Trained Transformer and Retrieval-Augmented Generation on Knowledge Graphs With Augmented Reality Cues
by: Bensch, Oliver, et al.
Published: (2024)
by: Bensch, Oliver, et al.
Published: (2024)
Morphological Synthesizer for Ge'ez Language: Addressing Morphological Complexity and Resource Limitations
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
by: Gebremariam, Gebrearegawi, et al.
Published: (2025)
Survey of Genetic and Differential Evolutionary Algorithm Approaches to Search Documents Based On Semantic Similarity
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
by: Muniyappa, Chandrashekar, et al.
Published: (2025)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
by: Arabov, Mullosharaf K.
Published: (2026)
by: Arabov, Mullosharaf K.
Published: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
by: Yao, Huaiyuan, et al.
Published: (2026)
by: Yao, Huaiyuan, et al.
Published: (2026)
Acceptable Use Policies for Foundation Models
by: Klyman, Kevin
Published: (2024)
by: Klyman, Kevin
Published: (2024)
The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
by: Johnson, Warren
Published: (2026)
by: Johnson, Warren
Published: (2026)
Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
by: Apostolopoulou, Alexandra, et al.
Published: (2025)
From Surface to Semantics: Semantic Structure Parsing for Table-Centric Document Analysis
by: Li, Xuan, et al.
Published: (2025)
by: Li, Xuan, et al.
Published: (2025)
ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation
by: Livieris, Ioannis E., et al.
Published: (2026)
by: Livieris, Ioannis E., et al.
Published: (2026)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
by: Consoli, Sergio, et al.
Published: (2025)
by: Consoli, Sergio, et al.
Published: (2025)
OptPO: Optimal Rollout Allocation for Test-time Policy Optimization
by: Wang, Youkang, et al.
Published: (2025)
by: Wang, Youkang, et al.
Published: (2025)
A transfer learning approach for automatic conflicts detection in software requirement sentence pairs based on dual encoders
by: Wang, Yizheng, et al.
Published: (2025)
by: Wang, Yizheng, et al.
Published: (2025)
Surfing the modeling of PoS taggers in low-resource scenarios
by: Ferro, Manuel Vilares, et al.
Published: (2024)
by: Ferro, Manuel Vilares, et al.
Published: (2024)
An alternative formulation of attention pooling function in translation
by: Conti, Eddie
Published: (2024)
by: Conti, Eddie
Published: (2024)
Generic Embedding-Based Lexicons for Transparent and Reproducible Text Scoring
by: Moez, Catherine
Published: (2024)
by: Moez, Catherine
Published: (2024)
Large Language Models are Inconsistent and Biased Evaluators
by: Stureborg, Rickard, et al.
Published: (2024)
by: Stureborg, Rickard, et al.
Published: (2024)
Triplètoile: Extraction of Knowledge from Microblogging Text
by: Zavarella, Vanni, et al.
Published: (2024)
by: Zavarella, Vanni, et al.
Published: (2024)
Epidemic Information Extraction for Event-Based Surveillance using Large Language Models
by: Consoli, Sergio, et al.
Published: (2024)
by: Consoli, Sergio, et al.
Published: (2024)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
by: Kim, Jaemin, et al.
Published: (2024)
by: Kim, Jaemin, et al.
Published: (2024)
Similar Items
-
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
by: Ferdousi, Rahatara, et al.
Published: (2025) -
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
by: Štefánik, Michal, et al.
Published: (2025) -
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
by: Kornilov, Albert, et al.
Published: (2024) -
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
by: Yu, Zhiqi, et al.
Published: (2026) -
Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
by: Souverain, Thomas
Published: (2025)