Silicon Bureaucracy and AI Test-Oriented Education: Contamination Sensitivity and Score Confidence in LLM Benchmarks
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Yiliang, An, Hongjun, Chen, Jiangan, Yan, Xuanchen, Song, Huan, Shao, Jiawei, Li, Xuelong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
por: An, Hongjun, et al.
Publicado: (2026)
por: An, Hongjun, et al.
Publicado: (2026)
CreditAudit: 2$^\text{nd}$ Dimension for LLM Evaluation and Selection
por: Song, Yiliang, et al.
Publicado: (2026)
por: Song, Yiliang, et al.
Publicado: (2026)
Ruyi2 Technical Report
por: Song, Huan, et al.
Publicado: (2026)
por: Song, Huan, et al.
Publicado: (2026)
Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence
por: An, Hongjun, et al.
Publicado: (2026)
por: An, Hongjun, et al.
Publicado: (2026)
Physics in Next-token Prediction
por: An, Hongjun, et al.
Publicado: (2024)
por: An, Hongjun, et al.
Publicado: (2024)
Theoretical Foundations of Scaling Law in Familial Models
por: Song, Huan, et al.
Publicado: (2025)
por: Song, Huan, et al.
Publicado: (2025)
PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models
por: Zhang, Huixuan, et al.
Publicado: (2024)
por: Zhang, Huixuan, et al.
Publicado: (2024)
AI Flow: Perspectives, Scenarios, and Approaches
por: An, Hongjun, et al.
Publicado: (2025)
por: An, Hongjun, et al.
Publicado: (2025)
SentenceVAE: Enable Next-sentence Prediction for Large Language Models with Faster Speed, Higher Accuracy and Longer Context
por: An, Hongjun, et al.
Publicado: (2024)
por: An, Hongjun, et al.
Publicado: (2024)
Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compression
por: Yuan, Cheng, et al.
Publicado: (2025)
por: Yuan, Cheng, et al.
Publicado: (2025)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
por: Sun, Yifan, et al.
Publicado: (2025)
por: Sun, Yifan, et al.
Publicado: (2025)
ScRPO: From Errors to Insights
por: Li, Lianrui, et al.
Publicado: (2025)
por: Li, Lianrui, et al.
Publicado: (2025)
Preference Leakage: A Contamination Problem in LLM-as-a-judge
por: Li, Dawei, et al.
Publicado: (2025)
por: Li, Dawei, et al.
Publicado: (2025)
Enhancing Marker Scoring Accuracy through Ordinal Confidence Modelling in Educational Assessments
por: Chakravarty, Abhirup, et al.
Publicado: (2025)
por: Chakravarty, Abhirup, et al.
Publicado: (2025)
Confident Rankings with Fewer Items: Adaptive LLM Evaluation with Continuous Scores
por: Balkır, Esma, et al.
Publicado: (2026)
por: Balkır, Esma, et al.
Publicado: (2026)
Confident RAG: Enhancing the Performance of LLMs for Mathematics Question Answering through Multi-Embedding and Confidence Scoring
por: Chen, Shiting, et al.
Publicado: (2025)
por: Chen, Shiting, et al.
Publicado: (2025)
Benchmark Test-Time Scaling of General LLM Agents
por: Li, Xiaochuan, et al.
Publicado: (2026)
por: Li, Xiaochuan, et al.
Publicado: (2026)
LiveBench: A Challenging, Contamination-Limited LLM Benchmark
por: White, Colin, et al.
Publicado: (2024)
por: White, Colin, et al.
Publicado: (2024)
Confident in a Confidence Score: Investigating the Sensitivity of Confidence Scores to Supervised Fine-Tuning
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2026)
por: Flores, Lorenzo Jaime Yu, et al.
Publicado: (2026)
AI Flow at the Network Edge
por: Shao, Jiawei, et al.
Publicado: (2024)
por: Shao, Jiawei, et al.
Publicado: (2024)
Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
por: Kostić, Bogdan, et al.
Publicado: (2026)
por: Kostić, Bogdan, et al.
Publicado: (2026)
Effectively Steer LLM To Follow Preference via Building Confident Directions
por: Song, Bingqing, et al.
Publicado: (2025)
por: Song, Bingqing, et al.
Publicado: (2025)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
por: Li, Yucheng, et al.
Publicado: (2023)
por: Li, Yucheng, et al.
Publicado: (2023)
Pipeline Parallelism is All You Need for Optimized Early-Exit Based Self-Speculative Decoding
por: Li, Ruanjun, et al.
Publicado: (2025)
por: Li, Ruanjun, et al.
Publicado: (2025)
Self-Assessment Tests are Unreliable Measures of LLM Personality
por: Gupta, Akshat, et al.
Publicado: (2023)
por: Gupta, Akshat, et al.
Publicado: (2023)
Confidence-aware Self-Semantic Distillation on Knowledge Graph Embedding
por: Liu, Yichen, et al.
Publicado: (2022)
por: Liu, Yichen, et al.
Publicado: (2022)
PredictaBoard: Benchmarking LLM Score Predictability
por: Pacchiardi, Lorenzo, et al.
Publicado: (2025)
por: Pacchiardi, Lorenzo, et al.
Publicado: (2025)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
por: Aksoy, Sinan G., et al.
Publicado: (2026)
por: Aksoy, Sinan G., et al.
Publicado: (2026)
Sensitivity of Small Language Models to Fine-tuning Data Contamination
por: Scaria, Nicy, et al.
Publicado: (2025)
por: Scaria, Nicy, et al.
Publicado: (2025)
Knowledge Distillation of LLM for Automatic Scoring of Science Education Assessments
por: Latif, Ehsan, et al.
Publicado: (2023)
por: Latif, Ehsan, et al.
Publicado: (2023)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
por: Song, Dingjie, et al.
Publicado: (2024)
por: Song, Dingjie, et al.
Publicado: (2024)
LLM-Confidence Reranker: A Training-Free Approach for Enhancing Retrieval-Augmented Generation Systems
por: Song, Zhipeng, et al.
Publicado: (2026)
por: Song, Zhipeng, et al.
Publicado: (2026)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
por: Pathak, Manas, et al.
Publicado: (2026)
por: Pathak, Manas, et al.
Publicado: (2026)
Contextualized Sequence Likelihood: Enhanced Confidence Scores for Natural Language Generation
por: Lin, Zhen, et al.
Publicado: (2024)
por: Lin, Zhen, et al.
Publicado: (2024)
Screen Before You Interpret: A Portable Validity Protocol for Benchmark-Based LLM Confidence Signals
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
por: Li, Peiyu, et al.
Publicado: (2025)
por: Li, Peiyu, et al.
Publicado: (2025)
TripScore: Benchmarking and rewarding real-world travel planning with fine-grained evaluation
por: Qu, Yincen, et al.
Publicado: (2025)
por: Qu, Yincen, et al.
Publicado: (2025)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
por: Wang, Guanghui, et al.
Publicado: (2025)
por: Wang, Guanghui, et al.
Publicado: (2025)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
por: Tang, Zeyu, et al.
Publicado: (2026)
por: Tang, Zeyu, et al.
Publicado: (2026)
HoneyComb: A Flexible LLM-Based Agent System for Materials Science
por: Zhang, Huan, et al.
Publicado: (2024)
por: Zhang, Huan, et al.
Publicado: (2024)
Ejemplares similares
-
Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity
por: An, Hongjun, et al.
Publicado: (2026) -
CreditAudit: 2$^\text{nd}$ Dimension for LLM Evaluation and Selection
por: Song, Yiliang, et al.
Publicado: (2026) -
Ruyi2 Technical Report
por: Song, Huan, et al.
Publicado: (2026) -
Single-Pixel Vision-Language Model for Intrinsic Privacy-Preserving Behavioral Intelligence
por: An, Hongjun, et al.
Publicado: (2026) -
Physics in Next-token Prediction
por: An, Hongjun, et al.
Publicado: (2024)