Data Contamination Can Cross Language Barriers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Feng, Zhuang, Yufan, Sun, Zihao, Xu, Sunan, Kumar, Animesh, Shang, Jingbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2025)
von: Tian, Runchu, et al.
Veröffentlicht: (2025)
Vector-ICL: In-context Learning with Continuous Vector Representations
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency Graph
von: Gao, Xiaochen Kev, et al.
Veröffentlicht: (2024)
von: Gao, Xiaochen Kev, et al.
Veröffentlicht: (2024)
Learning a Decision Tree Algorithm with Transformers
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)
Self-Taught Agentic Long Context Understanding
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025)
How Much Can We Forget about Data Contamination?
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
Sensitivity of Small Language Models to Fine-tuning Data Contamination
von: Scaria, Nicy, et al.
Veröffentlicht: (2025)
von: Scaria, Nicy, et al.
Veröffentlicht: (2025)
An Open Source Data Contamination Report for Large Language Models
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Investigating Data Contamination in Modern Benchmarks for Large Language Models
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
von: Deng, Chunyuan, et al.
Veröffentlicht: (2023)
READ: Improving Relation Extraction from an ADversarial Perspective
von: Li, Dawei, et al.
Veröffentlicht: (2024)
von: Li, Dawei, et al.
Veröffentlicht: (2024)
Debug like a Human: A Large Language Model Debugger via Verifying Runtime Execution Step-by-step
von: Zhong, Li, et al.
Veröffentlicht: (2024)
von: Zhong, Li, et al.
Veröffentlicht: (2024)
Training Language Models to Generate Quality Code with Program Analysis Feedback
von: Yao, Feng, et al.
Veröffentlicht: (2025)
von: Yao, Feng, et al.
Veröffentlicht: (2025)
Can Language Models Follow Multiple Turns of Entangled Instructions?
von: Han, Chi, et al.
Veröffentlicht: (2025)
von: Han, Chi, et al.
Veröffentlicht: (2025)
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2023)
Next-Token Prediction Task Assumes Optimal Data Ordering for LLM Training in Proof Generation
von: An, Chenyang, et al.
Veröffentlicht: (2024)
von: An, Chenyang, et al.
Veröffentlicht: (2024)
Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
von: Tong, Ziyi, et al.
Veröffentlicht: (2026)
Investigating Data Contamination for Pre-training Language Models
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
von: Jiang, Minhao, et al.
Veröffentlicht: (2024)
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
RecycleGPT: An Autoregressive Language Model with Recyclable Module
von: Jiang, Yufan, et al.
Veröffentlicht: (2023)
von: Jiang, Yufan, et al.
Veröffentlicht: (2023)
Towards Data Contamination Detection for Modern Large Language Models: Limitations, Inconsistencies, and Oracle Challenges
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
von: Samuel, Vinay, et al.
Veröffentlicht: (2024)
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
von: Fan, Yang
Veröffentlicht: (2025)
von: Fan, Yang
Veröffentlicht: (2025)
When is the consistent prediction likely to be a correct prediction?
von: Nguyen, Alex, et al.
Veröffentlicht: (2024)
von: Nguyen, Alex, et al.
Veröffentlicht: (2024)
Can Language Models Solve Olympiad Programming?
von: Shi, Quan, et al.
Veröffentlicht: (2024)
von: Shi, Quan, et al.
Veröffentlicht: (2024)
Navigating the Dual Facets: A Comprehensive Evaluation of Sequential Memory Editing in Large Language Models
von: Lin, Zihao, et al.
Veröffentlicht: (2024)
von: Lin, Zihao, et al.
Veröffentlicht: (2024)
Large Language Models for Time Series: A Survey
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xiyuan, et al.
Veröffentlicht: (2024)
LatestEval: Addressing Data Contamination in Language Model Evaluation through Dynamic and Time-Sensitive Test Construction
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
von: Li, Yucheng, et al.
Veröffentlicht: (2023)
Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
Reading between the Lines: Can LLMs Identify Cross-Cultural Communication Gaps?
von: Saha, Sougata, et al.
Veröffentlicht: (2025)
von: Saha, Sougata, et al.
Veröffentlicht: (2025)
Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations
von: Banerjee, Somnath, et al.
Veröffentlicht: (2025)
von: Banerjee, Somnath, et al.
Veröffentlicht: (2025)
Obscuring Data Contamination Through Translation: Evidence from Arabic Corpora
von: Abbas, Chaymaa, et al.
Veröffentlicht: (2026)
von: Abbas, Chaymaa, et al.
Veröffentlicht: (2026)
Can Language Models Solve Graph Problems in Natural Language?
von: Wang, Heng, et al.
Veröffentlicht: (2023)
von: Wang, Heng, et al.
Veröffentlicht: (2023)
No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
von: Yang, Tiankai, et al.
Veröffentlicht: (2026)
Model-diff: A Tool for Comparative Study of Language Models in the Input Space
von: Liu, Weitang, et al.
Veröffentlicht: (2024)
von: Liu, Weitang, et al.
Veröffentlicht: (2024)
Cross-Lingual Knowledge Editing in Large Language Models
von: Wang, Jiaan, et al.
Veröffentlicht: (2023)
von: Wang, Jiaan, et al.
Veröffentlicht: (2023)
Fine-tuning vs Prompting, Can Language Models Understand Human Values?
von: Sun, Pingwei
Veröffentlicht: (2024)
von: Sun, Pingwei
Veröffentlicht: (2024)
Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models
von: Xu, Zihao, et al.
Veröffentlicht: (2026)
von: Xu, Zihao, et al.
Veröffentlicht: (2026)
Foundations of Large Language Models
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
von: Balloccu, Simone, et al.
Veröffentlicht: (2024)
Towards Explainable Temporal Reasoning in Large Language Models: A Structure-Aware Generative Framework
von: Jiang, Zihao, et al.
Veröffentlicht: (2025)
von: Jiang, Zihao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Finish First, Perfect Later: Test-Time Token-Level Cross-Validation for Diffusion Large Language Models
von: Tian, Runchu, et al.
Veröffentlicht: (2025) -
Vector-ICL: In-context Learning with Continuous Vector Representations
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024) -
Text Generation Beyond Discrete Token Sampling
von: Zhuang, Yufan, et al.
Veröffentlicht: (2025) -
Beyond Scaling: Predicting Patent Approval with Domain-specific Fine-grained Claim Dependency Graph
von: Gao, Xiaochen Kev, et al.
Veröffentlicht: (2024) -
Learning a Decision Tree Algorithm with Transformers
von: Zhuang, Yufan, et al.
Veröffentlicht: (2024)