Peeking Behind Closed Doors: Risks of LLM Evaluation by Private Data Curators
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Bansal, Hritik, Maini, Pratyush |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
par: Bansal, Hritik, et autres
Publié: (2023)
par: Bansal, Hritik, et autres
Publié: (2023)
No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
par: Fama, Israel, et autres
Publié: (2024)
par: Fama, Israel, et autres
Publié: (2024)
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
par: Mohamed, Youssef, et autres
Publié: (2024)
par: Mohamed, Youssef, et autres
Publié: (2024)
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
par: Dong, Zhichen, et autres
Publié: (2024)
par: Dong, Zhichen, et autres
Publié: (2024)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
par: Suvarna, Ashima, et autres
Publié: (2026)
par: Suvarna, Ashima, et autres
Publié: (2026)
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning
par: Singhi, Nishad, et autres
Publié: (2025)
par: Singhi, Nishad, et autres
Publié: (2025)
Position: The Most Expensive Part of an LLM should be its Training Data
par: Kandpal, Nikhil, et autres
Publié: (2025)
par: Kandpal, Nikhil, et autres
Publié: (2025)
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
par: Lee, Jaehyeok, et autres
Publié: (2026)
par: Lee, Jaehyeok, et autres
Publié: (2026)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
par: Joshi, Abhinav, et autres
Publié: (2024)
par: Joshi, Abhinav, et autres
Publié: (2024)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
par: Husom, Erik Johannes, et autres
Publié: (2025)
par: Husom, Erik Johannes, et autres
Publié: (2025)
On the Detectability of LLM-Generated Text: What Exactly Is LLM-Generated Text?
par: Geng, Mingmeng, et autres
Publié: (2025)
par: Geng, Mingmeng, et autres
Publié: (2025)
Wikipedia in the Era of LLMs: Evolution and Risks
par: Huang, Siming, et autres
Publié: (2025)
par: Huang, Siming, et autres
Publié: (2025)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
par: Bansal, Hritik, et autres
Publié: (2024)
par: Bansal, Hritik, et autres
Publié: (2024)
Understanding and Mitigating Risks of Generative AI in Financial Services
par: Gehrmann, Sebastian, et autres
Publié: (2025)
par: Gehrmann, Sebastian, et autres
Publié: (2025)
A Multi-LLM Debiasing Framework
par: Owens, Deonna M., et autres
Publié: (2024)
par: Owens, Deonna M., et autres
Publié: (2024)
Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy
par: Tang, Xiangru, et autres
Publié: (2024)
par: Tang, Xiangru, et autres
Publié: (2024)
Towards Best Practices for Open Datasets for LLM Training
par: Baack, Stefan, et autres
Publié: (2025)
par: Baack, Stefan, et autres
Publié: (2025)
Urania: Differentially Private Insights into AI Use
par: Liu, Daogao, et autres
Publié: (2025)
par: Liu, Daogao, et autres
Publié: (2025)
LLM-Assisted Content Conditional Debiasing for Fair Text Embedding
par: Deng, Wenlong, et autres
Publié: (2024)
par: Deng, Wenlong, et autres
Publié: (2024)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
par: Labiad, Ismail, et autres
Publié: (2025)
par: Labiad, Ismail, et autres
Publié: (2025)
Reducing Large Language Model Safety Risks in Women's Health using Semantic Entropy
par: Penny-Dimri, Jahan C., et autres
Publié: (2025)
par: Penny-Dimri, Jahan C., et autres
Publié: (2025)
AI-University: An LLM-based platform for instructional alignment to scientific classrooms
par: Shojaei, Mostafa Faghih, et autres
Publié: (2025)
par: Shojaei, Mostafa Faghih, et autres
Publié: (2025)
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
par: Saha, Rounak, et autres
Publié: (2026)
par: Saha, Rounak, et autres
Publié: (2026)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
par: Schoenegger, Philipp, et autres
Publié: (2024)
par: Schoenegger, Philipp, et autres
Publié: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
par: Do, Heejin, et autres
Publié: (2026)
par: Do, Heejin, et autres
Publié: (2026)
LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences
par: Fu, Zhenxiao, et autres
Publié: (2024)
par: Fu, Zhenxiao, et autres
Publié: (2024)
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
par: Zhang, Andy K., et autres
Publié: (2024)
par: Zhang, Andy K., et autres
Publié: (2024)
An In-Depth Investigation of Data Collection in LLM App Ecosystems
par: Wu, Yuhao, et autres
Publié: (2024)
par: Wu, Yuhao, et autres
Publié: (2024)
LLM Unlearning Without an Expert Curated Dataset
par: Zhu, Xiaoyuan, et autres
Publié: (2025)
par: Zhu, Xiaoyuan, et autres
Publié: (2025)
Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation
par: Wang, Jiawei, et autres
Publié: (2024)
par: Wang, Jiawei, et autres
Publié: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
par: Schoenegger, Philipp, et autres
Publié: (2024)
par: Schoenegger, Philipp, et autres
Publié: (2024)
Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text
par: Islam, Tunazzina
Publié: (2026)
par: Islam, Tunazzina
Publié: (2026)
Evaluating the Performance of ChatGPT for Spam Email Detection
par: Si, Shijing, et autres
Publié: (2024)
par: Si, Shijing, et autres
Publié: (2024)
Leveraging Social Determinants of Health in Alzheimer's Research Using LLM-Augmented Literature Mining and Knowledge Graphs
par: Shang, Tianqi, et autres
Publié: (2024)
par: Shang, Tianqi, et autres
Publié: (2024)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
par: Guldimann, Philipp, et autres
Publié: (2024)
par: Guldimann, Philipp, et autres
Publié: (2024)
Not All Options Are Created Equal: Textual Option Weighting for Token-Efficient LLM-Based Knowledge Tracing
par: Kim, JongWoo, et autres
Publié: (2024)
par: Kim, JongWoo, et autres
Publié: (2024)
ArzEn-LLM: Code-Switched Egyptian Arabic-English Translation and Speech Recognition Using LLMs
par: Heakl, Ahmed, et autres
Publié: (2024)
par: Heakl, Ahmed, et autres
Publié: (2024)
Evaluating the Effectiveness of XAI Techniques for Encoder-Based Language Models
par: Mersha, Melkamu Abay, et autres
Publié: (2025)
par: Mersha, Melkamu Abay, et autres
Publié: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
par: van der Weij, Teun, et autres
Publié: (2024)
par: van der Weij, Teun, et autres
Publié: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
par: Bai, Xiaoyan, et autres
Publié: (2026)
par: Bai, Xiaoyan, et autres
Publié: (2026)
Documents similaires
-
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
par: Bansal, Hritik, et autres
Publié: (2023) -
No Argument Left Behind: Overlapping Chunks for Faster Processing of Arbitrarily Long Legal Texts
par: Fama, Israel, et autres
Publié: (2024) -
No Culture Left Behind: ArtELingo-28, a Benchmark of WikiArt with Captions in 28 Languages
par: Mohamed, Youssef, et autres
Publié: (2024) -
Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey
par: Dong, Zhichen, et autres
Publié: (2024) -
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
par: Suvarna, Ashima, et autres
Publié: (2026)