TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rajore, Tanmay, Chandran, Nishanth, Sitaram, Sunayana, Gupta, Divya, Sharma, Rahul, Mittal, Kashish, Swaminathan, Manohar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enterprise AI Must Enforce Participant-Aware Access Control
by: Bhatt, Shashank Shreedhar, et al.
Published: (2025)
by: Bhatt, Shashank Shreedhar, et al.
Published: (2025)
TRUCE: TRUsted Compliance Enforcement Service for Secure Health Data Exchange
by: Kim, Dae-young, et al.
Published: (2025)
by: Kim, Dae-young, et al.
Published: (2025)
TrustRate: A Decentralized Platform for Hijack-Resistant Anonymous Reviews
by: Dwivedula, Rohit, et al.
Published: (2024)
by: Dwivedula, Rohit, et al.
Published: (2024)
Redefining Website Fingerprinting Attacks With Multiagent LLMs
by: Song, Chuxu, et al.
Published: (2025)
by: Song, Chuxu, et al.
Published: (2025)
Supporting Socially Constrained Private Communications with SecureWhispers
by: Khandkar, Vinod, et al.
Published: (2025)
by: Khandkar, Vinod, et al.
Published: (2025)
Private, Efficient and Scalable Kernel Learning for Medical Image Analysis
by: Hannemann, Anika, et al.
Published: (2024)
by: Hannemann, Anika, et al.
Published: (2024)
Benchmarking Differentially Private Tabular Data Synthesis
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
Benchmarking Fraud Detectors on Private Graph Data
by: Goldberg, Alexander, et al.
Published: (2025)
by: Goldberg, Alexander, et al.
Published: (2025)
Detecting Benchmark Contamination Through Watermarking
by: Sander, Tom, et al.
Published: (2025)
by: Sander, Tom, et al.
Published: (2025)
An Attack to Break Permutation-Based Private Third-Party Inference Schemes for LLMs
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
Private Aggregate Queries to Untrusted Databases
by: Hafiz, Syed Mahbub, et al.
Published: (2024)
by: Hafiz, Syed Mahbub, et al.
Published: (2024)
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
by: Li, Junhao, et al.
Published: (2025)
by: Li, Junhao, et al.
Published: (2025)
Exploiting Page Faults for Covert Communication
by: Swaminathan, Sathvik
Published: (2025)
by: Swaminathan, Sathvik
Published: (2025)
GraMFedDHAR: Graph Based Multimodal Differentially Private Federated HAR
by: Halder, Labani, et al.
Published: (2025)
by: Halder, Labani, et al.
Published: (2025)
Improving Parameter-Efficient Federated Learning with Differentially Private Refactorization
by: Tran, Linh, et al.
Published: (2026)
by: Tran, Linh, et al.
Published: (2026)
DP-FlogTinyLLM: Differentially private federated log anomaly detection using Tiny LLMs
by: Thompson, Isaiah, et al.
Published: (2026)
by: Thompson, Isaiah, et al.
Published: (2026)
MAUI: Reconstructing Private Client Data in Federated Transfer Learning
by: Dabholkar, Ahaan, et al.
Published: (2025)
by: Dabholkar, Ahaan, et al.
Published: (2025)
ModelForge: Using GenAI to Improve the Development of Security Protocols
by: Duclos, Martin, et al.
Published: (2025)
by: Duclos, Martin, et al.
Published: (2025)
On Evaluating the Durability of Safeguards for Open-Weight LLMs
by: Qi, Xiangyu, et al.
Published: (2024)
by: Qi, Xiangyu, et al.
Published: (2024)
Entropy-Guided Attention for Private LLMs
by: Jha, Nandan Kumar, et al.
Published: (2025)
by: Jha, Nandan Kumar, et al.
Published: (2025)
Locally Differentially Private Embedding Models in Distributed Fraud Prevention Systems
by: Perez, Iker, et al.
Published: (2024)
by: Perez, Iker, et al.
Published: (2024)
Teach LLMs to Phish: Stealing Private Information from Language Models
by: Panda, Ashwinee, et al.
Published: (2024)
by: Panda, Ashwinee, et al.
Published: (2024)
Inferentially-Private Private Information
by: Wang, Shuaiqi, et al.
Published: (2024)
by: Wang, Shuaiqi, et al.
Published: (2024)
How to Privately Tune Hyperparameters in Federated Learning? Insights from a Benchmark Study
by: Mitic, Natalija, et al.
Published: (2024)
by: Mitic, Natalija, et al.
Published: (2024)
Benchmarking Private Population Data Release Mechanisms: Synthetic Data vs. TopDown
by: Maddi, Aadyaa, et al.
Published: (2024)
by: Maddi, Aadyaa, et al.
Published: (2024)
Towards a Benchmark for Dependency Decision-Making
by: Singla, Tanmay, et al.
Published: (2026)
by: Singla, Tanmay, et al.
Published: (2026)
In Specs we Trust? Conformance-Analysis of Implementation to Specifications in Node-RED and Associated Security Risks
by: Schneider, Simon, et al.
Published: (2025)
by: Schneider, Simon, et al.
Published: (2025)
LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks
by: Ullah, Saad, et al.
Published: (2023)
by: Ullah, Saad, et al.
Published: (2023)
Risks to Zero Trust in a Federated Mission Partner Environment
by: Strandell, Keith, et al.
Published: (2022)
by: Strandell, Keith, et al.
Published: (2022)
When FinTech Meets Privacy: Securing Financial LLMs with Differential Private Fine-Tuning
by: Zhu, Sichen, et al.
Published: (2025)
by: Zhu, Sichen, et al.
Published: (2025)
Differentially Private High-dimensional Variable Selection via Integer Programming
by: Prastakos, Petros, et al.
Published: (2025)
by: Prastakos, Petros, et al.
Published: (2025)
Enhancing MOTION2NX for Efficient, Scalable and Secure Image Inference using Convolutional Neural Networks
by: K, Haritha, et al.
Published: (2024)
by: K, Haritha, et al.
Published: (2024)
DPDSyn: Improving Differentially Private Dataset Synthesis for Model Training by Downstream Task Guidance
by: Jia, Mingxuan, et al.
Published: (2026)
by: Jia, Mingxuan, et al.
Published: (2026)
CTIBench: A Benchmark for Evaluating LLMs in Cyber Threat Intelligence
by: Alam, Md Tanvirul, et al.
Published: (2024)
by: Alam, Md Tanvirul, et al.
Published: (2024)
Improving Count-Mean Sketch as the Leading Locally Differentially Private Frequency Estimator for Large Dictionaries
by: Pan, Mingen
Published: (2024)
by: Pan, Mingen
Published: (2024)
On The Fragility of Benchmark Contamination Detection in Reasoning Models
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
LLM Benchmark Datasets Should Be Contamination-Resistant
by: Al-Lawati, Ali, et al.
Published: (2026)
by: Al-Lawati, Ali, et al.
Published: (2026)
Cascade: Token-Sharded Private LLM Inference
by: Thomas, Rahul, et al.
Published: (2025)
by: Thomas, Rahul, et al.
Published: (2025)
Private Map-Secure Reduce: Infrastructure for Efficient AI Data Markets
by: Wagh, Sameer, et al.
Published: (2025)
by: Wagh, Sameer, et al.
Published: (2025)
From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection
by: Han, Junxiao, et al.
Published: (2025)
by: Han, Junxiao, et al.
Published: (2025)
Similar Items
-
Enterprise AI Must Enforce Participant-Aware Access Control
by: Bhatt, Shashank Shreedhar, et al.
Published: (2025) -
TRUCE: TRUsted Compliance Enforcement Service for Secure Health Data Exchange
by: Kim, Dae-young, et al.
Published: (2025) -
TrustRate: A Decentralized Platform for Hijack-Resistant Anonymous Reviews
by: Dwivedula, Rohit, et al.
Published: (2024) -
Redefining Website Fingerprinting Attacks With Multiagent LLMs
by: Song, Chuxu, et al.
Published: (2025) -
Supporting Socially Constrained Private Communications with SecureWhispers
by: Khandkar, Vinod, et al.
Published: (2025)