PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
Fuente:
arXiv
Saved in:
| Main Authors: | Mukhopadhyay, Srija, Reddy, Sathwik, Muthukumar, Shruthi, An, Jisun, Kumaraguru, Ponnurangam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026)
by: Ovcharov, Volodymyr
Published: (2026)
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
by: Paech, Samuel J.
Published: (2023)
by: Paech, Samuel J.
Published: (2023)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
by: Kodali, Prashant, et al.
Published: (2025)
by: Kodali, Prashant, et al.
Published: (2025)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
UrduBench: An Urdu Reasoning Benchmark using Contextually Ensembled Translations with Human-in-the-Loop
by: Shafique, Muhammad Ali, et al.
Published: (2026)
by: Shafique, Muhammad Ali, et al.
Published: (2026)
Critical Insights into Leading Conversational AI Models
by: Kohli, Urja, et al.
Published: (2025)
by: Kohli, Urja, et al.
Published: (2025)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
by: Michail, Andrianos, et al.
Published: (2024)
by: Michail, Andrianos, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
ExpressivityBench: Can LLMs Communicate Implicitly?
by: Tint, Joshua, et al.
Published: (2024)
by: Tint, Joshua, et al.
Published: (2024)
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer
by: Sun, Huashan, et al.
Published: (2023)
by: Sun, Huashan, et al.
Published: (2023)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
by: Yuan, Weikang, et al.
Published: (2025)
by: Yuan, Weikang, et al.
Published: (2025)
Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?
by: Argese, Alex, et al.
Published: (2026)
by: Argese, Alex, et al.
Published: (2026)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
by: Tran, Hung, et al.
Published: (2026)
by: Tran, Hung, et al.
Published: (2026)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
by: Naeem, Numaan, et al.
Published: (2025)
by: Naeem, Numaan, et al.
Published: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
by: Er, Yakup Abrek, et al.
Published: (2025)
by: Er, Yakup Abrek, et al.
Published: (2025)
How to Evaluate Medical AI
by: Kopanichuk, Ilia, et al.
Published: (2025)
by: Kopanichuk, Ilia, et al.
Published: (2025)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
by: Baxi, Rahul
Published: (2025)
by: Baxi, Rahul
Published: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
by: Kytöniemi, Joona, et al.
Published: (2025)
by: Kytöniemi, Joona, et al.
Published: (2025)
TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus
by: Ferreira, Rafael, et al.
Published: (2023)
by: Ferreira, Rafael, et al.
Published: (2023)
GanitBench: A bi-lingual benchmark for evaluating mathematical reasoning in Vision Language Models
by: Bandooni, Ashutosh, et al.
Published: (2025)
by: Bandooni, Ashutosh, et al.
Published: (2025)
Swiss-Bench SBP-002: A Frontier Model Comparison on Swiss Legal and Regulatory Tasks
by: Uenal, Fatih
Published: (2026)
by: Uenal, Fatih
Published: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
by: Adapala, Sai Teja Reddy
Published: (2025)
by: Adapala, Sai Teja Reddy
Published: (2025)
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
by: Nourbakhsh, Aria, et al.
Published: (2026)
by: Nourbakhsh, Aria, et al.
Published: (2026)
LAraBench: Benchmarking Arabic AI with Large Language Models
by: Abdelali, Ahmed, et al.
Published: (2023)
by: Abdelali, Ahmed, et al.
Published: (2023)
Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations
by: Baskar, Sarvesh, et al.
Published: (2025)
by: Baskar, Sarvesh, et al.
Published: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
by: Kim, Soyeon, et al.
Published: (2026)
by: Kim, Soyeon, et al.
Published: (2026)
A Practical Approach for Building Production-Grade Conversational Agents with Workflow Graphs
by: Park, Chiwan, et al.
Published: (2025)
by: Park, Chiwan, et al.
Published: (2025)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
by: Simhi, Adi, et al.
Published: (2026)
by: Simhi, Adi, et al.
Published: (2026)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
by: Le, Van-Truong
Published: (2026)
by: Le, Van-Truong
Published: (2026)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
by: Gokdemir, Ozan, et al.
Published: (2025)
by: Gokdemir, Ozan, et al.
Published: (2025)
StyloAI: Distinguishing AI-Generated Content with Stylometric Analysis
by: Opara, Chidimma
Published: (2024)
by: Opara, Chidimma
Published: (2024)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
by: Schopf, Tim, et al.
Published: (2026)
by: Schopf, Tim, et al.
Published: (2026)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
by: Kautsar, Muhammad Dehan Al, et al.
Published: (2026)
SADAS: A Dialogue Assistant System Towards Remediating Norm Violations in Bilingual Socio-Cultural Conversations
by: Hua, Yuncheng, et al.
Published: (2024)
by: Hua, Yuncheng, et al.
Published: (2024)
Similar Items
-
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
by: Ovcharov, Volodymyr
Published: (2026) -
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024) -
EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
by: Paech, Samuel J.
Published: (2023) -
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)