Back to the Drawing Board for Fair Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pouget, Angéline, Jovanović, Nikola, Vero, Mark, Staab, Robin, Vechev, Martin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Watermark Stealing in Large Language Models
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
A Synthetic Dataset for Personal Attribute Inference
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
Discovering Spoofing Attempts on Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
Ward: Provable RAG Dataset Inference via LLM Watermarks
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)
Watermarking Diffusion Language Models
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
A Unified Framework for LLM Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Beyond Memorization: Violating Privacy Via Inference with Large Language Models
von: Staab, Robin, et al.
Veröffentlicht: (2023)
von: Staab, Robin, et al.
Veröffentlicht: (2023)
Private Attribute Inference from Images with Vision-Language Models
von: Tömekçe, Batuhan, et al.
Veröffentlicht: (2024)
von: Tömekçe, Batuhan, et al.
Veröffentlicht: (2024)
Widening the Gap: Exploiting LLM Quantization via Outlier Injection
von: Zhan, Xiaohua, et al.
Veröffentlicht: (2026)
von: Zhan, Xiaohua, et al.
Veröffentlicht: (2026)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
von: De Muri, Giovanni, et al.
Veröffentlicht: (2025)
von: De Muri, Giovanni, et al.
Veröffentlicht: (2025)
Exploiting LLM Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2024)
Fewer Weights, More Problems: A Practical Attack on LLM Pruning
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
Mind the Gap: A Practical Attack on GGUF Quantization
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2025)
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
von: Egashira, Kazuki, et al.
Veröffentlicht: (2026)
von: Egashira, Kazuki, et al.
Veröffentlicht: (2026)
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
von: Guldimann, Philipp, et al.
Veröffentlicht: (2024)
Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2026)
Black-Box Detection of Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024)
LLM Fingerprinting via Semantically Conditioned Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
Towards Watermarking of Open-Source LLMs
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025)
BaxBench: Can LLMs Generate Correct and Secure Backends?
von: Vero, Mark, et al.
Veröffentlicht: (2025)
von: Vero, Mark, et al.
Veröffentlicht: (2025)
Large Language Models are Advanced Anonymizers
von: Staab, Robin, et al.
Veröffentlicht: (2024)
von: Staab, Robin, et al.
Veröffentlicht: (2024)
Instruction Tuning for Secure Code Generation
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
von: He, Jingxuan, et al.
Veröffentlicht: (2024)
Suitability Filter: A Statistical Framework for Classifier Evaluation in Real-World Deployment Settings
von: Pouget, Angéline, et al.
Veröffentlicht: (2025)
von: Pouget, Angéline, et al.
Veröffentlicht: (2025)
AutoBaxBuilder: Bootstrapping Code Security Benchmarking
von: von Arx, Tobias, et al.
Veröffentlicht: (2025)
von: von Arx, Tobias, et al.
Veröffentlicht: (2025)
Watermarking Autoregressive Image Generation
von: Jovanović, Nikola, et al.
Veröffentlicht: (2025)
von: Jovanović, Nikola, et al.
Veröffentlicht: (2025)
SoK: Data Minimization in Machine Learning
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
MixAT: Combining Continuous and Discrete Adversarial Training for LLMs
von: Dékány, Csaba, et al.
Veröffentlicht: (2025)
von: Dékány, Csaba, et al.
Veröffentlicht: (2025)
Honeyval: A Comprehensive Evaluation Framework for LLM-powered HTTP Honeypots
von: Vero, Mark, et al.
Veröffentlicht: (2026)
von: Vero, Mark, et al.
Veröffentlicht: (2026)
Learning Compact Boolean Networks
von: Wang, Shengpu, et al.
Veröffentlicht: (2026)
von: Wang, Shengpu, et al.
Veröffentlicht: (2026)
CuTS: Customizable Tabular Synthetic Data Generation
von: Vero, Mark, et al.
Veröffentlicht: (2023)
von: Vero, Mark, et al.
Veröffentlicht: (2023)
Understanding Certified Training with Interval Bound Propagation
von: Mao, Yuhao, et al.
Veröffentlicht: (2023)
von: Mao, Yuhao, et al.
Veröffentlicht: (2023)
CTBENCH: A Library and Benchmark for Certified Training
von: Mao, Yuhao, et al.
Veröffentlicht: (2024)
von: Mao, Yuhao, et al.
Veröffentlicht: (2024)
Expressiveness of Multi-Neuron Convex Relaxations in Neural Network Certification
von: Mao, Yuhao, et al.
Veröffentlicht: (2024)
von: Mao, Yuhao, et al.
Veröffentlicht: (2024)
Dual Randomized Smoothing: Beyond Global Noise Variance
von: Sun, Chenhao, et al.
Veröffentlicht: (2025)
von: Sun, Chenhao, et al.
Veröffentlicht: (2025)
Differential Adjusted Parity for Learning Fair Representations
von: Sahyouni, Bucher, et al.
Veröffentlicht: (2025)
von: Sahyouni, Bucher, et al.
Veröffentlicht: (2025)
Generating Synthetic Fair Syntax-agnostic Data by Learning and Distilling Fair Representation
von: Sikder, Md Fahim, et al.
Veröffentlicht: (2024)
von: Sikder, Md Fahim, et al.
Veröffentlicht: (2024)
SWT-Bench: Testing and Validating Real-World Bug-Fixes with Code Agents
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
von: Mündler, Niels, et al.
Veröffentlicht: (2024)
Adaptive Generation of Bias-Eliciting Questions for LLMs
von: Staab, Robin, et al.
Veröffentlicht: (2025)
von: Staab, Robin, et al.
Veröffentlicht: (2025)
Fair CCA for Fair Representation Learning: An ADNI Study
von: Hou, Bojian, et al.
Veröffentlicht: (2025)
von: Hou, Bojian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Watermark Stealing in Large Language Models
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024) -
A Synthetic Dataset for Personal Attribute Inference
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024) -
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2025) -
Discovering Spoofing Attempts on Language Model Watermarks
von: Gloaguen, Thibaud, et al.
Veröffentlicht: (2024) -
Ward: Provable RAG Dataset Inference via LLM Watermarks
von: Jovanović, Nikola, et al.
Veröffentlicht: (2024)