A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
Fuente:
arXiv
Guardado en:
| Autores principales: | Kaneko, Masahiro, Baldwin, Timothy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
por: Kaneko, Masahiro, et al.
Publicado: (2025)
por: Kaneko, Masahiro, et al.
Publicado: (2025)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
por: Kaneko, Masahiro, et al.
Publicado: (2025)
por: Kaneko, Masahiro, et al.
Publicado: (2025)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
por: Kaneko, Masahiro, et al.
Publicado: (2024)
por: Kaneko, Masahiro, et al.
Publicado: (2024)
Eagle: Ethical Dataset Given from Real Interactions
por: Kaneko, Masahiro, et al.
Publicado: (2024)
por: Kaneko, Masahiro, et al.
Publicado: (2024)
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
por: Stuart, Harry, et al.
Publicado: (2026)
por: Stuart, Harry, et al.
Publicado: (2026)
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
por: Kaneko, Masahiro, et al.
Publicado: (2024)
por: Kaneko, Masahiro, et al.
Publicado: (2024)
Online Learning Defense against Iterative Jailbreak Attacks via Prompt Optimization
por: Kaneko, Masahiro, et al.
Publicado: (2025)
por: Kaneko, Masahiro, et al.
Publicado: (2025)
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
por: Kaneko, Masahiro, et al.
Publicado: (2026)
por: Kaneko, Masahiro, et al.
Publicado: (2026)
In-Contextual Gender Bias Suppression for Large Language Models
por: Oba, Daisuke, et al.
Publicado: (2023)
por: Oba, Daisuke, et al.
Publicado: (2023)
On the Alignment of Large Language Models with Global Human Opinion
por: Liu, Yang, et al.
Publicado: (2025)
por: Liu, Yang, et al.
Publicado: (2025)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
por: Hida, Rem, et al.
Publicado: (2024)
por: Hida, Rem, et al.
Publicado: (2024)
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2025)
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2025)
Psychometric Predictive Power of Large Language Models
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
Garbage Attention in Large Language Models: BOS Sink Heads and Sink-aware Pruning
por: Sok, Jaewon, et al.
Publicado: (2026)
por: Sok, Jaewon, et al.
Publicado: (2026)
Loose LIPS Sink Ships: Asking Questions in Battleship with Language-Informed Program Sampling
por: Grand, Gabriel, et al.
Publicado: (2024)
por: Grand, Gabriel, et al.
Publicado: (2024)
Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2023)
por: Anantaprayoon, Panatchakorn, et al.
Publicado: (2023)
First Finish Search: Efficient Test-Time Scaling in Large Language Models
por: Agarwal, Aradhye, et al.
Publicado: (2025)
por: Agarwal, Aradhye, et al.
Publicado: (2025)
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
por: Luo, Jiayun, et al.
Publicado: (2025)
por: Luo, Jiayun, et al.
Publicado: (2025)
Benchmarking Gender and Political Bias in Large Language Models
por: Yang, Jinrui, et al.
Publicado: (2025)
por: Yang, Jinrui, et al.
Publicado: (2025)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
por: Oi, Masanari, et al.
Publicado: (2024)
por: Oi, Masanari, et al.
Publicado: (2024)
Paraphrasing Adversarial Attack on LLM-as-a-Reviewer
por: Kaneko, Masahiro
Publicado: (2026)
por: Kaneko, Masahiro
Publicado: (2026)
Large Language Models Are Human-Like Internally
por: Kuribayashi, Tatsuki, et al.
Publicado: (2025)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2025)
Towards Transparent AI: A Survey on Explainable Large Language Models
por: Palikhe, Avash, et al.
Publicado: (2025)
por: Palikhe, Avash, et al.
Publicado: (2025)
Uncertainty Quantification for Large Language Diffusion Models
por: Vazhentsev, Artem, et al.
Publicado: (2026)
por: Vazhentsev, Artem, et al.
Publicado: (2026)
Leaking LoRa: An Evaluation of Password Leaks and Knowledge Storage in Large Language Models
por: Marinelli, Ryan, et al.
Publicado: (2025)
por: Marinelli, Ryan, et al.
Publicado: (2025)
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
por: Kuribayashi, Tatsuki, et al.
Publicado: (2023)
CTR-Sink: Attention Sink for Language Models in Click-Through Rate Prediction
por: Li, Zixuan, et al.
Publicado: (2025)
por: Li, Zixuan, et al.
Publicado: (2025)
Inference-Time Decontamination: Reusing Leaked Benchmarks for Large Language Model Evaluation
por: Zhu, Qin, et al.
Publicado: (2024)
por: Zhu, Qin, et al.
Publicado: (2024)
Large Language Models Lack Understanding of Character Composition of Words
por: Shin, Andrew, et al.
Publicado: (2024)
por: Shin, Andrew, et al.
Publicado: (2024)
Towards Transparent AI: A Survey on Explainable Language Models
por: Palikhe, Avash, et al.
Publicado: (2025)
por: Palikhe, Avash, et al.
Publicado: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
por: Binkowski, Jakub, et al.
Publicado: (2026)
por: Binkowski, Jakub, et al.
Publicado: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
por: Peng, Runyu, et al.
Publicado: (2026)
por: Peng, Runyu, et al.
Publicado: (2026)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
por: Qiu, Zihan, et al.
Publicado: (2025)
por: Qiu, Zihan, et al.
Publicado: (2025)
Attention Sinks in Diffusion Language Models
por: Rulli, Maximo Eduardo, et al.
Publicado: (2025)
por: Rulli, Maximo Eduardo, et al.
Publicado: (2025)
A Chinese Dataset for Evaluating the Safeguards in Large Language Models
por: Wang, Yuxia, et al.
Publicado: (2024)
por: Wang, Yuxia, et al.
Publicado: (2024)
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
por: Dasgupta, Sayantan, et al.
Publicado: (2026)
por: Dasgupta, Sayantan, et al.
Publicado: (2026)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
por: Shiotani, Taihei, et al.
Publicado: (2026)
por: Shiotani, Taihei, et al.
Publicado: (2026)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
por: Kaneko, Masahiro, et al.
Publicado: (2023)
por: Kaneko, Masahiro, et al.
Publicado: (2023)
SinkLoRA: Enhanced Efficiency and Chat Capabilities for Long-Context Large Language Models
por: Zhang, Hengyu
Publicado: (2024)
por: Zhang, Hengyu
Publicado: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
por: Ohi, Masanari, et al.
Publicado: (2024)
por: Ohi, Masanari, et al.
Publicado: (2024)
Ejemplares similares
-
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
por: Kaneko, Masahiro, et al.
Publicado: (2025) -
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
por: Kaneko, Masahiro, et al.
Publicado: (2025) -
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
por: Kaneko, Masahiro, et al.
Publicado: (2024) -
Eagle: Ethical Dataset Given from Real Interactions
por: Kaneko, Masahiro, et al.
Publicado: (2024) -
Beyond the Resumé: A Rubric-Aware Automatic Interview System for Information Elicitation
por: Stuart, Harry, et al.
Publicado: (2026)