Better Language Model Inversion by Compactly Representing Next-Token Distributions
Fuente:
arXiv
Guardado en:
| Autores principales: | Nazir, Murtaza, Finlayson, Matthew, Morris, John X., Ren, Xiang, Swayamdipta, Swabha |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Logits of API-Protected LLMs Leak Proprietary Information
por: Finlayson, Matthew, et al.
Publicado: (2024)
por: Finlayson, Matthew, et al.
Publicado: (2024)
Every Language Model Has a Forgery-Resistant Signature
por: Finlayson, Matthew, et al.
Publicado: (2025)
por: Finlayson, Matthew, et al.
Publicado: (2025)
Teaching Models to Understand (but not Generate) High-risk Data
por: Wang, Ryan, et al.
Publicado: (2025)
por: Wang, Ryan, et al.
Publicado: (2025)
Annotating FrameNet via Structure-Conditioned Language Generation
por: Cui, Xinyue, et al.
Publicado: (2024)
por: Cui, Xinyue, et al.
Publicado: (2024)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
por: Joshi, Brihi, et al.
Publicado: (2025)
por: Joshi, Brihi, et al.
Publicado: (2025)
How Reliable is Language Model Micro-Benchmarking?
por: Yauney, Gregory, et al.
Publicado: (2025)
por: Yauney, Gregory, et al.
Publicado: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
por: He, Keyu, et al.
Publicado: (2025)
por: He, Keyu, et al.
Publicado: (2025)
Side-by-side Comparison Amplifies Dialect Bias in Language Models
por: Kondapally, Kritee, et al.
Publicado: (2026)
por: Kondapally, Kritee, et al.
Publicado: (2026)
Disentangling Geometry, Performance, and Training in Language Models
por: Kulkarni, Atharva, et al.
Publicado: (2026)
por: Kulkarni, Atharva, et al.
Publicado: (2026)
Compare without Despair: Reliable Preference Evaluation with Generation Separability
por: Ghosh, Sayan, et al.
Publicado: (2024)
por: Ghosh, Sayan, et al.
Publicado: (2024)
Understanding Dataset Difficulty with $\mathcal{V}$-Usable Information
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
por: Ethayarajh, Kawin, et al.
Publicado: (2021)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
por: Joshi, Brihi, et al.
Publicado: (2025)
por: Joshi, Brihi, et al.
Publicado: (2025)
Evaluation Under Imperfect Benchmarks and Ratings: A Case Study in Text Simplification
por: Liu, Joseph, et al.
Publicado: (2025)
por: Liu, Joseph, et al.
Publicado: (2025)
Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?
por: Khurana, Urja, et al.
Publicado: (2024)
por: Khurana, Urja, et al.
Publicado: (2024)
Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge
por: Cui, Xinyue, et al.
Publicado: (2025)
por: Cui, Xinyue, et al.
Publicado: (2025)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
por: Sathe, Ashutosh, et al.
Publicado: (2024)
por: Sathe, Ashutosh, et al.
Publicado: (2024)
BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity
por: Diddee, Harshita, et al.
Publicado: (2026)
por: Diddee, Harshita, et al.
Publicado: (2026)
Sample, Align, Synthesize: Graph-Based Response Synthesis with ConGrs
por: Ghosh, Sayan, et al.
Publicado: (2025)
por: Ghosh, Sayan, et al.
Publicado: (2025)
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge
por: Howard, Phillip, et al.
Publicado: (2023)
por: Howard, Phillip, et al.
Publicado: (2023)
Generative Explanations for Program Synthesizers
por: Nazari, Amirmohammad, et al.
Publicado: (2024)
por: Nazari, Amirmohammad, et al.
Publicado: (2024)
Universal Zero-shot Embedding Inversion
por: Zhang, Collin, et al.
Publicado: (2025)
por: Zhang, Collin, et al.
Publicado: (2025)
Rethinking Tokenization: Crafting Better Tokenizers for Large Language Models
por: Yang, Jinbiao
Publicado: (2024)
por: Yang, Jinbiao
Publicado: (2024)
Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can Produce
por: Wang, Haojin, et al.
Publicado: (2025)
por: Wang, Haojin, et al.
Publicado: (2025)
Splintering Nonconcatenative Languages for Better Tokenization
por: Gazit, Bar, et al.
Publicado: (2025)
por: Gazit, Bar, et al.
Publicado: (2025)
Pre-Trained Language Models Represent Some Geographic Populations Better Than Others
por: Dunn, Jonathan, et al.
Publicado: (2024)
por: Dunn, Jonathan, et al.
Publicado: (2024)
Mitigating Gradient Inversion Risks in Language Models via Token Obfuscation
por: Feng, Xinguo, et al.
Publicado: (2026)
por: Feng, Xinguo, et al.
Publicado: (2026)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
por: Ranjit, Jaspreet, et al.
Publicado: (2024)
por: Ranjit, Jaspreet, et al.
Publicado: (2024)
Bigram Subnetworks: Mapping to Next Tokens in Transformer Language Models
por: Chang, Tyler A., et al.
Publicado: (2025)
por: Chang, Tyler A., et al.
Publicado: (2025)
ChEmREF: Evaluating Language Model Readiness for Chemical Emergency Response
por: Surana, Risha, et al.
Publicado: (2025)
por: Surana, Risha, et al.
Publicado: (2025)
Can Language Models Represent the Past without Anachronism?
por: Underwood, Ted, et al.
Publicado: (2025)
por: Underwood, Ted, et al.
Publicado: (2025)
Why Fine-Tuning Encourages Hallucinations and How to Fix It
por: Kaplan, Guy, et al.
Publicado: (2026)
por: Kaplan, Guy, et al.
Publicado: (2026)
Multimodal Latent Language Modeling with Next-Token Diffusion
por: Sun, Yutao, et al.
Publicado: (2024)
por: Sun, Yutao, et al.
Publicado: (2024)
NeoBERT: A Next-Generation BERT
por: Breton, Lola Le, et al.
Publicado: (2025)
por: Breton, Lola Le, et al.
Publicado: (2025)
Are We Automating the Joy Out of Work? Designing AI to Augment Work, Not Meaning
por: Ranjit, Jaspreet, et al.
Publicado: (2026)
por: Ranjit, Jaspreet, et al.
Publicado: (2026)
A Law of Next-Token Prediction in Large Language Models
por: He, Hangfeng, et al.
Publicado: (2024)
por: He, Hangfeng, et al.
Publicado: (2024)
Differentially Private Next-Token Prediction of Large Language Models
por: Flemings, James, et al.
Publicado: (2024)
por: Flemings, James, et al.
Publicado: (2024)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
por: Wen, Yuxin, et al.
Publicado: (2024)
por: Wen, Yuxin, et al.
Publicado: (2024)
Reasoning Bias of Next Token Prediction Training
por: Lin, Pengxiao, et al.
Publicado: (2025)
por: Lin, Pengxiao, et al.
Publicado: (2025)
From Next-Token to Mathematics: The Learning Dynamics of Mathematical Reasoning in Language Models
por: Mishra, Shubhra, et al.
Publicado: (2024)
por: Mishra, Shubhra, et al.
Publicado: (2024)
Checklists Are Better Than Reward Models For Aligning Language Models
por: Viswanathan, Vijay, et al.
Publicado: (2025)
por: Viswanathan, Vijay, et al.
Publicado: (2025)
Ejemplares similares
-
Logits of API-Protected LLMs Leak Proprietary Information
por: Finlayson, Matthew, et al.
Publicado: (2024) -
Every Language Model Has a Forgery-Resistant Signature
por: Finlayson, Matthew, et al.
Publicado: (2025) -
Teaching Models to Understand (but not Generate) High-risk Data
por: Wang, Ryan, et al.
Publicado: (2025) -
Annotating FrameNet via Structure-Conditioned Language Generation
por: Cui, Xinyue, et al.
Publicado: (2024) -
Improving Language Model Personas via Rationalization with Psychological Scaffolds
por: Joshi, Brihi, et al.
Publicado: (2025)