LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Sztwiertnia, Sebastian, Friedrich, Felix, Kersting, Kristian, Schramowski, Patrick, Deiseroth, Björn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
by: Härle, Ruben, et al.
Published: (2024)
by: Härle, Ruben, et al.
Published: (2024)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
Measuring and Guiding Monosemanticity
by: Härle, Ruben, et al.
Published: (2025)
by: Härle, Ruben, et al.
Published: (2025)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)
by: Höth, Max Henning, et al.
Published: (2026)
Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur Protocols
by: Deiseroth, Björn, et al.
Published: (2025)
by: Deiseroth, Björn, et al.
Published: (2025)
A Typology for Exploring the Mitigation of Shortcut Behavior
by: Friedrich, Felix, et al.
Published: (2022)
by: Friedrich, Felix, et al.
Published: (2022)
Divergent Token Metrics: Measuring degradation to prune away LLM components -- and optimize quantization
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
by: Hegde, Niharika, et al.
Published: (2025)
by: Hegde, Niharika, et al.
Published: (2025)
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024)
by: Helff, Lukas, et al.
Published: (2024)
SLR: Automated Synthesis for Scalable Logical Reasoning
by: Helff, Lukas, et al.
Published: (2025)
by: Helff, Lukas, et al.
Published: (2025)
Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information
by: Struppek, Lukas, et al.
Published: (2025)
by: Struppek, Lukas, et al.
Published: (2025)
Making Metadata More FAIR Using Large Language Models
by: Sundaram, Sowmya S., et al.
Published: (2023)
by: Sundaram, Sowmya S., et al.
Published: (2023)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
by: Delfosse, Quentin, et al.
Published: (2023)
by: Delfosse, Quentin, et al.
Published: (2023)
Core Tokensets for Data-efficient Sequential Training of Transformers
by: Paul, Subarnaduti, et al.
Published: (2024)
by: Paul, Subarnaduti, et al.
Published: (2024)
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022)
by: Struppek, Lukas, et al.
Published: (2022)
LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
by: Neitemeier, Pit, et al.
Published: (2025)
by: Neitemeier, Pit, et al.
Published: (2025)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
by: Friedrich, Felix, et al.
Published: (2025)
by: Friedrich, Felix, et al.
Published: (2025)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
LEDITS++: Limitless Image Editing using Text-to-Image Models
by: Brack, Manuel, et al.
Published: (2023)
by: Brack, Manuel, et al.
Published: (2023)
LIME-LLM: Probing Models with Fluent Counterfactuals, Not Broken Text
by: Mihaila, George, et al.
Published: (2026)
by: Mihaila, George, et al.
Published: (2026)
Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints
by: Jehu-Appiah, Rodney
Published: (2026)
by: Jehu-Appiah, Rodney
Published: (2026)
Soft Begging: Modular and Efficient Shielding of LLMs against Prompt Injection and Jailbreaking based on Prompt Tuning
by: Ostermann, Simon, et al.
Published: (2024)
by: Ostermann, Simon, et al.
Published: (2024)
EmoNet-Voice: A Fine-Grained, Expert-Verified Benchmark for Speech Emotion Detection
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
by: Paul, Indraneil, et al.
Published: (2025)
by: Paul, Indraneil, et al.
Published: (2025)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
by: Heyen, Henning, et al.
Published: (2024)
by: Heyen, Henning, et al.
Published: (2024)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025)
by: Burns, Thomas F, et al.
Published: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
ART: Adaptive Relation Tuning for Generalized Relation Prediction
by: Sudhakaran, Gopika, et al.
Published: (2025)
by: Sudhakaran, Gopika, et al.
Published: (2025)
Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You
by: Friedrich, Felix, et al.
Published: (2024)
by: Friedrich, Felix, et al.
Published: (2024)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
by: Watson, William, et al.
Published: (2026)
by: Watson, William, et al.
Published: (2026)
ALERT: A Comprehensive Benchmark for Assessing Large Language Models' Safety through Red Teaming
by: Tedeschi, Simone, et al.
Published: (2024)
by: Tedeschi, Simone, et al.
Published: (2024)
DeiSAM: Segment Anything with Deictic Prompting
by: Shindo, Hikaru, et al.
Published: (2024)
by: Shindo, Hikaru, et al.
Published: (2024)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
by: Ali, Mehdi, et al.
Published: (2025)
by: Ali, Mehdi, et al.
Published: (2025)
Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting
by: Shrestha, Gyanendra, et al.
Published: (2025)
by: Shrestha, Gyanendra, et al.
Published: (2025)
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems
by: Shindo, Hikaru, et al.
Published: (2026)
by: Shindo, Hikaru, et al.
Published: (2026)
An Investigation of Linguistic Biases in LLM-Based Recommendations
by: Venkateswaran, Nitin, et al.
Published: (2026)
by: Venkateswaran, Nitin, et al.
Published: (2026)
Linguistics-Aware Non-Distortionary LLM Watermarking
by: Park, Shinwoo, et al.
Published: (2026)
by: Park, Shinwoo, et al.
Published: (2026)
Output Embedding Centering for Stable LLM Pretraining
by: Stollenwerk, Felix, et al.
Published: (2026)
by: Stollenwerk, Felix, et al.
Published: (2026)
Similar Items
-
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024) -
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
by: Härle, Ruben, et al.
Published: (2024) -
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023) -
Measuring and Guiding Monosemanticity
by: Härle, Ruben, et al.
Published: (2025) -
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)