Core Tokensets for Data-efficient Sequential Training of Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Paul, Subarnaduti, Brack, Manuel, Schramowski, Patrick, Kersting, Kristian, Mundt, Martin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024)
by: Helff, Lukas, et al.
Published: (2024)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
by: Hegde, Niharika, et al.
Published: (2025)
by: Hegde, Niharika, et al.
Published: (2025)
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022)
by: Struppek, Lukas, et al.
Published: (2022)
DeiSAM: Segment Anything with Deictic Prompting
by: Shindo, Hikaru, et al.
Published: (2024)
by: Shindo, Hikaru, et al.
Published: (2024)
LEDITS++: Limitless Image Editing using Text-to-Image Models
by: Brack, Manuel, et al.
Published: (2023)
by: Brack, Manuel, et al.
Published: (2023)
ART: Adaptive Relation Tuning for Generalized Relation Prediction
by: Sudhakaran, Gopika, et al.
Published: (2025)
by: Sudhakaran, Gopika, et al.
Published: (2025)
Does CLIP Know My Face?
by: Hintersdorf, Dominik, et al.
Published: (2022)
by: Hintersdorf, Dominik, et al.
Published: (2022)
No Safe Dose: How Training Data Drives Unsafe Image Generation
by: Friedrich, Felix, et al.
Published: (2026)
by: Friedrich, Felix, et al.
Published: (2026)
The Cake that is Intelligence and Who Gets to Bake it: An AI Analogy and its Implications for Participation
by: Mundt, Martin, et al.
Published: (2025)
by: Mundt, Martin, et al.
Published: (2025)
T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
by: Deiseroth, Björn, et al.
Published: (2024)
by: Deiseroth, Björn, et al.
Published: (2024)
Learning Differentiable Logic Programs for Abstract Visual Reasoning
by: Shindo, Hikaru, et al.
Published: (2023)
by: Shindo, Hikaru, et al.
Published: (2023)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
OCAtari: Object-Centric Atari 2600 Reinforcement Learning Environments
by: Delfosse, Quentin, et al.
Published: (2023)
by: Delfosse, Quentin, et al.
Published: (2023)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
V-LoL: A Diagnostic Dataset for Visual Logical Learning
by: Helff, Lukas, et al.
Published: (2023)
by: Helff, Lukas, et al.
Published: (2023)
Finding DoRI: Discovery of Retained Images in Diffusion Models
by: Kowalczuk, Antoni, et al.
Published: (2025)
by: Kowalczuk, Antoni, et al.
Published: (2025)
Training Transitive and Commutative Multimodal Transformers with LoReTTa
by: Tran, Manuel, et al.
Published: (2023)
by: Tran, Manuel, et al.
Published: (2023)
UniFusion: Vision-Language Model as Unified Encoder in Image Generation
by: Li, Kevin, et al.
Published: (2025)
by: Li, Kevin, et al.
Published: (2025)
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
by: Schuhmann, Christoph, et al.
Published: (2025)
by: Schuhmann, Christoph, et al.
Published: (2025)
Fake & Square: Training Self-Supervised Vision Transformers with Synthetic Data and Synthetic Hard Negatives
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
by: Kohaut, Simon, et al.
Published: (2025)
by: Kohaut, Simon, et al.
Published: (2025)
Vision Transformers Don't Need Trained Registers
by: Jiang, Nick, et al.
Published: (2025)
by: Jiang, Nick, et al.
Published: (2025)
Unsupervised Training of Vision Transformers with Synthetic Negatives
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
by: Giakoumoglou, Nikolaos, et al.
Published: (2025)
RCTrans: Radar-Camera Transformer via Radar Densifier and Sequential Decoder for 3D Object Detection
by: Li, Yiheng, et al.
Published: (2024)
by: Li, Yiheng, et al.
Published: (2024)
TLCM: Training-efficient Latent Consistency Model for Image Generation with 2-8 Steps
by: Xie, Qingsong, et al.
Published: (2024)
by: Xie, Qingsong, et al.
Published: (2024)
MUG-V 10B: High-efficiency Training Pipeline for Large Video Generation Models
by: Zhang, Yongshun, et al.
Published: (2025)
by: Zhang, Yongshun, et al.
Published: (2025)
Q-DiT: Accurate Post-Training Quantization for Diffusion Transformers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Deep Classifier Mimicry without Data Access
by: Braun, Steven, et al.
Published: (2023)
by: Braun, Steven, et al.
Published: (2023)
LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
by: Sztwiertnia, Sebastian, et al.
Published: (2025)
Improved Ear Verification with Vision Transformers and Overlapping Patches
by: Arun, Deeksha, et al.
Published: (2025)
by: Arun, Deeksha, et al.
Published: (2025)
Pre-Trained CNN Architecture for Transformer-Based Image Caption Generation Model
by: Dufera, Amanuel Tafese
Published: (2025)
by: Dufera, Amanuel Tafese
Published: (2025)
Chipmunk: Training-Free Acceleration of Diffusion Transformers with Dynamic Column-Sparse Deltas
by: Silveria, Austin, et al.
Published: (2025)
by: Silveria, Austin, et al.
Published: (2025)
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
by: Federici, Marco, et al.
Published: (2025)
by: Federici, Marco, et al.
Published: (2025)
EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients
by: Wu, Meihan, et al.
Published: (2024)
by: Wu, Meihan, et al.
Published: (2024)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
by: Shi, Huihong, et al.
Published: (2024)
by: Shi, Huihong, et al.
Published: (2024)
Fast-iTPN: Integrally Pre-Trained Transformer Pyramid Network with Token Migration
by: Tian, Yunjie, et al.
Published: (2022)
by: Tian, Yunjie, et al.
Published: (2022)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
by: Zou, Dongyun, et al.
Published: (2026)
by: Zou, Dongyun, et al.
Published: (2026)
Dual-Channel Attention Guidance for Training-Free Image Editing Control in Diffusion Transformers
by: Li, Guandong
Published: (2026)
by: Li, Guandong
Published: (2026)
SeqPE: Transformer with Sequential Position Encoding
by: Li, Huayang, et al.
Published: (2025)
by: Li, Huayang, et al.
Published: (2025)
Similar Items
-
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models
by: Helff, Lukas, et al.
Published: (2024) -
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025) -
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
by: Hegde, Niharika, et al.
Published: (2025) -
Exploiting Cultural Biases via Homoglyphs in Text-to-Image Synthesis
by: Struppek, Lukas, et al.
Published: (2022) -
DeiSAM: Segment Anything with Deictic Prompting
by: Shindo, Hikaru, et al.
Published: (2024)