Single-pass Adaptive Image Tokenization for Minimum Program Search
Fuente:
arXiv
Salvato in:
| Autori principali: | Duggal, Shivam, Byun, Sanghyun, Freeman, William T., Torralba, Antonio, Isola, Phillip |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Adaptive Length Image Tokenization via Recurrent Allocation
di: Duggal, Shivam, et al.
Pubblicazione: (2024)
di: Duggal, Shivam, et al.
Pubblicazione: (2024)
End-to-End Training for Unified Tokenization and Latent Denoising
di: Duggal, Shivam, et al.
Pubblicazione: (2026)
di: Duggal, Shivam, et al.
Pubblicazione: (2026)
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
A Vision Check-up for Language Models
di: Sharma, Pratyusha, et al.
Pubblicazione: (2024)
di: Sharma, Pratyusha, et al.
Pubblicazione: (2024)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
di: Cazenavette, George, et al.
Pubblicazione: (2025)
di: Cazenavette, George, et al.
Pubblicazione: (2025)
Separating Knowledge and Perception with Procedural Data
di: Rodríguez-Muñoz, Adrián, et al.
Pubblicazione: (2025)
di: Rodríguez-Muñoz, Adrián, et al.
Pubblicazione: (2025)
The Platonic Representation Hypothesis
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
di: Huh, Minyoung, et al.
Pubblicazione: (2024)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
di: Wang, Ke, et al.
Pubblicazione: (2024)
di: Wang, Ke, et al.
Pubblicazione: (2024)
Life, Machine Learning, and the Search for Habitability: Predicting Biosignature Fluxes for the Habitable Worlds Observatory
di: Moussa, Mark, et al.
Pubblicazione: (2026)
di: Moussa, Mark, et al.
Pubblicazione: (2026)
DiffEM: Learning from Corrupted Data with Diffusion Models via Expectation Maximization
di: Hosseintabar, Danial, et al.
Pubblicazione: (2025)
di: Hosseintabar, Danial, et al.
Pubblicazione: (2025)
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
di: Shen, William, et al.
Pubblicazione: (2023)
di: Shen, William, et al.
Pubblicazione: (2023)
GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
di: Ghasemi, Narges, et al.
Pubblicazione: (2025)
Mixture of Nested Experts: Adaptive Processing of Visual Tokens
di: Jain, Gagan, et al.
Pubblicazione: (2024)
di: Jain, Gagan, et al.
Pubblicazione: (2024)
Language-Guided Image Tokenization for Generation
di: Zha, Kaiwen, et al.
Pubblicazione: (2024)
di: Zha, Kaiwen, et al.
Pubblicazione: (2024)
Diffusion Autoencoders are Scalable Image Tokenizers
di: Chen, Yinbo, et al.
Pubblicazione: (2025)
di: Chen, Yinbo, et al.
Pubblicazione: (2025)
(1D) Ordered Tokens Enable Efficient Test-Time Search
di: Gao, Zhitong, et al.
Pubblicazione: (2026)
di: Gao, Zhitong, et al.
Pubblicazione: (2026)
Communication-Inspired Tokenization for Structured Image Representations
di: Davtyan, Aram, et al.
Pubblicazione: (2026)
di: Davtyan, Aram, et al.
Pubblicazione: (2026)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
A More Word-like Image Tokenization for MLLMs
di: Lee, Hyun, et al.
Pubblicazione: (2026)
di: Lee, Hyun, et al.
Pubblicazione: (2026)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
di: Lin, Huawei, et al.
Pubblicazione: (2025)
di: Lin, Huawei, et al.
Pubblicazione: (2025)
Layer- and Timestep-Adaptive Differentiable Token Compression Ratios for Efficient Diffusion Transformers
di: You, Haoran, et al.
Pubblicazione: (2024)
di: You, Haoran, et al.
Pubblicazione: (2024)
ARTA: Adaptive Mixed-Resolution Token Allocation for Efficient Dense Feature Extraction
di: Hagerman, David, et al.
Pubblicazione: (2026)
di: Hagerman, David, et al.
Pubblicazione: (2026)
Beyond I-Con: Exploring New Dimension of Distance Measures in Representation Learning
di: Shone, Jasmine, et al.
Pubblicazione: (2025)
di: Shone, Jasmine, et al.
Pubblicazione: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
di: Chen, Dengsheng, et al.
Pubblicazione: (2024)
di: Chen, Dengsheng, et al.
Pubblicazione: (2024)
Not All Layers Are Created Equal: Adaptive LoRA Ranks for Personalized Image Generation
di: Shenaj, Donald, et al.
Pubblicazione: (2026)
di: Shenaj, Donald, et al.
Pubblicazione: (2026)
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
di: Smith, Ethan, et al.
Pubblicazione: (2024)
di: Smith, Ethan, et al.
Pubblicazione: (2024)
MultiFloodSynth: Multi-Annotated Flood Synthetic Dataset Generation
di: Kang, YoonJe, et al.
Pubblicazione: (2025)
di: Kang, YoonJe, et al.
Pubblicazione: (2025)
Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
di: Bahng, Hyojin, et al.
Pubblicazione: (2025)
di: Bahng, Hyojin, et al.
Pubblicazione: (2025)
MedNNS: Supernet-based Medical Task-Adaptive Neural Network Search
di: Mecharbat, Lotfi Abdelkrim, et al.
Pubblicazione: (2025)
di: Mecharbat, Lotfi Abdelkrim, et al.
Pubblicazione: (2025)
COMO: Closed-Loop Optical Molecule Recognition with Minimum Risk Training
di: Lyu, Zhuoqi, et al.
Pubblicazione: (2026)
di: Lyu, Zhuoqi, et al.
Pubblicazione: (2026)
Negative Token Merging: Image-based Adversarial Feature Guidance
di: Singh, Jaskirat, et al.
Pubblicazione: (2024)
di: Singh, Jaskirat, et al.
Pubblicazione: (2024)
Adaptive Patching for High-resolution Image Segmentation with Transformers
di: Zhang, Enzhi, et al.
Pubblicazione: (2024)
di: Zhang, Enzhi, et al.
Pubblicazione: (2024)
Scaling Image and Video Generation via Test-Time Evolutionary Search
di: He, Haoran, et al.
Pubblicazione: (2025)
di: He, Haoran, et al.
Pubblicazione: (2025)
A Lightweight Neural Architecture Search Model for Medical Image Classification
di: Xie, Lunchen, et al.
Pubblicazione: (2024)
di: Xie, Lunchen, et al.
Pubblicazione: (2024)
Peer-Ranked Precision: Creating a Foundational Dataset for Fine-Tuning Vision Models from DataSeeds' Annotated Imagery
di: Abdoli, Sajjad, et al.
Pubblicazione: (2025)
di: Abdoli, Sajjad, et al.
Pubblicazione: (2025)
Hybrid Deep Learning for Hyperspectral Single Image Super-Resolution
di: Muhammad, Usman, et al.
Pubblicazione: (2025)
di: Muhammad, Usman, et al.
Pubblicazione: (2025)
SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
di: Embacher, Felix, et al.
Pubblicazione: (2026)
di: Embacher, Felix, et al.
Pubblicazione: (2026)
Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction
di: Farahbakhsh, Mahdi, et al.
Pubblicazione: (2025)
di: Farahbakhsh, Mahdi, et al.
Pubblicazione: (2025)
CLASP: Adaptive Spectral Clustering for Unsupervised Per-Image Segmentation
di: Curie, Max, et al.
Pubblicazione: (2025)
di: Curie, Max, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Adaptive Length Image Tokenization via Recurrent Allocation
di: Duggal, Shivam, et al.
Pubblicazione: (2024) -
End-to-End Training for Unified Tokenization and Latent Denoising
di: Duggal, Shivam, et al.
Pubblicazione: (2026) -
Minecraft-ify: Minecraft Style Image Generation with Text-guided Image Editing for In-Game Application
di: Kim, Bumsoo, et al.
Pubblicazione: (2024) -
A Vision Check-up for Language Models
di: Sharma, Pratyusha, et al.
Pubblicazione: (2024) -
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
di: Huh, Minyoung, et al.
Pubblicazione: (2024)