Adaptive Length Image Tokenization via Recurrent Allocation
Fuente:
arXiv
Saved in:
| Main Authors: | Duggal, Shivam, Isola, Phillip, Torralba, Antonio, Freeman, William T. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026)
by: Duggal, Shivam, et al.
Published: (2026)
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
by: Shen, William, et al.
Published: (2023)
by: Shen, William, et al.
Published: (2023)
A Vision Check-up for Language Models
by: Sharma, Pratyusha, et al.
Published: (2024)
by: Sharma, Pratyusha, et al.
Published: (2024)
SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
by: Brown, Alexandre, et al.
Published: (2025)
by: Brown, Alexandre, et al.
Published: (2025)
TTF-VLA: Temporal Token Fusion via Pixel-Attention Integration for Vision-Language-Action Models
by: Liu, Chenghao, et al.
Published: (2025)
by: Liu, Chenghao, et al.
Published: (2025)
Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
by: Bendikas, Rokas, et al.
Published: (2025)
by: Bendikas, Rokas, et al.
Published: (2025)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
by: Zheng, Shuhong, et al.
Published: (2026)
by: Zheng, Shuhong, et al.
Published: (2026)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
Enhance Vision-based Tactile Sensors via Dynamic Illumination and Image Fusion
by: Redkin, Artemii, et al.
Published: (2025)
by: Redkin, Artemii, et al.
Published: (2025)
Training Neural Networks from Scratch with Parallel Low-Rank Adapters
by: Huh, Minyoung, et al.
Published: (2024)
by: Huh, Minyoung, et al.
Published: (2024)
ARTA: Adaptive Mixed-Resolution Token Allocation for Efficient Dense Feature Extraction
by: Hagerman, David, et al.
Published: (2026)
by: Hagerman, David, et al.
Published: (2026)
SilvaScenes: Tree Segmentation and Species Classification from Under-Canopy Images in Natural Forests
by: Duclos, David-Alexandre, et al.
Published: (2025)
by: Duclos, David-Alexandre, et al.
Published: (2025)
Leveraging Image Generators to Address Training Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
by: Jeanson, Gabriel, et al.
Published: (2026)
by: Jeanson, Gabriel, et al.
Published: (2026)
Adaptive Prediction Ensemble: Improving Out-of-Distribution Generalization of Motion Forecasting
by: Li, Jinning, et al.
Published: (2024)
by: Li, Jinning, et al.
Published: (2024)
Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Addressing the Waypoint-Action Gap in End-to-End Autonomous Driving via Vehicle Motion Models
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
Adapt3R: Adaptive 3D Scene Representation for Domain Transfer in Imitation Learning
by: Wilcox, Albert, et al.
Published: (2025)
by: Wilcox, Albert, et al.
Published: (2025)
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
by: Lai, Zhenglin, et al.
Published: (2026)
by: Lai, Zhenglin, et al.
Published: (2026)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
GHIL-Glue: Hierarchical Control with Filtered Subgoal Images
by: Hatch, Kyle B., et al.
Published: (2024)
by: Hatch, Kyle B., et al.
Published: (2024)
Aligning Text-to-Image Diffusion Models with Reward Backpropagation
by: Prabhudesai, Mihir, et al.
Published: (2023)
by: Prabhudesai, Mihir, et al.
Published: (2023)
3D Hand Pose Estimation in Everyday Egocentric Images
by: Prakash, Aditya, et al.
Published: (2023)
by: Prakash, Aditya, et al.
Published: (2023)
You Only Crash Once v2: Perceptually Consistent Strong Features for One-Stage Domain Adaptive Detection of Space Terrain
by: Chase Jr, Timothy, et al.
Published: (2025)
by: Chase Jr, Timothy, et al.
Published: (2025)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
by: Goyal, Divyanshu, et al.
Published: (2026)
by: Goyal, Divyanshu, et al.
Published: (2026)
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
by: Cazenavette, George, et al.
Published: (2025)
by: Cazenavette, George, et al.
Published: (2025)
Render and Diffuse: Aligning Image and Action Spaces for Diffusion-based Behaviour Cloning
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
by: Pinkovich, Barak, et al.
Published: (2025)
by: Pinkovich, Barak, et al.
Published: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radar
by: Ding, Fangqiang, et al.
Published: (2024)
by: Ding, Fangqiang, et al.
Published: (2024)
TempBEV: Improving Learned BEV Encoders with Combined Image and BEV Space Temporal Aggregation
by: Monninger, Thomas, et al.
Published: (2024)
by: Monninger, Thomas, et al.
Published: (2024)
Cross-Modal Instructions for Robot Motion Generation
by: Barron, William, et al.
Published: (2025)
by: Barron, William, et al.
Published: (2025)
The Platonic Representation Hypothesis
by: Huh, Minyoung, et al.
Published: (2024)
by: Huh, Minyoung, et al.
Published: (2024)
Video Diffusion Alignment via Reward Gradients
by: Prabhudesai, Mihir, et al.
Published: (2024)
by: Prabhudesai, Mihir, et al.
Published: (2024)
Unlocking Generalization for Robotics via Modularity and Scale
by: Dalal, Murtaza
Published: (2025)
by: Dalal, Murtaza
Published: (2025)
Robot-Enabled Machine Learning-Based Diagnosis of Gastric Cancer Polyps Using Partial Surface Tactile Imaging
by: Kapuria, Siddhartha, et al.
Published: (2024)
by: Kapuria, Siddhartha, et al.
Published: (2024)
Understanding Video Transformers via Universal Concept Discovery
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
Separating Knowledge and Perception with Procedural Data
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2025)
by: Rodríguez-Muñoz, Adrián, et al.
Published: (2025)
Instant Policy: In-Context Imitation Learning via Graph Diffusion
by: Vosylius, Vitalis, et al.
Published: (2024)
by: Vosylius, Vitalis, et al.
Published: (2024)
What's the Move? Hybrid Imitation Learning via Salient Points
by: Sundaresan, Priya, et al.
Published: (2024)
by: Sundaresan, Priya, et al.
Published: (2024)
Similar Items
-
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025) -
End-to-End Training for Unified Tokenization and Latent Denoising
by: Duggal, Shivam, et al.
Published: (2026) -
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
by: Shen, William, et al.
Published: (2023) -
A Vision Check-up for Language Models
by: Sharma, Pratyusha, et al.
Published: (2024) -
SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokens
by: Brown, Alexandre, et al.
Published: (2025)