GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Ghasemi, Narges, Ziashahabi, Amir, Avestimehr, Salman, Shahabi, Cyrus |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap Data
by: Ziashahabi, Amir, et al.
Published: (2025)
by: Ziashahabi, Amir, et al.
Published: (2025)
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024)
by: Chen, Dengsheng, et al.
Published: (2024)
Robust Multimodal Learning via Cross-Modal Proxy Tokens
by: Reza, Md Kaykobad, et al.
Published: (2025)
by: Reza, Md Kaykobad, et al.
Published: (2025)
Geo2Vec: Shape- and Distance-Aware Neural Representation of Geospatial Entities
by: Chu, Chen, et al.
Published: (2025)
by: Chu, Chen, et al.
Published: (2025)
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)
by: Zha, Kaiwen, et al.
Published: (2024)
Diffusion Autoencoders are Scalable Image Tokenizers
by: Chen, Yinbo, et al.
Published: (2025)
by: Chen, Yinbo, et al.
Published: (2025)
Communication-Inspired Tokenization for Structured Image Representations
by: Davtyan, Aram, et al.
Published: (2026)
by: Davtyan, Aram, et al.
Published: (2026)
Adaptive Length Image Tokenization via Recurrent Allocation
by: Duggal, Shivam, et al.
Published: (2024)
by: Duggal, Shivam, et al.
Published: (2024)
ATHENA: Adaptive Test-Time Steering for Improving Count Fidelity in Diffusion Models
by: Sepehri, Mohammad Shahab, et al.
Published: (2026)
by: Sepehri, Mohammad Shahab, et al.
Published: (2026)
H$_{2}$OT: Hierarchical Hourglass Tokenizer for Efficient Video Pose Transformers
by: Li, Wenhao, et al.
Published: (2025)
by: Li, Wenhao, et al.
Published: (2025)
A More Word-like Image Tokenization for MLLMs
by: Lee, Hyun, et al.
Published: (2026)
by: Lee, Hyun, et al.
Published: (2026)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
ULTra: Unveiling Latent Token Interpretability in Transformer-Based Understanding and Segmentation
by: Hosseini, Hesam, et al.
Published: (2024)
by: Hosseini, Hesam, et al.
Published: (2024)
Multi-Token Prediction Needs Registers
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
by: Gerontopoulos, Anastasios, et al.
Published: (2025)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
by: Wu, Xiaoping, et al.
Published: (2024)
by: Wu, Xiaoping, et al.
Published: (2024)
Single-pass Adaptive Image Tokenization for Minimum Program Search
by: Duggal, Shivam, et al.
Published: (2025)
by: Duggal, Shivam, et al.
Published: (2025)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
by: Zheng, Chenhao, et al.
Published: (2025)
by: Zheng, Chenhao, et al.
Published: (2025)
ToDo: Token Downsampling for Efficient Generation of High-Resolution Images
by: Smith, Ethan, et al.
Published: (2024)
by: Smith, Ethan, et al.
Published: (2024)
(1D) Ordered Tokens Enable Efficient Test-Time Search
by: Gao, Zhitong, et al.
Published: (2026)
by: Gao, Zhitong, et al.
Published: (2026)
GeoRC: A Benchmark for Geolocation Reasoning Chains
by: Talreja, Mohit, et al.
Published: (2026)
by: Talreja, Mohit, et al.
Published: (2026)
LLaVA-Zip: Adaptive Visual Token Compression with Intrinsic Image Information
by: Wang, Ke, et al.
Published: (2024)
by: Wang, Ke, et al.
Published: (2024)
Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning
by: Jeong, Wooseong, et al.
Published: (2025)
by: Jeong, Wooseong, et al.
Published: (2025)
Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
Identifiable Token Correspondence for World Models
by: Kim, Youngin, et al.
Published: (2026)
by: Kim, Youngin, et al.
Published: (2026)
Uncertainty-DTW for Sequences and Visual Tokens
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Taming Outlier Tokens in Diffusion Transformers
by: Wu, Xiaoyu, et al.
Published: (2026)
by: Wu, Xiaoyu, et al.
Published: (2026)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
by: Chen, Liang, et al.
Published: (2024)
by: Chen, Liang, et al.
Published: (2024)
Negative Token Merging: Image-based Adversarial Feature Guidance
by: Singh, Jaskirat, et al.
Published: (2024)
by: Singh, Jaskirat, et al.
Published: (2024)
Enhancing Worldwide Image Geolocation by Ensembling Satellite-Based Ground-Level Attribute Predictors
by: Bianco, Michael J., et al.
Published: (2024)
by: Bianco, Michael J., et al.
Published: (2024)
Efficient World Models with Context-Aware Tokenization
by: Micheli, Vincent, et al.
Published: (2024)
by: Micheli, Vincent, et al.
Published: (2024)
Robustness Tokens: Towards Adversarial Robustness of Transformers
by: Pulfer, Brian, et al.
Published: (2025)
by: Pulfer, Brian, et al.
Published: (2025)
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization
by: Nguyen, Son, et al.
Published: (2025)
by: Nguyen, Son, et al.
Published: (2025)
DisCoRD: Discrete Tokens to Continuous Motion via Rectified Flow Decoding
by: Cho, Jungbin, et al.
Published: (2024)
by: Cho, Jungbin, et al.
Published: (2024)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
SkipSR: Faster Super Resolution with Token Skipping
by: Choudhury, Rohan, et al.
Published: (2025)
by: Choudhury, Rohan, et al.
Published: (2025)
Accelerating Diffusion Transformers with Token-wise Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
Token Activation Map to Visually Explain Multimodal LLMs
by: Li, Yi, et al.
Published: (2025)
by: Li, Yi, et al.
Published: (2025)
Random Token Fusion for Multi-View Medical Diagnosis
by: Guo, Jingyu, et al.
Published: (2024)
by: Guo, Jingyu, et al.
Published: (2024)
Exploring Token Pruning in Vision State Space Models
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
Similar Items
-
OSMGen: Highly Controllable Satellite Image Synthesis using OpenStreetMap Data
by: Ziashahabi, Amir, et al.
Published: (2025) -
High-Resolution Image Synthesis via Next-Token Prediction
by: Chen, Dengsheng, et al.
Published: (2024) -
Robust Multimodal Learning via Cross-Modal Proxy Tokens
by: Reza, Md Kaykobad, et al.
Published: (2025) -
Geo2Vec: Shape- and Distance-Aware Neural Representation of Geospatial Entities
by: Chu, Chen, et al.
Published: (2025) -
Language-Guided Image Tokenization for Generation
by: Zha, Kaiwen, et al.
Published: (2024)