Fractal Autoregressive Depth Estimation with Continuous Token Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Jinchang, Kang, Xinrou, Lu, Guoyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vision-Language Embodiment for Monocular Depth Estimation
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
Depth Estimation Based on 3D Gaussian Splatting Siamese Defocus
by: Zhang, Jinchang, et al.
Published: (2024)
by: Zhang, Jinchang, et al.
Published: (2024)
Language-Depth Navigated Thermal and Visible Image Fusion
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
Underground Mapping and Localization Based on Ground-Penetrating Radar
by: Zhang, Jinchang, et al.
Published: (2024)
by: Zhang, Jinchang, et al.
Published: (2024)
Embodiment: Self-Supervised Depth Estimation Based on Camera Models
by: Zhang, Jinchang, et al.
Published: (2024)
by: Zhang, Jinchang, et al.
Published: (2024)
Keypoint Detection and Description for Raw Bayer Images
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Graph Integrated Multimodal Concept Bottleneck Model
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Adaptive Event Stream Slicing for Open-Vocabulary Event-Based Object Detection via Vision-Language Knowledge Distillation
by: Zhang, Jinchang, et al.
Published: (2025)
by: Zhang, Jinchang, et al.
Published: (2025)
Automated Genomic Interpretation via Concept Bottleneck Models for Medical Robotics
by: Li, Zijun, et al.
Published: (2025)
by: Li, Zijun, et al.
Published: (2025)
3D Plant Root Skeleton Detection and Extraction
by: Lin, Jiakai, et al.
Published: (2025)
by: Lin, Jiakai, et al.
Published: (2025)
Scalable Autoregressive Monocular Depth Estimation
by: Wang, Jinhong, et al.
Published: (2024)
by: Wang, Jinhong, et al.
Published: (2024)
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
by: Wang, Bohan, et al.
Published: (2025)
by: Wang, Bohan, et al.
Published: (2025)
DepthART: Monocular Depth Estimation as Autoregressive Refinement Task
by: Gabdullin, Bulat, et al.
Published: (2024)
by: Gabdullin, Bulat, et al.
Published: (2024)
Visual Autoregressive Modelling for Monocular Depth Estimation
by: El-Ghoussani, Amir, et al.
Published: (2025)
by: El-Ghoussani, Amir, et al.
Published: (2025)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
by: Cheng, Tianle, et al.
Published: (2025)
by: Cheng, Tianle, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
LINA: Linear Autoregressive Image Generative Models with Continuous Tokens
by: Wang, Jiahao, et al.
Published: (2026)
by: Wang, Jiahao, et al.
Published: (2026)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
by: NextStep Team, et al.
Published: (2025)
by: NextStep Team, et al.
Published: (2025)
DepthMaster: Taming Diffusion Models for Monocular Depth Estimation
by: Song, Ziyang, et al.
Published: (2025)
by: Song, Ziyang, et al.
Published: (2025)
Frequency Autoregressive Image Generation with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Rethinking Discrete Tokens: Treating Them as Conditions for Continuous Autoregressive Image Synthesis
by: Zheng, Peng, et al.
Published: (2025)
by: Zheng, Peng, et al.
Published: (2025)
VideoMAR: Autoregressive Video Generatio with Continuous Tokens
by: Yu, Hu, et al.
Published: (2025)
by: Yu, Hu, et al.
Published: (2025)
Hyperspherical Latents Improve Continuous-Token Autoregressive Generation
by: Ke, Guolin, et al.
Published: (2025)
by: Ke, Guolin, et al.
Published: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Interpretable Traffic Responsibility from Dashcam Video via Legal Multi Agent Reasoning
by: Yang, Jingchun, et al.
Published: (2026)
by: Yang, Jingchun, et al.
Published: (2026)
D2C: Unlocking the Potential of Continuous Autoregressive Image Generation with Discrete Tokens
by: Wang, Panpan, et al.
Published: (2025)
by: Wang, Panpan, et al.
Published: (2025)
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
by: Dong, Yue-Jiang, et al.
Published: (2025)
by: Dong, Yue-Jiang, et al.
Published: (2025)
Improving Depth Gradient Continuity in Transformers: A Comparative Study on Monocular Depth Estimation with CNN
by: Yao, Jiawei, et al.
Published: (2023)
by: Yao, Jiawei, et al.
Published: (2023)
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
by: Sun, Haiyue, et al.
Published: (2025)
by: Sun, Haiyue, et al.
Published: (2025)
BetterDepth: Plug-and-Play Diffusion Refiner for Zero-Shot Monocular Depth Estimation
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
by: Zhen, Dingcheng, et al.
Published: (2025)
by: Zhen, Dingcheng, et al.
Published: (2025)
Infinite Gaze Generation for Videos with Autoregressive Diffusion
by: Kang, Jenna, et al.
Published: (2026)
by: Kang, Jenna, et al.
Published: (2026)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
by: Li, Yizhuo, et al.
Published: (2024)
by: Li, Yizhuo, et al.
Published: (2024)
Depth Adaptive Efficient Visual Autoregressive Modeling
by: Li, Chunliang, et al.
Published: (2026)
by: Li, Chunliang, et al.
Published: (2026)
ResTok: Learning Hierarchical Residuals in 1D Visual Tokenizers for Autoregressive Image Generation
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
by: Lin, Xin, et al.
Published: (2025)
by: Lin, Xin, et al.
Published: (2025)
PrimeDepth: Efficient Monocular Depth Estimation with a Stable Diffusion Preimage
by: Zavadski, Denis, et al.
Published: (2024)
by: Zavadski, Denis, et al.
Published: (2024)
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024)
by: Fan, Lijie, et al.
Published: (2024)
ScaleDepth: Decomposing Metric Depth Estimation into Scale Prediction and Relative Depth Estimation
by: Zhu, Ruijie, et al.
Published: (2024)
by: Zhu, Ruijie, et al.
Published: (2024)
Similar Items
-
Vision-Language Embodiment for Monocular Depth Estimation
by: Zhang, Jinchang, et al.
Published: (2025) -
Depth Estimation Based on 3D Gaussian Splatting Siamese Defocus
by: Zhang, Jinchang, et al.
Published: (2024) -
Language-Depth Navigated Thermal and Visible Image Fusion
by: Zhang, Jinchang, et al.
Published: (2025) -
Underground Mapping and Localization Based on Ground-Penetrating Radar
by: Zhang, Jinchang, et al.
Published: (2024) -
Embodiment: Self-Supervised Depth Estimation Based on Camera Models
by: Zhang, Jinchang, et al.
Published: (2024)