Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hansen-Estruch, Philippe, Yan, David, Chung, Ching-Yao, Zohar, Orr, Wang, Jialiang, Hou, Tingbo, Xu, Tao, Vishwanath, Sriram, Vajda, Peter, Chen, Xinlei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
di: Adiban, Mohammad, et al.
Pubblicazione: (2023)
di: Adiban, Mohammad, et al.
Pubblicazione: (2023)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
FLD+: Data-efficient Evaluation Metric for Generative Models
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
Normalizing Flow-Based Metric for Image Generation
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
Differentiable Hierarchical Visual Tokenization
di: Aasan, Marius, et al.
Pubblicazione: (2025)
di: Aasan, Marius, et al.
Pubblicazione: (2025)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
di: Zhao, Brian Nlong, et al.
Pubblicazione: (2025)
di: Zhao, Brian Nlong, et al.
Pubblicazione: (2025)
Task Singular Vectors: Reducing Task Interference in Model Merging
di: Gargiulo, Antonio Andrea, et al.
Pubblicazione: (2024)
di: Gargiulo, Antonio Andrea, et al.
Pubblicazione: (2024)
Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers
di: Zhao, Yiming
Pubblicazione: (2026)
di: Zhao, Yiming
Pubblicazione: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
di: Gupta, Sunny, et al.
Pubblicazione: (2024)
di: Gupta, Sunny, et al.
Pubblicazione: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
di: Kashyap, Pankhi, et al.
Pubblicazione: (2024)
di: Kashyap, Pankhi, et al.
Pubblicazione: (2024)
Is Single-View Mesh Reconstruction Ready for Robotics?
di: Nolte, Frederik, et al.
Pubblicazione: (2025)
di: Nolte, Frederik, et al.
Pubblicazione: (2025)
Revealing an Unattractivity Bias in Mental Reconstruction of Occluded Faces using Generative Image Models
di: Riedmann, Frederik, et al.
Pubblicazione: (2024)
di: Riedmann, Frederik, et al.
Pubblicazione: (2024)
Revisiting SVD and Wavelet Difference Reduction for Lossy Image Compression: A Reproducibility Study
di: Makarova, Alena
Pubblicazione: (2025)
di: Makarova, Alena
Pubblicazione: (2025)
Efficient Temporally-Aware DeepFake Detection using H.264 Motion Vectors
di: Grönquist, Peter, et al.
Pubblicazione: (2023)
di: Grönquist, Peter, et al.
Pubblicazione: (2023)
WaveMix: A Resource-efficient Neural Network for Image Analysis
di: Jeevan, Pranav, et al.
Pubblicazione: (2022)
di: Jeevan, Pranav, et al.
Pubblicazione: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
Removing Motion Artifact in MRI by Using a Perceptual Loss Driven Deep Learning Framework
di: Guo, Ziheng, et al.
Pubblicazione: (2026)
di: Guo, Ziheng, et al.
Pubblicazione: (2026)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
di: Benfeghoul, Martin, et al.
Pubblicazione: (2024)
di: Benfeghoul, Martin, et al.
Pubblicazione: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
di: Portelance, Eva, et al.
Pubblicazione: (2024)
di: Portelance, Eva, et al.
Pubblicazione: (2024)
Generalization performance of neural mapping schemes for the space-time interpolation of satellite-derived ocean colour datasets
di: Nguyen, Thi Thuy Nga, et al.
Pubblicazione: (2025)
di: Nguyen, Thi Thuy Nga, et al.
Pubblicazione: (2025)
Observation-only learning of neural mapping schemes for gappy satellite-derived ocean colour parameters
di: Dorffer, Clément, et al.
Pubblicazione: (2025)
di: Dorffer, Clément, et al.
Pubblicazione: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
3DGS-to-PC: Convert a 3D Gaussian Splatting Scene into a Dense Point Cloud or Mesh
di: Stuart, Lewis A G, et al.
Pubblicazione: (2025)
di: Stuart, Lewis A G, et al.
Pubblicazione: (2025)
FACT: Multinomial Misalignment Classification for Point Cloud Registration
di: Dillén, Ludvig, et al.
Pubblicazione: (2025)
di: Dillén, Ludvig, et al.
Pubblicazione: (2025)
Illumination and Shadows in Head Rotation: experiments with Denoising Diffusion Models
di: Asperti, Andrea, et al.
Pubblicazione: (2023)
di: Asperti, Andrea, et al.
Pubblicazione: (2023)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
di: Wang, Fusang, et al.
Pubblicazione: (2026)
di: Wang, Fusang, et al.
Pubblicazione: (2026)
Flexible-weighted Chamfer Distance: Enhanced Objective Function for Point Cloud Completion
di: Li, Jie, et al.
Pubblicazione: (2025)
di: Li, Jie, et al.
Pubblicazione: (2025)
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
di: Fang, Tiancheng, et al.
Pubblicazione: (2026)
di: Fang, Tiancheng, et al.
Pubblicazione: (2026)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
di: Hu, Chenggong, et al.
Pubblicazione: (2026)
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
di: Yu, Sicheng, et al.
Pubblicazione: (2025)
di: Yu, Sicheng, et al.
Pubblicazione: (2025)
J-NeuS: Joint field optimization for Neural Surface reconstruction in urban scenes with limited image overlap
di: Wang, Fusang, et al.
Pubblicazione: (2025)
di: Wang, Fusang, et al.
Pubblicazione: (2025)
FEDTAIL: Federated Long-Tailed Domain Generalization with Sharpness-Guided Gradient Matching
di: Gupta, Sunny, et al.
Pubblicazione: (2025)
di: Gupta, Sunny, et al.
Pubblicazione: (2025)
A Landmark-Aware Visual Navigation Dataset
di: Johnson, Faith, et al.
Pubblicazione: (2024)
di: Johnson, Faith, et al.
Pubblicazione: (2024)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
di: Dokme, Atahan, et al.
Pubblicazione: (2026)
di: Dokme, Atahan, et al.
Pubblicazione: (2026)
Classification of Cattle Behavior and Detection of Heat (Estrus) using Sensor Data
di: Dhakshinamoorthy, Druva, et al.
Pubblicazione: (2025)
di: Dhakshinamoorthy, Druva, et al.
Pubblicazione: (2025)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
Evaluation Metric for Quality Control and Generative Models in Histopathology Images
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)
Documenti analoghi
-
S-HR-VQVAE: Sequential Hierarchical Residual Learning Vector Quantized Variational Autoencoder for Video Prediction
di: Adiban, Mohammad, et al.
Pubblicazione: (2023) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
FLD+: Data-efficient Evaluation Metric for Generative Models
di: Jeevan, Pranav, et al.
Pubblicazione: (2024) -
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
di: Jeevan, Pranav, et al.
Pubblicazione: (2024) -
Normalizing Flow-Based Metric for Image Generation
di: Jeevan, Pranav, et al.
Pubblicazione: (2024)