Attentive VQ-VAE
Fuente:
arXiv
Saved in:
| Main Authors: | Hoyos, Angello, Rivera, Mariano |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
by: Rivera, Mariano, et al.
Published: (2025)
by: Rivera, Mariano, et al.
Published: (2025)
Evaluation of Environmental Conditions on Object Detection using Oriented Bounding Boxes for AR Applications
by: Li, Vladislav, et al.
Published: (2023)
by: Li, Vladislav, et al.
Published: (2023)
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
by: Herashchenko, Dmytro, et al.
Published: (2023)
by: Herashchenko, Dmytro, et al.
Published: (2023)
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
by: Ma, Jie, et al.
Published: (2023)
by: Ma, Jie, et al.
Published: (2023)
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
by: Agarwal, Rachit, et al.
Published: (2026)
by: Agarwal, Rachit, et al.
Published: (2026)
SITUATE -- Synthetic Object Counting Dataset for VLM training
by: Peinl, René, et al.
Published: (2026)
by: Peinl, René, et al.
Published: (2026)
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
Image Segmentation and Classification of E-waste for Training Robots for Waste Segregation
by: Tripathi, Prakriti
Published: (2025)
by: Tripathi, Prakriti
Published: (2025)
ProtoFlow: Interpretable and Robust Surgical Workflow Modeling with Learned Dynamic Scene Graph Prototypes
by: Holm, Felix, et al.
Published: (2025)
by: Holm, Felix, et al.
Published: (2025)
Siamese Networks for Cat Re-Identification: Exploring Neural Models for Cat Instance Recognition
by: Trein, Tobias, et al.
Published: (2025)
by: Trein, Tobias, et al.
Published: (2025)
From Prompt to Production:Automating Brand-Safe Marketing Imagery with Text-to-Image Models
by: Atighehchian, Parmida, et al.
Published: (2026)
by: Atighehchian, Parmida, et al.
Published: (2026)
MFTF: Mask-free Training-free Object Level Layout Control Diffusion Model
by: Yang, Shan
Published: (2024)
by: Yang, Shan
Published: (2024)
Sora as a World Model? A Complete Survey on Text-to-Video Generation
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
by: Puspitasari, Fachrina Dewi, et al.
Published: (2024)
TexTailor: Customized Text-aligned Texturing via Effective Resampling
by: Lee, Suin, et al.
Published: (2025)
by: Lee, Suin, et al.
Published: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
by: Chen, Zhangquan, et al.
Published: (2026)
by: Chen, Zhangquan, et al.
Published: (2026)
CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier
by: Ou, Ziyang
Published: (2025)
by: Ou, Ziyang
Published: (2025)
Rethinking Multimodal Point Cloud Completion: A Completion-by-Correction Perspective
by: Luo, Wang, et al.
Published: (2025)
by: Luo, Wang, et al.
Published: (2025)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
by: Safdar, Aon, et al.
Published: (2025)
by: Safdar, Aon, et al.
Published: (2025)
Next-Generation License Plate Detection and Recognition System using YOLOv8
by: Amin, Arslan, et al.
Published: (2025)
by: Amin, Arslan, et al.
Published: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
by: Chen, Yiping, et al.
Published: (2026)
by: Chen, Yiping, et al.
Published: (2026)
VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding
by: He, Jianxiang, et al.
Published: (2025)
by: He, Jianxiang, et al.
Published: (2025)
Instruction-based Image Editing with Planning, Reasoning, and Generation
by: Ji, Liya, et al.
Published: (2026)
by: Ji, Liya, et al.
Published: (2026)
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
by: Feng, Yigui, et al.
Published: (2026)
by: Feng, Yigui, et al.
Published: (2026)
Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based Customization
by: Liu, Yisu, et al.
Published: (2024)
by: Liu, Yisu, et al.
Published: (2024)
Unified Auto-Encoding with Masked Diffusion
by: Hansen-Estruch, Philippe, et al.
Published: (2024)
by: Hansen-Estruch, Philippe, et al.
Published: (2024)
Supervised Contrastive Learning for Few-Shot AI-Generated Image Detection and Attribution
by: Urueña, Jaime Álvarez, et al.
Published: (2025)
by: Urueña, Jaime Álvarez, et al.
Published: (2025)
Interpretable Tau-PET Synthesis from Multimodal T1-Weighted and FLAIR MRI Using Partial Information Decomposition Guided Disentangled Quantized Half-UNet
by: Chopra, Agamdeep S., et al.
Published: (2026)
by: Chopra, Agamdeep S., et al.
Published: (2026)
Invariant Representation via Decoupling Style and Spurious Features from Images
by: Li, Ruimeng, et al.
Published: (2023)
by: Li, Ruimeng, et al.
Published: (2023)
AI-Dentify: Deep learning for proximal caries detection on bitewing x-ray -- HUNT4 Oral Health Study
by: de Frutos, Javier Pérez, et al.
Published: (2023)
by: de Frutos, Javier Pérez, et al.
Published: (2023)
Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments
by: Gonzalez, Laura Alejandra Encinar, et al.
Published: (2025)
by: Gonzalez, Laura Alejandra Encinar, et al.
Published: (2025)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025)
by: Wang, Zhaohui, et al.
Published: (2025)
Deep Learning methodology for the identification of wood species using high-resolution macroscopic images
by: Herrera-Poyatos, David, et al.
Published: (2024)
by: Herrera-Poyatos, David, et al.
Published: (2024)
Vectra: A New Metric, Dataset, and Model for Visual Quality Assessment in E-Commerce In-Image Machine Translation
by: Wu, Qingyu, et al.
Published: (2026)
by: Wu, Qingyu, et al.
Published: (2026)
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges
by: Oliveira, Daniel A. P., et al.
Published: (2024)
by: Oliveira, Daniel A. P., et al.
Published: (2024)
ESCAPE: Energy-based Selective Adaptive Correction for Out-of-distribution 3D Human Pose Estimation
by: Bidulka, Luke, et al.
Published: (2024)
by: Bidulka, Luke, et al.
Published: (2024)
Unsupervised Decomposition and Recombination with Discriminator-Driven Diffusion Models
by: Wang, Archer, et al.
Published: (2026)
by: Wang, Archer, et al.
Published: (2026)
Learning to Seek Evidence: A Verifiable Reasoning Agent with Causal Faithfulness Analysis
by: Huang, Yuhang, et al.
Published: (2025)
by: Huang, Yuhang, et al.
Published: (2025)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
by: Bu, Weijue, et al.
Published: (2025)
by: Bu, Weijue, et al.
Published: (2025)
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
by: Khurdula, Harsha Vardhan, et al.
Published: (2024)
Similar Items
-
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
by: Rivera, Mariano, et al.
Published: (2025) -
Evaluation of Environmental Conditions on Object Detection using Oriented Bounding Boxes for AR Applications
by: Li, Vladislav, et al.
Published: (2023) -
Appearance-based gaze estimation enhanced with synthetic images using deep neural networks
by: Herashchenko, Dmytro, et al.
Published: (2023) -
Robust Visual Question Answering: Datasets, Methods, and Future Challenges
by: Ma, Jie, et al.
Published: (2023) -
Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation
by: Agarwal, Rachit, et al.
Published: (2026)