VUGEN: Visual Understanding priors for GENeration
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Xiangyi, Vallaeys, Théophane, Elbayad, Maha, Nguyen, John, Verbeek, Jakob |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
di: Vallaeys, Théophane, et al.
Pubblicazione: (2024)
di: Vallaeys, Théophane, et al.
Pubblicazione: (2024)
Text-Guided Semantic Image Encoder
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2025)
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2025)
Beyond Language Modeling: An Exploration of Multimodal Pretraining
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
di: Tong, Shengbang, et al.
Pubblicazione: (2026)
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
di: Meric, Adil, et al.
Pubblicazione: (2026)
di: Meric, Adil, et al.
Pubblicazione: (2026)
Flowception: Temporally Expansive Flow Matching for Video Generation
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
Better (pseudo-)labels for semi-supervised instance segmentation
di: Porcher, François, et al.
Pubblicazione: (2024)
di: Porcher, François, et al.
Pubblicazione: (2024)
TV2TV: A Unified Framework for Interleaved Language and Video Generation
di: Han, Xiaochuang, et al.
Pubblicazione: (2025)
di: Han, Xiaochuang, et al.
Pubblicazione: (2025)
Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025)
Towards image compression with perfect realism at ultra-low bitrates
di: Careil, Marlène, et al.
Pubblicazione: (2023)
di: Careil, Marlène, et al.
Pubblicazione: (2023)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
di: Berrada, Tariq, et al.
Pubblicazione: (2023)
ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
di: Pham, Huy Quang, et al.
Pubblicazione: (2024)
MoME: Mixture of Visual Language Medical Experts for Medical Imaging Segmentation
di: Rezvani, Arghavan, et al.
Pubblicazione: (2025)
di: Rezvani, Arghavan, et al.
Pubblicazione: (2025)
Entropy Rectifying Guidance for Diffusion and Flow Models
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
di: Ifriqi, Tariq Berrada, et al.
Pubblicazione: (2025)
Analysis of 3D Urticaceae Pollen Classification Using Deep Learning Models
di: Konijn, Tijs, et al.
Pubblicazione: (2025)
di: Konijn, Tijs, et al.
Pubblicazione: (2025)
Insect-Foundation: A Foundation Model and Large-scale 1M Dataset for Visual Insect Understanding
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2023)
di: Nguyen, Hoang-Quan, et al.
Pubblicazione: (2023)
Increasing the Utility of Synthetic Images through Chamfer Guidance
di: Dall'Asen, Nicola, et al.
Pubblicazione: (2025)
di: Dall'Asen, Nicola, et al.
Pubblicazione: (2025)
Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments
di: Nguyen, Van Quang
Pubblicazione: (2026)
di: Nguyen, Van Quang
Pubblicazione: (2026)
Semantically-aware Neural Radiance Fields for Visual Scene Understanding: A Comprehensive Review
di: Nguyen, Thang-Anh-Quan, et al.
Pubblicazione: (2024)
di: Nguyen, Thang-Anh-Quan, et al.
Pubblicazione: (2024)
Boosting Latent Diffusion with Perceptual Objectives
di: Berrada, Tariq, et al.
Pubblicazione: (2024)
di: Berrada, Tariq, et al.
Pubblicazione: (2024)
Understanding and Evaluating Hallucinations in 3D Visual Language Models
di: Peng, Ruiying, et al.
Pubblicazione: (2025)
di: Peng, Ruiying, et al.
Pubblicazione: (2025)
Towards Visual Syntactical Understanding
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2024)
di: Chowdhury, Sayeed Shafayet, et al.
Pubblicazione: (2024)
Towards More Unified In-context Visual Understanding
di: Sheng, Dianmo, et al.
Pubblicazione: (2023)
di: Sheng, Dianmo, et al.
Pubblicazione: (2023)
DeepTraverse: A Depth-First Search Inspired Network for Algorithmic Visual Understanding
di: Guo, Bin, et al.
Pubblicazione: (2025)
di: Guo, Bin, et al.
Pubblicazione: (2025)
Exploring Clustering Capability of Inpainting Model Embeddings for Pattern-based Individual Identification
di: van Bijsterveld, Jens, et al.
Pubblicazione: (2026)
di: van Bijsterveld, Jens, et al.
Pubblicazione: (2026)
Multi-view Remote Sensing Image Segmentation With SAM priors
di: Qi, Zipeng, et al.
Pubblicazione: (2024)
di: Qi, Zipeng, et al.
Pubblicazione: (2024)
Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
di: Chen, Yizhu, et al.
Pubblicazione: (2025)
di: Chen, Yizhu, et al.
Pubblicazione: (2025)
VisKnow: Constructing Visual Knowledge Base for Object Understanding
di: Yao, Ziwei, et al.
Pubblicazione: (2025)
di: Yao, Ziwei, et al.
Pubblicazione: (2025)
VirDA: Reusing Backbone for Unsupervised Domain Adaptation with Visual Reprogramming
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
Video Understanding: Through A Temporal Lens
di: Nguyen, Thong Thanh
Pubblicazione: (2026)
di: Nguyen, Thong Thanh
Pubblicazione: (2026)
SynthDoc: Bilingual Documents Synthesis for Visual Document Understanding
di: Ding, Chuanghao, et al.
Pubblicazione: (2024)
di: Ding, Chuanghao, et al.
Pubblicazione: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
di: Wu, Size, et al.
Pubblicazione: (2025)
di: Wu, Size, et al.
Pubblicazione: (2025)
MosaicDoc: A Large-Scale Bilingual Benchmark for Visually Rich Document Understanding
di: Chen, Ketong, et al.
Pubblicazione: (2025)
di: Chen, Ketong, et al.
Pubblicazione: (2025)
Visual Context Window Extension: A New Perspective for Long Video Understanding
di: Wei, Hongchen, et al.
Pubblicazione: (2024)
di: Wei, Hongchen, et al.
Pubblicazione: (2024)
Domain Generalization through Spatial Relation Induction over Visual Primitives
di: Nguyen, Dat, et al.
Pubblicazione: (2026)
di: Nguyen, Dat, et al.
Pubblicazione: (2026)
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding
di: Xu, Pengxin, et al.
Pubblicazione: (2026)
di: Xu, Pengxin, et al.
Pubblicazione: (2026)
EVLM: An Efficient Vision-Language Model for Visual Understanding
di: Chen, Kaibing, et al.
Pubblicazione: (2024)
di: Chen, Kaibing, et al.
Pubblicazione: (2024)
HIG: Hierarchical Interlacement Graph Approach to Scene Graph Generation in Video Understanding
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2023)
di: Nguyen, Trong-Thuan, et al.
Pubblicazione: (2023)
Exploiting Text-Image Latent Spaces for the Description of Visual Concepts
di: Schmalwasser, Laines, et al.
Pubblicazione: (2024)
di: Schmalwasser, Laines, et al.
Pubblicazione: (2024)
Surgical Visual Understanding (SurgVU) Dataset
di: Zia, Aneeq, et al.
Pubblicazione: (2025)
di: Zia, Aneeq, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
di: Vallaeys, Théophane, et al.
Pubblicazione: (2025) -
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
di: Vallaeys, Théophane, et al.
Pubblicazione: (2024) -
Text-Guided Semantic Image Encoder
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2025) -
Beyond Language Modeling: An Exploration of Multimodal Pretraining
di: Tong, Shengbang, et al.
Pubblicazione: (2026) -
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
di: Meric, Adil, et al.
Pubblicazione: (2026)