Diversity Has Always Been There in Your Visual Autoregressive Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Tong, Yang, Guanyu, Liu, Nian, Wang, Kai, Wang, Yaxing, Shaker, Abdelrahman M, Khan, Salman, Khan, Fahad Shahbaz, Li, Senmao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
GroupMamba: Efficient Group-Based Visual State Space Model
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
von: Wang, Tong, et al.
Veröffentlicht: (2025)
von: Wang, Tong, et al.
Veröffentlicht: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2022)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2022)
WorldCache: Content-Aware Caching for Accelerated Video World Models
von: Nawaz, Umair, et al.
Veröffentlicht: (2026)
von: Nawaz, Umair, et al.
Veröffentlicht: (2026)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2025)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
von: Ahmad, Ghazi Shazan, et al.
Veröffentlicht: (2025)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2024)
von: Li, Senmao, et al.
Veröffentlicht: (2024)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2026)
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2023)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
Towards Evaluating the Robustness of Visual State Space Models
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
von: Maaz, Muhammad, et al.
Veröffentlicht: (2023)
CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
von: Heakl, Ahmed, et al.
Veröffentlicht: (2026)
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
von: Luo, Ziyang, et al.
Veröffentlicht: (2025)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
von: Khattak, Muhammad Uzair, et al.
Veröffentlicht: (2024)
Language Guided Domain Generalized Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
von: Dharmasiri, Amaya, et al.
Veröffentlicht: (2024)
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
von: Li, Senmao, et al.
Veröffentlicht: (2023)
von: Li, Senmao, et al.
Veröffentlicht: (2023)
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
von: Thawakar, Omkar, et al.
Veröffentlicht: (2025)
Not All Modalities Are Equal: Instruction-Aware Gating for Multimodal Videos
von: Ding, Bonan, et al.
Veröffentlicht: (2026)
von: Ding, Bonan, et al.
Veröffentlicht: (2026)
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
von: Li, Senmao, et al.
Veröffentlicht: (2023)
von: Li, Senmao, et al.
Veröffentlicht: (2023)
Enhancing Novel Object Detection via Cooperative Foundational Models
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
von: Bharadwaj, Rohit, et al.
Veröffentlicht: (2023)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
von: Chen, Shiming, et al.
Veröffentlicht: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
von: Chen, Shiming, et al.
Veröffentlicht: (2024)
von: Chen, Shiming, et al.
Veröffentlicht: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
von: Liu, Tao, et al.
Veröffentlicht: (2025)
von: Liu, Tao, et al.
Veröffentlicht: (2025)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
von: Hassan, Jameel, et al.
Veröffentlicht: (2023)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
von: Malik, Hashmat Shadab, et al.
Veröffentlicht: (2025)
Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and Method
von: Ye, Shuo, et al.
Veröffentlicht: (2023)
von: Ye, Shuo, et al.
Veröffentlicht: (2023)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
von: Mahmood, Ahmad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
von: Li, Senmao, et al.
Veröffentlicht: (2025) -
GroupMamba: Efficient Group-Based Visual State Space Model
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2024) -
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
von: Shaker, Abdelrahman, et al.
Veröffentlicht: (2025) -
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
von: Wang, Tong, et al.
Veröffentlicht: (2025) -
Learnable Weight Initialization for Volumetric Medical Image Segmentation
von: Kunhimon, Shahina, et al.
Veröffentlicht: (2023)