Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Dianyi, Song, Wei, Wang, Yikun, Wang, Siyuan, Yu, Kaicheng, Wei, Zhongyu, Wang, Jiaqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
di: Wang, Dianyi, et al.
Pubblicazione: (2026)
di: Wang, Dianyi, et al.
Pubblicazione: (2026)
AdaCodec: A Predictive Visual Code for Video MLLMs
di: Hou, Haowen, et al.
Pubblicazione: (2026)
di: Hou, Haowen, et al.
Pubblicazione: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
di: Song, Wei, et al.
Pubblicazione: (2025)
di: Song, Wei, et al.
Pubblicazione: (2025)
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
di: Wang, Dianyi, et al.
Pubblicazione: (2026)
di: Wang, Dianyi, et al.
Pubblicazione: (2026)
Evolutionary Negative Module Pruning for Better LoRA Merging
di: Cao, Anda, et al.
Pubblicazione: (2026)
di: Cao, Anda, et al.
Pubblicazione: (2026)
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
di: Zhang, Jiwen, et al.
Pubblicazione: (2026)
OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding
di: Yang, Dianyi, et al.
Pubblicazione: (2025)
di: Yang, Dianyi, et al.
Pubblicazione: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
di: Lin, Bin, et al.
Pubblicazione: (2025)
di: Lin, Bin, et al.
Pubblicazione: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
di: Qin, Yiming, et al.
Pubblicazione: (2025)
di: Qin, Yiming, et al.
Pubblicazione: (2025)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
di: Li, Zejun, et al.
Pubblicazione: (2024)
di: Li, Zejun, et al.
Pubblicazione: (2024)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
di: Song, Zijia, et al.
Pubblicazione: (2024)
di: Song, Zijia, et al.
Pubblicazione: (2024)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
di: Lu, Meng, et al.
Pubblicazione: (2025)
di: Lu, Meng, et al.
Pubblicazione: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
di: Du, Mengfei, et al.
Pubblicazione: (2024)
di: Du, Mengfei, et al.
Pubblicazione: (2024)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
di: Mayer, Julius, et al.
Pubblicazione: (2025)
di: Mayer, Julius, et al.
Pubblicazione: (2025)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
di: Zhou, Guanyu, et al.
Pubblicazione: (2026)
Can VLMs Recall Factual Associations From Visual References?
di: Ashok, Dhananjay, et al.
Pubblicazione: (2025)
di: Ashok, Dhananjay, et al.
Pubblicazione: (2025)
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
BabyVLM: Data-Efficient Pretraining of VLMs Inspired by Infant Learning
di: Wang, Shengao, et al.
Pubblicazione: (2025)
di: Wang, Shengao, et al.
Pubblicazione: (2025)
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
di: Ma, Yiyang, et al.
Pubblicazione: (2024)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
di: Li, Zejun, et al.
Pubblicazione: (2025)
di: Li, Zejun, et al.
Pubblicazione: (2025)
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
Valley: Video Assistant with Large Language model Enhanced abilitY
di: Luo, Ruipu, et al.
Pubblicazione: (2023)
di: Luo, Ruipu, et al.
Pubblicazione: (2023)
Channel-wise Vector Quantization
di: Song, Wei, et al.
Pubblicazione: (2026)
di: Song, Wei, et al.
Pubblicazione: (2026)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
di: Li, Sifan, et al.
Pubblicazione: (2025)
di: Li, Sifan, et al.
Pubblicazione: (2025)
Semantic Is Enough: Only Semantic Information For NeRF Reconstruction
di: Wang, Ruibo, et al.
Pubblicazione: (2024)
di: Wang, Ruibo, et al.
Pubblicazione: (2024)
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
di: Feng, Sicheng, et al.
Pubblicazione: (2025)
di: Feng, Sicheng, et al.
Pubblicazione: (2025)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
di: Chen, Haonan, et al.
Pubblicazione: (2025)
di: Chen, Haonan, et al.
Pubblicazione: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
di: Yu, Hao, et al.
Pubblicazione: (2025)
di: Yu, Hao, et al.
Pubblicazione: (2025)
Reconstructive Visual Instruction Tuning
di: Wang, Haochen, et al.
Pubblicazione: (2024)
di: Wang, Haochen, et al.
Pubblicazione: (2024)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
di: Zhang, Yue, et al.
Pubblicazione: (2026)
di: Zhang, Yue, et al.
Pubblicazione: (2026)
Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
di: Choi, Dasol, et al.
Pubblicazione: (2025)
di: Choi, Dasol, et al.
Pubblicazione: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
di: Wang, Qiuchen, et al.
Pubblicazione: (2025)
Universal Approximation of Visual Autoregressive Transformers
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
di: Zhou, Kaiwen, et al.
Pubblicazione: (2023)
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
di: Wang, Siyuan, et al.
Pubblicazione: (2024)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
di: Avogaro, Niccolo, et al.
Pubblicazione: (2026)
di: Avogaro, Niccolo, et al.
Pubblicazione: (2026)
Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
di: Wang, Haochen, et al.
Pubblicazione: (2025)
di: Wang, Haochen, et al.
Pubblicazione: (2025)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
di: Dong, Shuai, et al.
Pubblicazione: (2025)
di: Dong, Shuai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
di: Wang, Dianyi, et al.
Pubblicazione: (2026) -
AdaCodec: A Predictive Visual Code for Video MLLMs
di: Hou, Haowen, et al.
Pubblicazione: (2026) -
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
di: Song, Wei, et al.
Pubblicazione: (2025) -
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing
di: Wang, Dianyi, et al.
Pubblicazione: (2026) -
Evolutionary Negative Module Pruning for Better LoRA Merging
di: Cao, Anda, et al.
Pubblicazione: (2026)