ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Yin, Yang, Kaicheng, Liang, Peirou, An, Xiang, Zhao, Yongle, Wang, Yumeng, Feng, Ziyong, Miles, Roy, Elezi, Ismail, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
by: Miles, Roy, et al.
Published: (2024)
by: Miles, Roy, et al.
Published: (2024)
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
by: Miles, Roy, et al.
Published: (2024)
by: Miles, Roy, et al.
Published: (2024)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025)
by: Toker, Aysim, et al.
Published: (2025)
G3DR: Generative 3D Reconstruction in ImageNet
by: Reddy, Pradyumna, et al.
Published: (2024)
by: Reddy, Pradyumna, et al.
Published: (2024)
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
Deep Active Learning: A Reality Check
by: Gashi, Edrina, et al.
Published: (2024)
by: Gashi, Edrina, et al.
Published: (2024)
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
by: Ye-Bin, Moon, et al.
Published: (2025)
by: Ye-Bin, Moon, et al.
Published: (2025)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
by: Choi, Yura, et al.
Published: (2026)
by: Choi, Yura, et al.
Published: (2026)
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
by: Miles, Roy, et al.
Published: (2026)
by: Miles, Roy, et al.
Published: (2026)
PaCo-FR: Patch-Pixel Aligned End-to-End Codebook Learning for Facial Representation Pre-training
by: Xie, Yin, et al.
Published: (2025)
by: Xie, Yin, et al.
Published: (2025)
From Attention to Activation: Unravelling the Enigmas of Large Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
Efficient, Validation-Free Intrinsic Quality Estimation for Large-Scale Face Recognition Datasets
by: Chen, Zhichao, et al.
Published: (2026)
by: Chen, Zhichao, et al.
Published: (2026)
IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models
by: Cui, Siying, et al.
Published: (2024)
by: Cui, Siying, et al.
Published: (2024)
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
by: Ma, Chengcheng, et al.
Published: (2023)
by: Ma, Chengcheng, et al.
Published: (2023)
"Principal Components" Enable A New Language of Images
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
RWKV-CLIP: A Robust Vision-Language Representation Learner
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
Fractal Calibration for long-tailed object detection
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
by: Alexandridis, Konstantinos Panagiotis, et al.
Published: (2024)
CLIP-CID: Efficient CLIP Distillation via Cluster-Instance Discrimination
by: Yang, Kaicheng, et al.
Published: (2024)
by: Yang, Kaicheng, et al.
Published: (2024)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
High-Fidelity Facial Albedo Estimation via Texture Quantization
by: Ran, Zimin, et al.
Published: (2024)
by: Ran, Zimin, et al.
Published: (2024)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
MidSteer: Optimal Affine Framework for Steering Generative Models
by: Gaintseva, Tatiana, et al.
Published: (2026)
by: Gaintseva, Tatiana, et al.
Published: (2026)
DeepStack: Deeply Stacking Visual Tokens is Surprisingly Simple and Effective for LMMs
by: Meng, Lingchen, et al.
Published: (2024)
by: Meng, Lingchen, et al.
Published: (2024)
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
CASteer: Cross-Attention Steering for Controllable Concept Erasure
by: Gaintseva, Tatiana, et al.
Published: (2025)
by: Gaintseva, Tatiana, et al.
Published: (2025)
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
by: Khan, Mohammad Sadil, et al.
Published: (2026)
by: Khan, Mohammad Sadil, et al.
Published: (2026)
MMSearch-R1: Incentivizing LMMs to Search
by: Wu, Jinming, et al.
Published: (2025)
by: Wu, Jinming, et al.
Published: (2025)
Decoupled Global-Local Alignment for Improving Compositional Understanding
by: Hu, Xiaoxing, et al.
Published: (2025)
by: Hu, Xiaoxing, et al.
Published: (2025)
Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval
by: Zheng, Tianlu, et al.
Published: (2025)
by: Zheng, Tianlu, et al.
Published: (2025)
ORID: Organ-Regional Information Driven Framework for Radiology Report Generation
by: Gu, Tiancheng, et al.
Published: (2024)
by: Gu, Tiancheng, et al.
Published: (2024)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
by: Yang, Sihan, et al.
Published: (2025)
by: Yang, Sihan, et al.
Published: (2025)
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos
by: Zhang, Jinglei, et al.
Published: (2025)
by: Zhang, Jinglei, et al.
Published: (2025)
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
by: Zhan, Wengyi, et al.
Published: (2025)
by: Zhan, Wengyi, et al.
Published: (2025)
Fusing Pretrained ViTs with TCNet for Enhanced EEG Regression
by: Modesitt, Eric, et al.
Published: (2024)
by: Modesitt, Eric, et al.
Published: (2024)
1st Place Solution to the 1st SkatingVerse Challenge
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
Similar Items
-
Region-based Cluster Discrimination for Visual Representation Learning
by: Xie, Yin, et al.
Published: (2025) -
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
by: Miles, Roy, et al.
Published: (2024) -
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
by: Miles, Roy, et al.
Published: (2024) -
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025) -
G3DR: Generative 3D Reconstruction in ImageNet
by: Reddy, Pradyumna, et al.
Published: (2024)