ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Juntian, Jin, Song, Cheng, Chuanqi, Liu, Yuhan, Lin, Yankai, Zhang, Xun, Zhang, Yufei, Jiang, Fei, Yin, Guojun, Lin, Wei, Yan, Rui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
by: Jin, Song, et al.
Published: (2025)
by: Jin, Song, et al.
Published: (2025)
Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
by: Jin, Song, et al.
Published: (2025)
by: Jin, Song, et al.
Published: (2025)
DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain
by: Jin, Song, et al.
Published: (2026)
by: Jin, Song, et al.
Published: (2026)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
by: Zhang, Juntian, et al.
Published: (2025)
by: Zhang, Juntian, et al.
Published: (2025)
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
by: Wang, Yubo, et al.
Published: (2026)
by: Wang, Yubo, et al.
Published: (2026)
Semantic Convergence: Harmonizing Recommender Systems via Two-Stage Alignment and Behavioral Semantic Tokenization
by: Li, Guanghan, et al.
Published: (2024)
by: Li, Guanghan, et al.
Published: (2024)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
by: Wu, Songhao, et al.
Published: (2025)
by: Wu, Songhao, et al.
Published: (2025)
VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
by: Wang, Qiuchen, et al.
Published: (2025)
by: Wang, Qiuchen, et al.
Published: (2025)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
by: Jin, Peng, et al.
Published: (2023)
by: Jin, Peng, et al.
Published: (2023)
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
by: Zhang, Jiayu, et al.
Published: (2025)
by: Zhang, Jiayu, et al.
Published: (2025)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
From Experience to Strategy: Empowering LLM Agents with Trainable Graph Memory
by: Xia, Siyu, et al.
Published: (2025)
by: Xia, Siyu, et al.
Published: (2025)
ViSymRe: Vision Multimodal Symbolic Regression
by: Li, Da, et al.
Published: (2024)
by: Li, Da, et al.
Published: (2024)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
by: Kamoi, Ryo, et al.
Published: (2024)
by: Kamoi, Ryo, et al.
Published: (2024)
DARC: Decoupled Asymmetric Reasoning Curriculum for LLM Evolution
by: Fan, Shengda, et al.
Published: (2026)
by: Fan, Shengda, et al.
Published: (2026)
ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction
by: Feng, Yi, et al.
Published: (2024)
by: Feng, Yi, et al.
Published: (2024)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
by: Li, Zhuohao, et al.
Published: (2026)
by: Li, Zhuohao, et al.
Published: (2026)
Beyond the Surface: Measuring Self-Preference in LLM Judgments
by: Chen, Zhi-Yuan, et al.
Published: (2025)
by: Chen, Zhi-Yuan, et al.
Published: (2025)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
by: Zhang, Kaiyi, et al.
Published: (2026)
by: Zhang, Kaiyi, et al.
Published: (2026)
FineViT: Progressively Unlocking Fine-Grained Perception with Dense Recaptions
by: Zhao, Peisen, et al.
Published: (2026)
by: Zhao, Peisen, et al.
Published: (2026)
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
by: Zhu, Yiqi, et al.
Published: (2025)
by: Zhu, Yiqi, et al.
Published: (2025)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025)
by: Zheng, Guanghao, et al.
Published: (2025)
EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning
by: Lin, Bingqian, et al.
Published: (2025)
by: Lin, Bingqian, et al.
Published: (2025)
Predicting Emergent Abilities with Infinite Resolution Evaluation
by: Hu, Shengding, et al.
Published: (2023)
by: Hu, Shengding, et al.
Published: (2023)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
by: You, Haoran, et al.
Published: (2022)
by: You, Haoran, et al.
Published: (2022)
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
by: Wang, Xiyao, et al.
Published: (2025)
by: Wang, Xiyao, et al.
Published: (2025)
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
by: Song, Yuehao, et al.
Published: (2024)
by: Song, Yuehao, et al.
Published: (2024)
Emergence of Painting Ability via Recognition-Driven Evolution
by: Lin, Yi, et al.
Published: (2025)
by: Lin, Yi, et al.
Published: (2025)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
by: Zhang, Yaocheng, et al.
Published: (2026)
by: Zhang, Yaocheng, et al.
Published: (2026)
ICLEval: Evaluating In-Context Learning Ability of Large Language Models
by: Chen, Wentong, et al.
Published: (2024)
by: Chen, Wentong, et al.
Published: (2024)
Lyrics: Boosting Fine-grained Language-Vision Alignment and Comprehension via Semantic-aware Visual Objects
by: Lu, Junyu, et al.
Published: (2023)
by: Lu, Junyu, et al.
Published: (2023)
Synergistic Prompting for Robust Visual Recognition with Missing Modalities
by: Zhang, Zhihui, et al.
Published: (2025)
by: Zhang, Zhihui, et al.
Published: (2025)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
by: Tang, Zicong, et al.
Published: (2025)
by: Tang, Zicong, et al.
Published: (2025)
From the Least to the Most: Building a Plug-and-Play Visual Reasoner via Data Synthesis
by: Cheng, Chuanqi, et al.
Published: (2024)
by: Cheng, Chuanqi, et al.
Published: (2024)
Empowering Segmentation Ability to Multi-modal Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
by: Zhou, Guanyu, et al.
Published: (2026)
by: Zhou, Guanyu, et al.
Published: (2026)
Similar Items
-
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
by: Jin, Song, et al.
Published: (2025) -
Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning
by: Jin, Song, et al.
Published: (2025) -
DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain
by: Jin, Song, et al.
Published: (2026) -
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
by: Zhang, Juntian, et al.
Published: (2025) -
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
by: Wang, Yubo, et al.
Published: (2026)