Saved in:
| Main Authors: | Li, Haotian, Jiao, Jianbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.14149 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
by: Sun, Qiyue, et al.
Published: (2025)
by: Sun, Qiyue, et al.
Published: (2025)
Exploring Image Representation with Decoupled Classical Visual Descriptors
by: Qu, Chenyuan, et al.
Published: (2025)
by: Qu, Chenyuan, et al.
Published: (2025)
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
by: Lin, Dongheng, et al.
Published: (2025)
by: Lin, Dongheng, et al.
Published: (2025)
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
by: Zheng, Xiaoyun, et al.
Published: (2025)
by: Zheng, Xiaoyun, et al.
Published: (2025)
One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition
by: Ismagilov, Timur, et al.
Published: (2026)
by: Ismagilov, Timur, et al.
Published: (2026)
Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression
by: Zhang, Haotian, et al.
Published: (2026)
by: Zhang, Haotian, et al.
Published: (2026)
Dynamic in Static: Hybrid Visual Correspondence for Self-Supervised Video Object Segmentation
by: Pei, Gensheng, et al.
Published: (2024)
by: Pei, Gensheng, et al.
Published: (2024)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Frame Interpolation with Consecutive Brownian Bridge Diffusion
by: Lyu, Zonglin, et al.
Published: (2024)
by: Lyu, Zonglin, et al.
Published: (2024)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
by: Ma, Yan, et al.
Published: (2025)
by: Ma, Yan, et al.
Published: (2025)
Enhancing Visual In-Context Learning by Multi-Faceted Fusion
by: Liao, Wenwen, et al.
Published: (2026)
by: Liao, Wenwen, et al.
Published: (2026)
Mimicking Human Visual Development for Learning Robust Image Representations
by: Raj, Ankita, et al.
Published: (2025)
by: Raj, Ankita, et al.
Published: (2025)
Exposing Hallucinations To Suppress Them: VLMs Representation Editing With Generative Anchors
by: Shi, Youxu, et al.
Published: (2025)
by: Shi, Youxu, et al.
Published: (2025)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
by: Vaishnav, Mohit, et al.
Published: (2026)
by: Vaishnav, Mohit, et al.
Published: (2026)
Image Generation Diversity Issues and How to Tame Them
by: Dombrowski, Mischa, et al.
Published: (2024)
by: Dombrowski, Mischa, et al.
Published: (2024)
Scaling Learned Image Compression Models up to 1 Billion
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
PVLR: Prompt-driven Visual-Linguistic Representation Learning for Multi-Label Image Recognition
by: Tan, Hao, et al.
Published: (2024)
by: Tan, Hao, et al.
Published: (2024)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
by: Guo, Wei, et al.
Published: (2024)
by: Guo, Wei, et al.
Published: (2024)
Structure-Aware Feature Rectification with Region Adjacency Graphs for Training-Free Open-Vocabulary Semantic Segmentation
by: Huang, Qiming, et al.
Published: (2025)
by: Huang, Qiming, et al.
Published: (2025)
LeOCLR: Leveraging Original Images for Contrastive Learning of Visual Representations
by: Alkhalefi, Mohammad, et al.
Published: (2024)
by: Alkhalefi, Mohammad, et al.
Published: (2024)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
by: Zhou, Jiazhou, et al.
Published: (2023)
by: Zhou, Jiazhou, et al.
Published: (2023)
AlignedCut: Visual Concepts Discovery on Brain-Guided Universal Feature Space
by: Yang, Huzheng, et al.
Published: (2024)
by: Yang, Huzheng, et al.
Published: (2024)
Neural Clustering based Visual Representation Learning
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
by: Xiao, Kelaiti, et al.
Published: (2025)
by: Xiao, Kelaiti, et al.
Published: (2025)
Few Exemplar-Based General Medical Image Segmentation via Domain-Aware Selective Adaptation
by: Xu, Chen, et al.
Published: (2024)
by: Xu, Chen, et al.
Published: (2024)
V2M: Visual 2-Dimensional Mamba for Image Representation Learning
by: Wang, Chengkun, et al.
Published: (2024)
by: Wang, Chengkun, et al.
Published: (2024)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
by: Lyu, Yuanhuiyi, et al.
Published: (2024)
Detecting Origin Attribution for Text-to-Image Diffusion Models
by: Xu, Katherine, et al.
Published: (2024)
by: Xu, Katherine, et al.
Published: (2024)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
Generalized Gaussian Model for Learned Image Compression
by: Zhang, Haotian, et al.
Published: (2024)
by: Zhang, Haotian, et al.
Published: (2024)
Towards Training-free Multimodal Hate Localisation with Large Language Models
by: Sun, Yueming, et al.
Published: (2026)
by: Sun, Yueming, et al.
Published: (2026)
MVD-HuGaS: Human Gaussians from a Single Image via 3D Human Multi-view Diffusion Prior
by: Xiong, Kaiqiang, et al.
Published: (2025)
by: Xiong, Kaiqiang, et al.
Published: (2025)
Paint by Inpaint: Learning to Add Image Objects by Removing Them First
by: Wasserman, Navve, et al.
Published: (2024)
by: Wasserman, Navve, et al.
Published: (2024)
CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization
by: Ding, Ziyang, et al.
Published: (2026)
by: Ding, Ziyang, et al.
Published: (2026)
ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
by: Liao, Wenjie, et al.
Published: (2025)
by: Liao, Wenjie, et al.
Published: (2025)
Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
by: Zhang, Wenqi, et al.
Published: (2024)
by: Zhang, Wenqi, et al.
Published: (2024)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
Multimodal Hate Detection Using Dual-Stream Graph Neural Networks
by: Yue, Jiangbei, et al.
Published: (2025)
by: Yue, Jiangbei, et al.
Published: (2025)
Similar Items
-
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
by: Sun, Qiyue, et al.
Published: (2025) -
Exploring Image Representation with Decoupled Classical Visual Descriptors
by: Qu, Chenyuan, et al.
Published: (2025) -
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
by: Lin, Dongheng, et al.
Published: (2025) -
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
by: Zheng, Xiaoyun, et al.
Published: (2025) -
One Channel to Rule Them All: Rethinking Input Representation for Visual Place Recognition
by: Ismagilov, Timur, et al.
Published: (2026)