TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
Fuente:
arXiv
Saved in:
| Main Author: | Stergiou, Alexandros |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAVIB: A Large-scale Video Interpolation Benchmark
by: Stergiou, Alexandros
Published: (2024)
by: Stergiou, Alexandros
Published: (2024)
MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
by: Jain, Animesh, et al.
Published: (2025)
by: Jain, Animesh, et al.
Published: (2025)
About Time: Advances, Challenges, and Outlooks of Action Understanding
by: Stergiou, Alexandros, et al.
Published: (2024)
by: Stergiou, Alexandros, et al.
Published: (2024)
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
by: Sinha, Saptarshi, et al.
Published: (2024)
by: Sinha, Saptarshi, et al.
Published: (2024)
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)
by: Qiu, Jason, et al.
Published: (2026)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
Does VLM Classification Benefit from LLM Description Semantics?
by: Ma, Pingchuan, et al.
Published: (2024)
by: Ma, Pingchuan, et al.
Published: (2024)
Geometrically Regularized Transfer Learning with On-Manifold and Off-Manifold Perturbation
by: Satou, Hana, et al.
Published: (2025)
by: Satou, Hana, et al.
Published: (2025)
Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
by: Wang, Wentao, et al.
Published: (2025)
by: Wang, Wentao, et al.
Published: (2025)
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
by: Bao, Wentao, et al.
Published: (2024)
by: Bao, Wentao, et al.
Published: (2024)
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
by: Xu, Runsen, et al.
Published: (2024)
by: Xu, Runsen, et al.
Published: (2024)
PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
by: Chang, Zewei, et al.
Published: (2025)
by: Chang, Zewei, et al.
Published: (2025)
Can Better Text Semantics in Prompt Tuning Improve VLM Generalization?
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
by: Kuchibhotla, Hari Chandana, et al.
Published: (2024)
PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
by: Li, Yuliang, et al.
Published: (2026)
by: Li, Yuliang, et al.
Published: (2026)
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
by: Xie, Rongchang, et al.
Published: (2024)
by: Xie, Rongchang, et al.
Published: (2024)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
edgeVLM: Cloud-edge Collaborative Real-time VLM based on Context Transfer
by: Qian, Chen, et al.
Published: (2025)
by: Qian, Chen, et al.
Published: (2025)
What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing
by: Lin, Hangyu, et al.
Published: (2026)
by: Lin, Hangyu, et al.
Published: (2026)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
by: Hu, Junshan, et al.
Published: (2025)
by: Hu, Junshan, et al.
Published: (2025)
REF-VLM: Triplet-Based Referring Paradigm for Unified Visual Decoding
by: Tai, Yan, et al.
Published: (2025)
by: Tai, Yan, et al.
Published: (2025)
CogVLM2: Visual Language Models for Image and Video Understanding
by: Hong, Wenyi, et al.
Published: (2024)
by: Hong, Wenyi, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Enhancing End-to-End Autonomous Driving with Risk Semantic Distillaion from VLM
by: Qin, Jack, et al.
Published: (2025)
by: Qin, Jack, et al.
Published: (2025)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
A VLM-based Method for Visual Anomaly Detection in Robotic Scientific Laboratories
by: Lin, Shiwei, et al.
Published: (2025)
by: Lin, Shiwei, et al.
Published: (2025)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models
by: Qiu, Weijie, et al.
Published: (2026)
by: Qiu, Weijie, et al.
Published: (2026)
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
by: Wang, Pan, et al.
Published: (2026)
by: Wang, Pan, et al.
Published: (2026)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
by: Zhao, Yi, et al.
Published: (2026)
by: Zhao, Yi, et al.
Published: (2026)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
by: Xue, Xizhe, et al.
Published: (2024)
by: Xue, Xizhe, et al.
Published: (2024)
SemFaceEdit: Semantic Face Editing on Generative Radiance Manifolds
by: Verma, Shashikant, et al.
Published: (2025)
by: Verma, Shashikant, et al.
Published: (2025)
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
by: Waheed, Sania, et al.
Published: (2025)
by: Waheed, Sania, et al.
Published: (2025)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM
by: Chen, Tao, et al.
Published: (2025)
by: Chen, Tao, et al.
Published: (2025)
Similar Items
-
LAVIB: A Large-scale Video Interpolation Benchmark
by: Stergiou, Alexandros
Published: (2024) -
MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
by: Jain, Animesh, et al.
Published: (2025) -
About Time: Advances, Challenges, and Outlooks of Action Understanding
by: Stergiou, Alexandros, et al.
Published: (2024) -
Every Shot Counts: Using Exemplars for Repetition Counting in Videos
by: Sinha, Saptarshi, et al.
Published: (2024) -
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
by: Qiu, Jason, et al.
Published: (2026)