Revisiting Multi-Task Visual Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Di, Shangzhe, Zhai, Zhonghua, Xie, Weidi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025)
by: Li, Zeqian, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024)
by: Shi, Yudi, et al.
Published: (2024)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
Learning Streaming Video Representation via Multitask Training
by: Yan, Yibin, et al.
Published: (2025)
by: Yan, Yibin, et al.
Published: (2025)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
by: Yan, Yibin, et al.
Published: (2024)
by: Yan, Yibin, et al.
Published: (2024)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
Language-Informed Visual Concept Learning
by: Lee, Sharon, et al.
Published: (2023)
by: Lee, Sharon, et al.
Published: (2023)
Can Visual Foundation Models Achieve Long-term Point Tracking?
by: Aydemir, Görkay, et al.
Published: (2024)
by: Aydemir, Görkay, et al.
Published: (2024)
Appearance-Based Refinement for Object-Centric Motion Segmentation
by: Xie, Junyu, et al.
Published: (2023)
by: Xie, Junyu, et al.
Published: (2023)
ELIP: Enhanced Visual-Language Foundation Models for Image Retrieval
by: Zhan, Guanqi, et al.
Published: (2025)
by: Zhan, Guanqi, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Composable Visual Tokenizers with Generator-Free Diagnostics of Learnability
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
by: Aydemir, Görkay, et al.
Published: (2026)
by: Aydemir, Görkay, et al.
Published: (2026)
Count Anything at Any Granularity
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
Track-On2: Enhancing Online Point Tracking with Memory
by: Aydemir, Görkay, et al.
Published: (2025)
by: Aydemir, Görkay, et al.
Published: (2025)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
by: Liu, Chang, et al.
Published: (2023)
by: Liu, Chang, et al.
Published: (2023)
ImLoc: Revisiting Visual Localization with Image-based Representation
by: Jiang, Xudong, et al.
Published: (2026)
by: Jiang, Xudong, et al.
Published: (2026)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
Moving Object Segmentation: All You Need Is SAM (and Flow)
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
by: Sun, Luoyi, et al.
Published: (2023)
by: Sun, Luoyi, et al.
Published: (2023)
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
by: Ma, Yeyao, et al.
Published: (2026)
by: Ma, Yeyao, et al.
Published: (2026)
Multi-Agent System for Comprehensive Soccer Understanding
by: Rao, Jiayuan, et al.
Published: (2025)
by: Rao, Jiayuan, et al.
Published: (2025)
Multi-label Cluster Discrimination for Visual Representation Learning
by: An, Xiang, et al.
Published: (2024)
by: An, Xiang, et al.
Published: (2024)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
by: Wu, Size, et al.
Published: (2025)
by: Wu, Size, et al.
Published: (2025)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
Task-Guided Multi-Annotation Triplet Learning for Remote Sensing Representations
by: Zhou, Meilun, et al.
Published: (2026)
by: Zhou, Meilun, et al.
Published: (2026)
Decoding Visual Neural Representations by Multimodal with Dynamic Balancing
by: sun, Kaili, et al.
Published: (2025)
by: sun, Kaili, et al.
Published: (2025)
Cell Variational Information Bottleneck Network
by: Zhai, Zhonghua, et al.
Published: (2024)
by: Zhai, Zhonghua, et al.
Published: (2024)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
A Sanity Check on Composed Image Retrieval
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
by: Ge, Mingji, et al.
Published: (2026)
by: Ge, Mingji, et al.
Published: (2026)
Amodal Ground Truth and Completion in the Wild
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Track-On: Transformer-based Online Point Tracking with Memory
by: Aydemir, Görkay, et al.
Published: (2025)
by: Aydemir, Görkay, et al.
Published: (2025)
A General Protocol to Probe Large Vision Models for 3D Physical Understanding
by: Zhan, Guanqi, et al.
Published: (2023)
by: Zhan, Guanqi, et al.
Published: (2023)
Scaling Language-Free Visual Representation Learning
by: Fan, David, et al.
Published: (2025)
by: Fan, David, et al.
Published: (2025)
Similar Items
-
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
by: Li, Zeqian, et al.
Published: (2025) -
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023) -
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024) -
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
by: Shi, Yudi, et al.
Published: (2024) -
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)