Boosting Visual Instruction Tuning with Self-Supervised Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sirko-Galouchenko, Sophia, Wysoczanska, Monika, Bursuc, Andrei, Thome, Nicolas, Gidaris, Spyros |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2025)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2025)
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2024)
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
von: Simoncini, Walter, et al.
Veröffentlicht: (2024)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2023)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2023)
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
von: Vobecky, Antonin, et al.
Veröffentlicht: (2022)
von: Vobecky, Antonin, et al.
Veröffentlicht: (2022)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
von: Vobecky, Antonin, et al.
Veröffentlicht: (2024)
von: Vobecky, Antonin, et al.
Veröffentlicht: (2024)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2024)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2024)
MOCA: Self-supervised Representation Learning by Predicting Masked Online Codebook Assignments
von: Gidaris, Spyros, et al.
Veröffentlicht: (2023)
von: Gidaris, Spyros, et al.
Veröffentlicht: (2023)
Unsupervised Object Localization in the Era of Self-Supervised ViTs: A Survey
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
von: Siméoni, Oriane, et al.
Veröffentlicht: (2023)
Three Pillars improving Vision Foundation Model Distillation for Lidar
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2025)
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2025)
Boosting Generative Image Modeling via Joint Image-Feature Synthesis
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2025)
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2025)
SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
von: Kakogeorgiou, Ioannis, et al.
Veröffentlicht: (2023)
von: Kakogeorgiou, Ioannis, et al.
Veröffentlicht: (2023)
Coevolving Representations in Joint Image-Feature Diffusion
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2026)
von: Kouzelis, Theodoros, et al.
Veröffentlicht: (2026)
Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2026)
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2026)
DINO-Foresight: Looking into the Future with DINO
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
Attention, May I Have Your Decision? Localizing Generative Choices in Diffusion Models
von: Zaleska, Katarzyna, et al.
Veröffentlicht: (2026)
von: Zaleska, Katarzyna, et al.
Veröffentlicht: (2026)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2025)
Multi-Token Prediction Needs Registers
von: Gerontopoulos, Anastasios, et al.
Veröffentlicht: (2025)
von: Gerontopoulos, Anastasios, et al.
Veröffentlicht: (2025)
SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering
von: Hayes, Seamie, et al.
Veröffentlicht: (2025)
von: Hayes, Seamie, et al.
Veröffentlicht: (2025)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
von: Fahes, Mohammad, et al.
Veröffentlicht: (2024)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
Valeo4Cast: A Modular Approach to End-to-End Forecasting
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
von: Xu, Yihong, et al.
Veröffentlicht: (2024)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
BIGFix: Bidirectional Image Generation with Token Fixing
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
von: Besnier, Victor, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Generative Visual Instruction Tuning
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Self-Guidance: Boosting Flow and Diffusion Generation on Their Own
von: Li, Tiancheng, et al.
Veröffentlicht: (2024)
von: Li, Tiancheng, et al.
Veröffentlicht: (2024)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
Visual Instruction Tuning with Chain of Region-of-Interest
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
Driving on Registers
von: Kirby, Ellington, et al.
Veröffentlicht: (2026)
von: Kirby, Ellington, et al.
Veröffentlicht: (2026)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
von: Zohar, Orr, et al.
Veröffentlicht: (2024)
LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
von: de Moreau, Simon, et al.
Veröffentlicht: (2025)
von: de Moreau, Simon, et al.
Veröffentlicht: (2025)
A Simple Recipe for Language-guided Domain Generalized Segmentation
von: Fahes, Mohammad, et al.
Veröffentlicht: (2023)
von: Fahes, Mohammad, et al.
Veröffentlicht: (2023)
RealCamo: Boosting Real Camouflage Synthesis with Layout Controls and Textual-Visual Guidance
von: Chen, Chunyuan, et al.
Veröffentlicht: (2025)
von: Chen, Chunyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2025) -
OccFeat: Self-supervised Occupancy Feature Prediction for Pretraining BEV Segmentation Networks
von: Sirko-Galouchenko, Sophia, et al.
Veröffentlicht: (2024) -
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
von: Simoncini, Walter, et al.
Veröffentlicht: (2024) -
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
von: Wysoczańska, Monika, et al.
Veröffentlicht: (2023) -
Franca: Nested Matryoshka Clustering for Scalable Visual Representation Learning
von: Venkataramanan, Shashanka, et al.
Veröffentlicht: (2025)