SSL-R1: Self-Supervised Visual Reinforcement Post-Training for Multimodal Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Xie, Jiahao, Tonioni, Alessio, Rauschmayr, Nathalie, Tombari, Federico, Schiele, Bernt |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
di: Xie, Jiahao, et al.
Pubblicazione: (2026)
Test-Time Visual In-Context Tuning
di: Xie, Jiahao, et al.
Pubblicazione: (2025)
di: Xie, Jiahao, et al.
Pubblicazione: (2025)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
di: Kuzucu, Selim, et al.
Pubblicazione: (2025)
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
di: Kuzucu, Selim, et al.
Pubblicazione: (2026)
di: Kuzucu, Selim, et al.
Pubblicazione: (2026)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
di: Kukleva, Anna, et al.
Pubblicazione: (2025)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
di: Nath, Mriganka, et al.
Pubblicazione: (2026)
di: Nath, Mriganka, et al.
Pubblicazione: (2026)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
di: Simsar, Enis, et al.
Pubblicazione: (2023)
di: Simsar, Enis, et al.
Pubblicazione: (2023)
Extracting Training Data from Document-Based VQA Models
di: Pinto, Francesco, et al.
Pubblicazione: (2024)
di: Pinto, Francesco, et al.
Pubblicazione: (2024)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
di: Simsar, Enis, et al.
Pubblicazione: (2024)
di: Simsar, Enis, et al.
Pubblicazione: (2024)
Text-Conditioned Resampler For Long Form Video Understanding
di: Korbar, Bruno, et al.
Pubblicazione: (2023)
di: Korbar, Bruno, et al.
Pubblicazione: (2023)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
di: Plizzari, Chiara, et al.
Pubblicazione: (2025)
di: Plizzari, Chiara, et al.
Pubblicazione: (2025)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
di: Alshami, Eyad, et al.
Pubblicazione: (2025)
Reading Between the Lines: Abstaining from VLM-Generated OCR Errors via Latent Representation Probes
di: Yao, Jihan, et al.
Pubblicazione: (2025)
di: Yao, Jihan, et al.
Pubblicazione: (2025)
Active Data Curation Effectively Distills Large-Scale Multimodal Models
di: Udandarao, Vishaal, et al.
Pubblicazione: (2024)
di: Udandarao, Vishaal, et al.
Pubblicazione: (2024)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
di: Gairola, Siddhartha, et al.
Pubblicazione: (2025)
di: Gairola, Siddhartha, et al.
Pubblicazione: (2025)
SSL-SLR: Self-Supervised Representation Learning for Sign Language Recognition
di: Madjoukeng, Ariel Basso, et al.
Pubblicazione: (2025)
di: Madjoukeng, Ariel Basso, et al.
Pubblicazione: (2025)
Adversarial Training against Location-Optimized Adversarial Patches
di: Rao, Sukrut, et al.
Pubblicazione: (2020)
di: Rao, Sukrut, et al.
Pubblicazione: (2020)
BRAVE: Broadening the visual encoding of vision-language models
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2024)
di: Kar, Oğuzhan Fatih, et al.
Pubblicazione: (2024)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
di: Das, Anurag, et al.
Pubblicazione: (2024)
di: Das, Anurag, et al.
Pubblicazione: (2024)
Scribbles for All: Benchmarking Scribble Supervised Segmentation Across Datasets
di: Boettcher, Wolfgang, et al.
Pubblicazione: (2024)
di: Boettcher, Wolfgang, et al.
Pubblicazione: (2024)
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
di: Das, Anurag, et al.
Pubblicazione: (2026)
di: Das, Anurag, et al.
Pubblicazione: (2026)
SimNP: Learning Self-Similarity Priors Between Neural Points
di: Wewer, Christopher, et al.
Pubblicazione: (2023)
di: Wewer, Christopher, et al.
Pubblicazione: (2023)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
di: Ahmed, Noor, et al.
Pubblicazione: (2024)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
di: Dong, Jiangxin, et al.
Pubblicazione: (2021)
di: Dong, Jiangxin, et al.
Pubblicazione: (2021)
Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation
di: Ernst, Jonas, et al.
Pubblicazione: (2026)
di: Ernst, Jonas, et al.
Pubblicazione: (2026)
AnyUp: Universal Feature Upsampling
di: Wimmer, Thomas, et al.
Pubblicazione: (2025)
di: Wimmer, Thomas, et al.
Pubblicazione: (2025)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
di: Shahbazi, Mohamad, et al.
Pubblicazione: (2024)
di: Shahbazi, Mohamad, et al.
Pubblicazione: (2024)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
di: Shi, Shaoshuai, et al.
Pubblicazione: (2023)
di: Shi, Shaoshuai, et al.
Pubblicazione: (2023)
Training-free Online Video Step Grounding
di: Zanella, Luca, et al.
Pubblicazione: (2025)
di: Zanella, Luca, et al.
Pubblicazione: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
di: Khattak, Muhammad Uzair, et al.
Pubblicazione: (2024)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
di: Böhle, Moritz, et al.
Pubblicazione: (2021)
di: Böhle, Moritz, et al.
Pubblicazione: (2021)
Towards Better Understanding Attribution Methods
di: Rao, Sukrut, et al.
Pubblicazione: (2022)
di: Rao, Sukrut, et al.
Pubblicazione: (2022)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
di: Böhle, Moritz, et al.
Pubblicazione: (2023)
FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation
di: Bill, Eric Tillmann, et al.
Pubblicazione: (2026)
di: Bill, Eric Tillmann, et al.
Pubblicazione: (2026)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
di: Gorgun, Ada, et al.
Pubblicazione: (2025)
PP-SSL : Priority-Perception Self-Supervised Learning for Fine-Grained Recognition
di: Li, ShuaiHeng, et al.
Pubblicazione: (2024)
di: Li, ShuaiHeng, et al.
Pubblicazione: (2024)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
di: Wang, Hanqing, et al.
Pubblicazione: (2025)
di: Wang, Hanqing, et al.
Pubblicazione: (2025)
Documenti analoghi
-
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
di: Xie, Jiahao, et al.
Pubblicazione: (2026) -
Test-Time Visual In-Context Tuning
di: Xie, Jiahao, et al.
Pubblicazione: (2025) -
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
di: Kuzucu, Selim, et al.
Pubblicazione: (2025) -
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
di: Kuzucu, Selim, et al.
Pubblicazione: (2026) -
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
di: Kukleva, Anna, et al.
Pubblicazione: (2025)