Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Nekrasov, Alexey, Athar, Ali, de Geus, Daan, Hermans, Alexander, Leibe, Bastian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How Important are Videos for Training Video LLMs?
di: Lydakis, George, et al.
Pubblicazione: (2025)
di: Lydakis, George, et al.
Pubblicazione: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
di: Knoche, Markus, et al.
Pubblicazione: (2025)
di: Knoche, Markus, et al.
Pubblicazione: (2025)
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
di: Gong, Dengxian, et al.
Pubblicazione: (2026)
di: Gong, Dengxian, et al.
Pubblicazione: (2026)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
di: Knaebel, Karim, et al.
Pubblicazione: (2025)
di: Knaebel, Karim, et al.
Pubblicazione: (2025)
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
di: Wang, Zhiyu, et al.
Pubblicazione: (2026)
di: Wang, Zhiyu, et al.
Pubblicazione: (2026)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
di: Garcia, Gonzalo Martin, et al.
Pubblicazione: (2024)
di: Garcia, Gonzalo Martin, et al.
Pubblicazione: (2024)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
di: Nekrasov, Alexey, et al.
Pubblicazione: (2024)
di: Nekrasov, Alexey, et al.
Pubblicazione: (2024)
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
di: Niu, Quanzhu, et al.
Pubblicazione: (2025)
di: Niu, Quanzhu, et al.
Pubblicazione: (2025)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
di: Yilmaz, Kadir, et al.
Pubblicazione: (2023)
di: Yilmaz, Kadir, et al.
Pubblicazione: (2023)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
Your ViT is Secretly an Image Segmentation Model
di: Kerssies, Tommie, et al.
Pubblicazione: (2025)
di: Kerssies, Tommie, et al.
Pubblicazione: (2025)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
di: Yilmaz, Kadir, et al.
Pubblicazione: (2026)
di: Yilmaz, Kadir, et al.
Pubblicazione: (2026)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
di: Knaebel, Karim, et al.
Pubblicazione: (2023)
di: Knaebel, Karim, et al.
Pubblicazione: (2023)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
di: Uslu, Jan-Lucas, et al.
Pubblicazione: (2024)
di: Uslu, Jan-Lucas, et al.
Pubblicazione: (2024)
SurGe: Improved Surface Geometry in Point Maps
di: Knaebel, Karim, et al.
Pubblicazione: (2026)
di: Knaebel, Karim, et al.
Pubblicazione: (2026)
OCCUQ: Exploring Efficient Uncertainty Quantification for 3D Occupancy Prediction
di: Heidrich, Severin, et al.
Pubblicazione: (2025)
di: Heidrich, Severin, et al.
Pubblicazione: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
di: Yuan, Haobo, et al.
Pubblicazione: (2025)
Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
di: Nekrasov, Alexey, et al.
Pubblicazione: (2025)
di: Nekrasov, Alexey, et al.
Pubblicazione: (2025)
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
di: Hong, Ran, et al.
Pubblicazione: (2025)
di: Hong, Ran, et al.
Pubblicazione: (2025)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
di: Norouzi, Narges, et al.
Pubblicazione: (2026)
di: Norouzi, Narges, et al.
Pubblicazione: (2026)
Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
di: Beemelmanns, Till, et al.
Pubblicazione: (2026)
di: Beemelmanns, Till, et al.
Pubblicazione: (2026)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
di: de Geus, Daan, et al.
Pubblicazione: (2024)
di: de Geus, Daan, et al.
Pubblicazione: (2024)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
di: Piekenbrinck, Jens, et al.
Pubblicazione: (2025)
di: Piekenbrinck, Jens, et al.
Pubblicazione: (2025)
Panoptic-CUDAL: Rural Australia Point Cloud Dataset in Rainy Conditions
di: Tseng, Tzu-Yun, et al.
Pubblicazione: (2025)
di: Tseng, Tzu-Yun, et al.
Pubblicazione: (2025)
MoSa: Motion Generation with Scalable Autoregressive Modeling
di: Liu, Mengyuan, et al.
Pubblicazione: (2025)
di: Liu, Mengyuan, et al.
Pubblicazione: (2025)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
di: Wienholt, Patrick, et al.
Pubblicazione: (2024)
di: Wienholt, Patrick, et al.
Pubblicazione: (2024)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
di: Schmidt, Christian, et al.
Pubblicazione: (2024)
di: Schmidt, Christian, et al.
Pubblicazione: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
SaENeRF: Suppressing Artifacts in Event-based Neural Radiance Fields
di: Wang, Yuanjian, et al.
Pubblicazione: (2025)
di: Wang, Yuanjian, et al.
Pubblicazione: (2025)
DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
di: Talemi, Niloufar Alipour, et al.
Pubblicazione: (2025)
di: Talemi, Niloufar Alipour, et al.
Pubblicazione: (2025)
SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
di: Liu, Hongda, et al.
Pubblicazione: (2025)
di: Liu, Hongda, et al.
Pubblicazione: (2025)
SaLF: Sparse Local Fields for Multi-Sensor Rendering in Real-Time
di: Chen, Yun, et al.
Pubblicazione: (2025)
di: Chen, Yun, et al.
Pubblicazione: (2025)
DiSa: Saliency-Aware Foreground-Background Disentangled Framework for Open-Vocabulary Semantic Segmentation
di: Yao, Zhen, et al.
Pubblicazione: (2026)
di: Yao, Zhen, et al.
Pubblicazione: (2026)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
di: Li, Zhengqi, et al.
Pubblicazione: (2024)
di: Li, Zhengqi, et al.
Pubblicazione: (2024)
Point-VOS: Pointing Up Video Object Segmentation
di: Zulfikar, Idil Esen, et al.
Pubblicazione: (2024)
di: Zulfikar, Idil Esen, et al.
Pubblicazione: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
di: Hu, Teng, et al.
Pubblicazione: (2024)
di: Hu, Teng, et al.
Pubblicazione: (2024)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
How Important are Videos for Training Video LLMs?
di: Lydakis, George, et al.
Pubblicazione: (2025) -
DONUT: A Decoder-Only Model for Trajectory Prediction
di: Knoche, Markus, et al.
Pubblicazione: (2025) -
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
di: Gong, Dengxian, et al.
Pubblicazione: (2026) -
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
di: Knaebel, Karim, et al.
Pubblicazione: (2025) -
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
di: Wang, Zhiyu, et al.
Pubblicazione: (2026)