Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Nekrasov, Alexey, Athar, Ali, de Geus, Daan, Hermans, Alexander, Leibe, Bastian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Important are Videos for Training Video LLMs?
por: Lydakis, George, et al.
Publicado: (2025)
por: Lydakis, George, et al.
Publicado: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
por: Knoche, Markus, et al.
Publicado: (2025)
por: Knoche, Markus, et al.
Publicado: (2025)
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
por: Gong, Dengxian, et al.
Publicado: (2026)
por: Gong, Dengxian, et al.
Publicado: (2026)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
por: Knaebel, Karim, et al.
Publicado: (2025)
por: Knaebel, Karim, et al.
Publicado: (2025)
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
por: Wang, Zhiyu, et al.
Publicado: (2026)
por: Wang, Zhiyu, et al.
Publicado: (2026)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
por: Garcia, Gonzalo Martin, et al.
Publicado: (2024)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
por: Nekrasov, Alexey, et al.
Publicado: (2024)
por: Nekrasov, Alexey, et al.
Publicado: (2024)
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
por: Niu, Quanzhu, et al.
Publicado: (2025)
por: Niu, Quanzhu, et al.
Publicado: (2025)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
por: Yilmaz, Kadir, et al.
Publicado: (2023)
por: Yilmaz, Kadir, et al.
Publicado: (2023)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
por: Yuan, Haobo, et al.
Publicado: (2025)
por: Yuan, Haobo, et al.
Publicado: (2025)
Your ViT is Secretly an Image Segmentation Model
por: Kerssies, Tommie, et al.
Publicado: (2025)
por: Kerssies, Tommie, et al.
Publicado: (2025)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
por: Yilmaz, Kadir, et al.
Publicado: (2026)
por: Yilmaz, Kadir, et al.
Publicado: (2026)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
por: Knaebel, Karim, et al.
Publicado: (2023)
por: Knaebel, Karim, et al.
Publicado: (2023)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
por: Uslu, Jan-Lucas, et al.
Publicado: (2024)
por: Uslu, Jan-Lucas, et al.
Publicado: (2024)
SurGe: Improved Surface Geometry in Point Maps
por: Knaebel, Karim, et al.
Publicado: (2026)
por: Knaebel, Karim, et al.
Publicado: (2026)
OCCUQ: Exploring Efficient Uncertainty Quantification for 3D Occupancy Prediction
por: Heidrich, Severin, et al.
Publicado: (2025)
por: Heidrich, Severin, et al.
Publicado: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
por: Yuan, Haobo, et al.
Publicado: (2025)
por: Yuan, Haobo, et al.
Publicado: (2025)
Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
por: Nekrasov, Alexey, et al.
Publicado: (2025)
por: Nekrasov, Alexey, et al.
Publicado: (2025)
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
por: Hong, Ran, et al.
Publicado: (2025)
por: Hong, Ran, et al.
Publicado: (2025)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
por: Norouzi, Narges, et al.
Publicado: (2026)
por: Norouzi, Narges, et al.
Publicado: (2026)
Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
por: Beemelmanns, Till, et al.
Publicado: (2026)
por: Beemelmanns, Till, et al.
Publicado: (2026)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
por: de Geus, Daan, et al.
Publicado: (2024)
por: de Geus, Daan, et al.
Publicado: (2024)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
por: Piekenbrinck, Jens, et al.
Publicado: (2025)
por: Piekenbrinck, Jens, et al.
Publicado: (2025)
Panoptic-CUDAL: Rural Australia Point Cloud Dataset in Rainy Conditions
por: Tseng, Tzu-Yun, et al.
Publicado: (2025)
por: Tseng, Tzu-Yun, et al.
Publicado: (2025)
MoSa: Motion Generation with Scalable Autoregressive Modeling
por: Liu, Mengyuan, et al.
Publicado: (2025)
por: Liu, Mengyuan, et al.
Publicado: (2025)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
por: Wienholt, Patrick, et al.
Publicado: (2024)
por: Wienholt, Patrick, et al.
Publicado: (2024)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
por: Schmidt, Christian, et al.
Publicado: (2024)
por: Schmidt, Christian, et al.
Publicado: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
por: Norouzi, Narges, et al.
Publicado: (2024)
por: Norouzi, Narges, et al.
Publicado: (2024)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
por: Cavagnero, Niccolò, et al.
Publicado: (2026)
por: Cavagnero, Niccolò, et al.
Publicado: (2026)
SaENeRF: Suppressing Artifacts in Event-based Neural Radiance Fields
por: Wang, Yuanjian, et al.
Publicado: (2025)
por: Wang, Yuanjian, et al.
Publicado: (2025)
DiSa: Directional Saliency-Aware Prompt Learning for Generalizable Vision-Language Models
por: Talemi, Niloufar Alipour, et al.
Publicado: (2025)
por: Talemi, Niloufar Alipour, et al.
Publicado: (2025)
SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
por: Liu, Hongda, et al.
Publicado: (2025)
por: Liu, Hongda, et al.
Publicado: (2025)
SaLF: Sparse Local Fields for Multi-Sensor Rendering in Real-Time
por: Chen, Yun, et al.
Publicado: (2025)
por: Chen, Yun, et al.
Publicado: (2025)
DiSa: Saliency-Aware Foreground-Background Disentangled Framework for Open-Vocabulary Semantic Segmentation
por: Yao, Zhen, et al.
Publicado: (2026)
por: Yao, Zhen, et al.
Publicado: (2026)
MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos
por: Li, Zhengqi, et al.
Publicado: (2024)
por: Li, Zhengqi, et al.
Publicado: (2024)
Point-VOS: Pointing Up Video Object Segmentation
por: Zulfikar, Idil Esen, et al.
Publicado: (2024)
por: Zulfikar, Idil Esen, et al.
Publicado: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
por: Kerssies, Tommie, et al.
Publicado: (2024)
por: Kerssies, Tommie, et al.
Publicado: (2024)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
por: Hu, Teng, et al.
Publicado: (2024)
por: Hu, Teng, et al.
Publicado: (2024)
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering
por: Yang, Tianyu, et al.
Publicado: (2024)
por: Yang, Tianyu, et al.
Publicado: (2024)
Ejemplares similares
-
How Important are Videos for Training Video LLMs?
por: Lydakis, George, et al.
Publicado: (2025) -
DONUT: A Decoder-Only Model for Trajectory Prediction
por: Knoche, Markus, et al.
Publicado: (2025) -
SaSaSaSa2VA: 2nd Place of the 5th PVUW MeViS-Text Track
por: Gong, Dengxian, et al.
Publicado: (2026) -
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
por: Knaebel, Karim, et al.
Publicado: (2025) -
2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA
por: Wang, Zhiyu, et al.
Publicado: (2026)