Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
Fuente:
arXiv
Saved in:
| Main Authors: | Cheshmi, Leila, Siam, Mennatullah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)
by: Karim, Rezaul, et al.
Published: (2023)
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and Benchmark
by: Broni-Bediako, Clifford, et al.
Published: (2024)
by: Broni-Bediako, Clifford, et al.
Published: (2024)
Weakly and Self-Supervised Class-Agnostic Motion Prediction for Autonomous Driving
by: Li, Ruibo, et al.
Published: (2025)
by: Li, Ruibo, et al.
Published: (2025)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
Configurable Embodied Data Generation for Class-Agnostic RGB-D Video Segmentation
by: Opipari, Anthony, et al.
Published: (2024)
by: Opipari, Anthony, et al.
Published: (2024)
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving
by: Hassani, Hossein, et al.
Published: (2025)
by: Hassani, Hossein, et al.
Published: (2025)
A Bottom-Up Approach to Class-Agnostic Image Segmentation
by: Dille, Sebastian, et al.
Published: (2024)
by: Dille, Sebastian, et al.
Published: (2024)
SAM-Enhanced Segmentation on Road Datasets: Balancing Critical Classes in Autonomous Driving
by: Tahves, Toomas, et al.
Published: (2026)
by: Tahves, Toomas, et al.
Published: (2026)
CSL: Class-Agnostic Structure-Constrained Learning for Segmentation Including the Unseen
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)
by: Kowal, Matthew, et al.
Published: (2022)
CLFT: Camera-LiDAR Fusion Transformer for Semantic Segmentation in Autonomous Driving
by: Gu, Junyi, et al.
Published: (2024)
by: Gu, Junyi, et al.
Published: (2024)
Region-Transformer: Self-Attention Region Based Class-Agnostic Point Cloud Segmentation
by: Gyawali, Dipesh, et al.
Published: (2024)
by: Gyawali, Dipesh, et al.
Published: (2024)
Enforcing View-Consistency in Class-Agnostic 3D Segmentation Fields
by: Dumery, Corentin, et al.
Published: (2024)
by: Dumery, Corentin, et al.
Published: (2024)
Class-Agnostic Visio-Temporal Scene Sketch Semantic Segmentation
by: Kütük, Aleyna, et al.
Published: (2024)
by: Kütük, Aleyna, et al.
Published: (2024)
Domain-Incremental Semantic Segmentation for Autonomous Driving under Adverse Driving Conditions
by: Muralidhara, Shishir, et al.
Published: (2025)
by: Muralidhara, Shishir, et al.
Published: (2025)
Visual Implicit Geometry Transformer for Autonomous Driving
by: Shirokov, Arsenii, et al.
Published: (2026)
by: Shirokov, Arsenii, et al.
Published: (2026)
DriveDiTFit: Fine-tuning Diffusion Transformers for Autonomous Driving
by: Tu, Jiahang, et al.
Published: (2024)
by: Tu, Jiahang, et al.
Published: (2024)
ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
by: Zhou, Shengchao, et al.
Published: (2025)
by: Zhou, Shengchao, et al.
Published: (2025)
DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
by: Hu, Xiaotao, et al.
Published: (2024)
by: Hu, Xiaotao, et al.
Published: (2024)
Panacea+: Panoramic and Controllable Video Generation for Autonomous Driving
by: Wen, Yuqing, et al.
Published: (2024)
by: Wen, Yuqing, et al.
Published: (2024)
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
by: Zheng, Mi, et al.
Published: (2025)
by: Zheng, Mi, et al.
Published: (2025)
Lightweight Temporal Transformer Decomposition for Federated Autonomous Driving
by: Do, Tuong, et al.
Published: (2025)
by: Do, Tuong, et al.
Published: (2025)
Any3DIS: Class-Agnostic 3D Instance Segmentation by 2D Mask Tracking
by: Nguyen, Phuc, et al.
Published: (2024)
by: Nguyen, Phuc, et al.
Published: (2024)
Vision Transformer Off-the-Shelf: A Surprising Baseline for Few-Shot Class-Agnostic Counting
by: Wang, Zhicheng, et al.
Published: (2023)
by: Wang, Zhicheng, et al.
Published: (2023)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
by: Xiang, Xinhao, et al.
Published: (2025)
by: Xiang, Xinhao, et al.
Published: (2025)
FAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving
by: Li, Yaoru, et al.
Published: (2026)
by: Li, Yaoru, et al.
Published: (2026)
CIT: Rethinking Class-incremental Semantic Segmentation with a Class Independent Transformation
by: Ge, Jinchao, et al.
Published: (2024)
by: Ge, Jinchao, et al.
Published: (2024)
EA-Swin: An Embedding-Agnostic Swin Transformer for AI-Generated Video Detection
by: Mai, Hung, et al.
Published: (2026)
by: Mai, Hung, et al.
Published: (2026)
DriveVGGT: Calibration-Constrained Visual Geometry Transformers for Multi-Camera Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
Unsupervised Monocular Road Segmentation for Autonomous Driving via Scene Geometry
by: Rostami, Sara Hatami, et al.
Published: (2025)
by: Rostami, Sara Hatami, et al.
Published: (2025)
ALISE: Annotation-Free LiDAR Instance Segmentation for Autonomous Driving
by: Lyu, Yongxuan, et al.
Published: (2025)
by: Lyu, Yongxuan, et al.
Published: (2025)
Calibrating the Full Predictive Class Distribution of 3D Object Detectors for Autonomous Driving
by: Schröder, Cornelius, et al.
Published: (2025)
by: Schröder, Cornelius, et al.
Published: (2025)
Class-Distribution Guided Active Learning for 3D Occupancy Prediction in Autonomous Driving
by: Kim, Wonjune, et al.
Published: (2026)
by: Kim, Wonjune, et al.
Published: (2026)
Similar Items
-
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025) -
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023) -
PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
by: Siam, Mennatullah
Published: (2025) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023) -
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)