Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Segu, Mattia, Piccinelli, Luigi, Li, Siyuan, Yang, Yung-Hsu, Schiele, Bernt, Van Gool, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
by: Yang, Yung-Hsu, et al.
Published: (2025)
by: Yang, Yung-Hsu, et al.
Published: (2025)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
Video Depth Propagation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
by: Mazzucco, Silvio, et al.
Published: (2025)
by: Mazzucco, Silvio, et al.
Published: (2025)
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Beyond SOT: Tracking Multiple Generic Objects at Once
by: Mayer, Christoph, et al.
Published: (2022)
by: Mayer, Christoph, et al.
Published: (2022)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025)
by: Broedermannn, Tim, et al.
Published: (2025)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Contrastive Learning for Multi-Object Tracking with Transformers
by: De Plaen, Pierre-François, et al.
Published: (2023)
by: De Plaen, Pierre-François, et al.
Published: (2023)
VOID: Video Object and Interaction Deletion
by: Motamed, Saman, et al.
Published: (2026)
by: Motamed, Saman, et al.
Published: (2026)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
Studying How to Efficiently and Effectively Guide Models with Explanations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Single-Model and Any-Modality for Video Object Tracking
by: Wu, Zongwei, et al.
Published: (2023)
by: Wu, Zongwei, et al.
Published: (2023)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026)
by: Sheludzko, Siarhei, et al.
Published: (2026)
Learning Generative Interactive Environments By Trained Agent Exploration
by: Kazemi, Naser, et al.
Published: (2024)
by: Kazemi, Naser, et al.
Published: (2024)
One2Any: One-Reference 6D Pose Estimation for Any Object
by: Liu, Mengya, et al.
Published: (2025)
by: Liu, Mengya, et al.
Published: (2025)
Pixel-level Certified Explanations via Randomized Smoothing
by: Anani, Alaa, et al.
Published: (2025)
by: Anani, Alaa, et al.
Published: (2025)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
by: Arya, Shreyash, et al.
Published: (2024)
by: Arya, Shreyash, et al.
Published: (2024)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
by: Rao, Sukrut, et al.
Published: (2024)
by: Rao, Sukrut, et al.
Published: (2024)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
CFM: Language-aligned Concept Foundation Model for Vision
by: Wittenmayer, Kai, et al.
Published: (2026)
by: Wittenmayer, Kai, et al.
Published: (2026)
What Matters for Scalable and Robust Learning in End-to-End Driving Planners?
by: Holtz, David, et al.
Published: (2026)
by: Holtz, David, et al.
Published: (2026)
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives
by: Fu, Yuqian, et al.
Published: (2024)
by: Fu, Yuqian, et al.
Published: (2024)
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation
by: Mahdi, Mohammad, et al.
Published: (2026)
by: Mahdi, Mohammad, et al.
Published: (2026)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
by: Parchami-Araghi, Amin, et al.
Published: (2024)
by: Parchami-Araghi, Amin, et al.
Published: (2024)
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
by: Parchami-Araghi, Amin, et al.
Published: (2025)
by: Parchami-Araghi, Amin, et al.
Published: (2025)
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
by: Pham, Nhi, et al.
Published: (2025)
by: Pham, Nhi, et al.
Published: (2025)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024)
by: Ahmed, Noor, et al.
Published: (2024)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
by: Dong, Jiangxin, et al.
Published: (2021)
by: Dong, Jiangxin, et al.
Published: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
R-CoV: Region-Aware Chain-of-Verification for Alleviating Object Hallucinations in LVLMs
by: Xie, Jiahao, et al.
Published: (2026)
by: Xie, Jiahao, et al.
Published: (2026)
Into the Fog: Evaluating Robustness of Multiple Object Tracking
by: Kirillova, Nadezda, et al.
Published: (2024)
by: Kirillova, Nadezda, et al.
Published: (2024)
Self-Explainable Affordance Learning with Embodied Caption
by: Zhang, Zhipeng, et al.
Published: (2024)
by: Zhang, Zhipeng, et al.
Published: (2024)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
Similar Items
-
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
by: Segu, Mattia, et al.
Published: (2024) -
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
by: Li, Siyuan, et al.
Published: (2024) -
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025) -
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025) -
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)