LaSSM: Efficient Semantic-Spatial Query Decoding via Local Aggregation and State Space Models for 3D Instance Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Lei, Wang, Yi, Cui, Yawen, Liu, Moyun, Chau, Lap-Pui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
von: Yao, Lei, et al.
Veröffentlicht: (2024)
von: Yao, Lei, et al.
Veröffentlicht: (2024)
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025)
von: Yao, Lei, et al.
Veröffentlicht: (2025)
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
von: Yao, Lei, et al.
Veröffentlicht: (2026)
von: Yao, Lei, et al.
Veröffentlicht: (2026)
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
GVSynergy-Det: Synergistic Gaussian-Voxel Representations for Multi-View 3D Object Detection
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Symmetric Multi-Similarity Loss for EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2024
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2024)
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2024)
SP-MoMamba: Superpixel-driven Mixture of State Space Experts for Efficient Image Super-Resolution
von: Zou, Wenbin, et al.
Veröffentlicht: (2026)
von: Zou, Wenbin, et al.
Veröffentlicht: (2026)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2025)
CaRe-Ego: Contact-aware Relationship Modeling for Egocentric Interactive Hand-object Segmentation
von: Su, Yuejiao, et al.
Veröffentlicht: (2024)
von: Su, Yuejiao, et al.
Veröffentlicht: (2024)
Weakly-supervised Part-Attention and Mentored Networks for Vehicle Re-Identification
von: Tang, Lisha, et al.
Veröffentlicht: (2021)
von: Tang, Lisha, et al.
Veröffentlicht: (2021)
Egocentric Human-Object Interaction Detection: A New Benchmark and Method
von: Deng, Kunyuan, et al.
Veröffentlicht: (2025)
von: Deng, Kunyuan, et al.
Veröffentlicht: (2025)
HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
von: Chen, Mingjin, et al.
Veröffentlicht: (2026)
von: Chen, Mingjin, et al.
Veröffentlicht: (2026)
RS-SSM: Refining Forgotten Specifics in State Space Model for Video Semantic Segmentation
von: Zhu, Kai, et al.
Veröffentlicht: (2026)
von: Zhu, Kai, et al.
Veröffentlicht: (2026)
OccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework
von: Chen, Junliang, et al.
Veröffentlicht: (2025)
von: Chen, Junliang, et al.
Veröffentlicht: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
von: Zheng, Ying, et al.
Veröffentlicht: (2025)
ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation
von: Zheng, Ying, et al.
Veröffentlicht: (2026)
von: Zheng, Ying, et al.
Veröffentlicht: (2026)
Evolution-Inspired Sample Competition for Deep Neural Network Optimization
von: Zheng, Ying, et al.
Veröffentlicht: (2026)
von: Zheng, Ying, et al.
Veröffentlicht: (2026)
Semantic Representation Attack against Aligned Large Language Models
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
LLM-Agnostic Semantic Representation Attack
von: Lian, Jiawei, et al.
Veröffentlicht: (2026)
von: Lian, Jiawei, et al.
Veröffentlicht: (2026)
PEM: Perception Error Model for Virtual Testing of Autonomous Vehicles
von: Piazzoni, Andrea, et al.
Veröffentlicht: (2023)
von: Piazzoni, Andrea, et al.
Veröffentlicht: (2023)
A Survey on Occupancy Perception for Autonomous Driving: The Information Fusion Perspective
von: Xu, Huaiyuan, et al.
Veröffentlicht: (2024)
von: Xu, Huaiyuan, et al.
Veröffentlicht: (2024)
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction
von: Su, Yuejiao, et al.
Veröffentlicht: (2025)
von: Su, Yuejiao, et al.
Veröffentlicht: (2025)
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
von: Zheng, Ying, et al.
Veröffentlicht: (2024)
von: Zheng, Ying, et al.
Veröffentlicht: (2024)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
Prompt-Driven Lightweight Foundation Model for Instance Segmentation-Based Fault Detection in Freight Trains
von: Sun, Guodong, et al.
Veröffentlicht: (2026)
von: Sun, Guodong, et al.
Veröffentlicht: (2026)
MASS: Mesh-inellipse Aligned Deformable Surfel Splatting for Hand Reconstruction and Rendering from Egocentric Monocular Video
von: Zhu, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhu, Haoyu, et al.
Veröffentlicht: (2026)
PADetBench: Towards Benchmarking Physical Attacks against Object Detection
von: Lian, Jiawei, et al.
Veröffentlicht: (2024)
von: Lian, Jiawei, et al.
Veröffentlicht: (2024)
HSNet: Heterogeneous Subgraph Network for Single Image Super-resolution
von: Hu, Qiongyang, et al.
Veröffentlicht: (2025)
von: Hu, Qiongyang, et al.
Veröffentlicht: (2025)
Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
von: Lian, Jiawei, et al.
Veröffentlicht: (2025)
GEM: Boost Simple Network for Glass Surface Segmentation via Segment Anything Model and Data Synthesis
von: Hao, Jing, et al.
Veröffentlicht: (2024)
von: Hao, Jing, et al.
Veröffentlicht: (2024)
PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers
von: Gal, Eshed, et al.
Veröffentlicht: (2026)
von: Gal, Eshed, et al.
Veröffentlicht: (2026)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
von: Liu, Wenyang, et al.
Veröffentlicht: (2024)
Attack Anything: Blind DNNs via Universal Background Adversarial Attack
von: Lian, Jiawei, et al.
Veröffentlicht: (2024)
von: Lian, Jiawei, et al.
Veröffentlicht: (2024)
TCP-SSM: Efficient Vision State Space Models with Token-Conditioned Poles
von: Shoouri, Sara, et al.
Veröffentlicht: (2026)
von: Shoouri, Sara, et al.
Veröffentlicht: (2026)
Cortical-SSM: A Deep State Space Model for EEG and ECoG Motor Imagery Decoding
von: Suzuki, Shuntaro, et al.
Veröffentlicht: (2025)
von: Suzuki, Shuntaro, et al.
Veröffentlicht: (2025)
MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
von: Chharia, Aviral, et al.
Veröffentlicht: (2025)
von: Chharia, Aviral, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SGIFormer: Semantic-guided and Geometric-enhanced Interleaving Transformer for 3D Instance Segmentation
von: Yao, Lei, et al.
Veröffentlicht: (2024) -
GaussianCross: Cross-modal Self-supervised 3D Representation Learning via Gaussian Splatting
von: Yao, Lei, et al.
Veröffentlicht: (2025) -
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
von: Yao, Lei, et al.
Veröffentlicht: (2026) -
3DGeoDet: General-purpose Geometry-aware Image-based 3D Object Detection
von: Zhang, Yi, et al.
Veröffentlicht: (2025) -
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)