STS-Mixer: Spatio-Temporal-Spectral Mixer for 4D Point Cloud Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Wenhao, Jiang, Xueying, Zhang, Gongjie, Zhang, Xiaoqin, Shao, Ling, Lu, Shijian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
by: Jiang, Xueying, et al.
Published: (2025)
by: Jiang, Xueying, et al.
Published: (2025)
A Survey of Label-Efficient Deep Learning for 3D Point Clouds
by: Xiao, Aoran, et al.
Published: (2023)
by: Xiao, Aoran, et al.
Published: (2023)
MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
TripleMixer: A 3D Point Cloud Denoising Model for Adverse Weather
by: Zhao, Xiongwei, et al.
Published: (2024)
by: Zhao, Xiongwei, et al.
Published: (2024)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
by: Jiang, Xueying, et al.
Published: (2026)
by: Jiang, Xueying, et al.
Published: (2026)
Modeling Continuous Motion for 3D Point Cloud Object Tracking
by: Luo, Zhipeng, et al.
Published: (2023)
by: Luo, Zhipeng, et al.
Published: (2023)
Multimodal 3D Reasoning Segmentation with Complex Scenes
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
Weakly Supervised Monocular 3D Detection with a Single-View Image
by: Jiang, Xueying, et al.
Published: (2024)
by: Jiang, Xueying, et al.
Published: (2024)
L3DR: 3D-aware LiDAR Diffusion and Rectification
by: Liu, Quan, et al.
Published: (2026)
by: Liu, Quan, et al.
Published: (2026)
ChebMixer: Efficient Graph Representation Learning with MLP Mixer
by: Kui, Xiaoyan, et al.
Published: (2024)
by: Kui, Xiaoyan, et al.
Published: (2024)
MixerCSeg: An Efficient Mixer Architecture for Crack Segmentation via Decoupled Mamba Attention
by: Zhao, Zilong, et al.
Published: (2026)
by: Zhao, Zilong, et al.
Published: (2026)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
PCR-GS: COLMAP-Free 3D Gaussian Splatting via Pose Co-Regularizations
by: Wei, Yu, et al.
Published: (2025)
by: Wei, Yu, et al.
Published: (2025)
MixerFlow: MLP-Mixer meets Normalising Flows
by: English, Eshant, et al.
Published: (2023)
by: English, Eshant, et al.
Published: (2023)
MuSASplat: Efficient Sparse-View 3D Gaussian Splats via Lightweight Multi-Scale Adaptation
by: Xu, Muyu, et al.
Published: (2025)
by: Xu, Muyu, et al.
Published: (2025)
SIAM: A Simple Alternating Mixer for Video Prediction
by: Zheng, Xin, et al.
Published: (2023)
by: Zheng, Xin, et al.
Published: (2023)
GroupedMixer: An Entropy Model with Group-wise Token-Mixers for Learned Image Compression
by: Li, Daxin, et al.
Published: (2024)
by: Li, Daxin, et al.
Published: (2024)
Hyperspectral Image Classification using Spectral-Spatial Mixer Network
by: Alkhatib, Mohammed Q.
Published: (2025)
by: Alkhatib, Mohammed Q.
Published: (2025)
ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
by: Li, Duo, et al.
Published: (2025)
by: Li, Duo, et al.
Published: (2025)
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
by: Yun, Seokju, et al.
Published: (2024)
by: Yun, Seokju, et al.
Published: (2024)
SCHEME: Scalable Channel Mixer for Vision Transformers
by: Sridhar, Deepak, et al.
Published: (2023)
by: Sridhar, Deepak, et al.
Published: (2023)
Novel View Extrapolation with Video Diffusion Priors
by: Liu, Kunhao, et al.
Published: (2024)
by: Liu, Kunhao, et al.
Published: (2024)
ATOM: Attention Mixer for Efficient Dataset Distillation
by: Khaki, Samir, et al.
Published: (2024)
by: Khaki, Samir, et al.
Published: (2024)
QN-Mixer: A Quasi-Newton MLP-Mixer Model for Sparse-View CT Reconstruction
by: Ayad, Ishak, et al.
Published: (2024)
by: Ayad, Ishak, et al.
Published: (2024)
Historical Test-time Prompt Tuning for Vision Foundation Models
by: Zhang, Jingyi, et al.
Published: (2024)
by: Zhang, Jingyi, et al.
Published: (2024)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Modality-Aware and Shift Mixer for Multi-modal Brain Tumor Segmentation
by: Huang, Zhongzhen, et al.
Published: (2024)
by: Huang, Zhongzhen, et al.
Published: (2024)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
by: Xu, Wenhao, et al.
Published: (2025)
by: Xu, Wenhao, et al.
Published: (2025)
PoM: Efficient Image and Video Generation with the Polynomial Mixer
by: Picard, David, et al.
Published: (2024)
by: Picard, David, et al.
Published: (2024)
On the Generalization Capacities of MLLMs for Spatial Intelligence
by: Zhang, Gongjie, et al.
Published: (2026)
by: Zhang, Gongjie, et al.
Published: (2026)
MMRel: Benchmarking Relation Understanding in Multi-Modal Large Language Models
by: Nie, Jiahao, et al.
Published: (2024)
by: Nie, Jiahao, et al.
Published: (2024)
DeGMix: Efficient Multi-Task Dense Prediction with Deformable and Gating Mixer
by: Xu, Yangyang, et al.
Published: (2023)
by: Xu, Yangyang, et al.
Published: (2023)
E.M.Ground: A Temporal Grounding Vid-LLM with Holistic Event Perception and Matching
by: Nie, Jiahao, et al.
Published: (2026)
by: Nie, Jiahao, et al.
Published: (2026)
MixerMDM: Learnable Composition of Human Motion Diffusion Models
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
by: Ruiz-Ponce, Pablo, et al.
Published: (2025)
HyPCV-Former: Hyperbolic Spatio-Temporal Transformer for 3D Point Cloud Video Anomaly Detection
by: Cao, Jiaping, et al.
Published: (2025)
by: Cao, Jiaping, et al.
Published: (2025)
MAMBA4D: Efficient Long-Sequence Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
by: Liu, Jiuming, et al.
Published: (2024)
by: Liu, Jiuming, et al.
Published: (2024)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
MixerSENet: A Lightweight Framework for Efficient Hyperspectral Image Classification
by: Alkhatib, Mohammed Q., et al.
Published: (2026)
by: Alkhatib, Mohammed Q., et al.
Published: (2026)
iMixer: hierarchical Hopfield network implies an invertible, implicit and iterative MLP-Mixer
by: Ota, Toshihiro, et al.
Published: (2023)
by: Ota, Toshihiro, et al.
Published: (2023)
Similar Items
-
Exploring 3D Reasoning-Driven Planning: From Implicit Human Intentions to Route-Aware Activity Planning
by: Jiang, Xueying, et al.
Published: (2025) -
A Survey of Label-Efficient Deep Learning for 3D Point Clouds
by: Xiao, Aoran, et al.
Published: (2023) -
MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked Autoencoders
by: Jiang, Xueying, et al.
Published: (2024) -
TripleMixer: A 3D Point Cloud Denoising Model for Adverse Weather
by: Zhao, Xiongwei, et al.
Published: (2024) -
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
by: Jiang, Xueying, et al.
Published: (2026)