SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Qi, Liu, Xinchen, Liu, Kun, Gu, Xiaoyan, Liu, Wu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026)
by: He, Allen, et al.
Published: (2026)
Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer
by: Liu, Ziyi, et al.
Published: (2025)
by: Liu, Ziyi, et al.
Published: (2025)
SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis
by: Feng, Bin, et al.
Published: (2025)
by: Feng, Bin, et al.
Published: (2025)
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025)
by: Zahid, Azizul, et al.
Published: (2025)
Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
by: Liu, Junlin, et al.
Published: (2024)
by: Liu, Junlin, et al.
Published: (2024)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
by: Guo, Jianing, et al.
Published: (2025)
by: Guo, Jianing, et al.
Published: (2025)
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
by: Dong, Zhe, et al.
Published: (2024)
by: Dong, Zhe, et al.
Published: (2024)
AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless Workflows
by: Zhang, RuiQiang, et al.
Published: (2025)
by: Zhang, RuiQiang, et al.
Published: (2025)
VGTS: Visually Guided Text Spotting for Novel Categories in Historical Manuscripts
by: Hu, Wenbo, et al.
Published: (2023)
by: Hu, Wenbo, et al.
Published: (2023)
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
by: Liu, Peng
Published: (2021)
by: Liu, Peng
Published: (2021)
SPACT18: Spiking Human Action Recognition Benchmark Dataset with Complementary RGB and Thermal Modalities
by: Ashraf, Yasser, et al.
Published: (2025)
by: Ashraf, Yasser, et al.
Published: (2025)
TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
by: Qi, Qiang, et al.
Published: (2025)
by: Qi, Qiang, et al.
Published: (2025)
Ideal Registration? Segmentation is All You Need
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
WaveFormer: Frequency-Time Decoupled Vision Modeling with Wave Equation
by: Shu, Zishan, et al.
Published: (2026)
by: Shu, Zishan, et al.
Published: (2026)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
ContourFormer: Real-Time Contour-Based End-to-End Instance Segmentation Transformer
by: Yao, Weiwei, et al.
Published: (2025)
by: Yao, Weiwei, et al.
Published: (2025)
SegMaFormer: A Hybrid State-Space and Transformer Model for Efficient Segmentation
by: Nguyen, Duy D., et al.
Published: (2026)
by: Nguyen, Duy D., et al.
Published: (2026)
Graph Relation Distillation for Efficient Biomedical Instance Segmentation
by: Liu, Xiaoyu, et al.
Published: (2024)
by: Liu, Xiaoyu, et al.
Published: (2024)
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
by: Liao, Zhichao, et al.
Published: (2025)
by: Liao, Zhichao, et al.
Published: (2025)
S3T-Former: A Purely Spike-Driven State-Space Topology Transformer for Skeleton Action Recognition
by: Zheng, Naichuan, et al.
Published: (2026)
by: Zheng, Naichuan, et al.
Published: (2026)
Multi-Scale Correlation-Aware Transformer for Maritime Vessel Re-Identification
by: Liu, Yunhe
Published: (2025)
by: Liu, Yunhe
Published: (2025)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
PG-SAM: Prior-Guided SAM with Medical for Multi-organ Segmentation
by: Zhong, Yiheng, et al.
Published: (2025)
by: Zhong, Yiheng, et al.
Published: (2025)
MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for Multi-Modal Video Temporal Grounding
by: Zhu, Zhiyi, et al.
Published: (2025)
by: Zhu, Zhiyi, et al.
Published: (2025)
AMBER -- Advanced SegFormer for Multi-Band Image Segmentation: an application to Hyperspectral Imaging
by: Dosi, Andrea, et al.
Published: (2024)
by: Dosi, Andrea, et al.
Published: (2024)
Hydra: Accurate Multi-Modal Leaf Wetness Sensing with mm-Wave and Camera Fusion
by: Liu, Yimeng, et al.
Published: (2025)
by: Liu, Yimeng, et al.
Published: (2025)
HUR-MACL: High-Uncertainty Region-Guided Multi-Architecture Collaborative Learning for Head and Neck Multi-Organ Segmentation
by: Liu, Xiaoyu, et al.
Published: (2026)
by: Liu, Xiaoyu, et al.
Published: (2026)
All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment
by: Zhang, Chunhui, et al.
Published: (2023)
by: Zhang, Chunhui, et al.
Published: (2023)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
by: Hong, Yuyang, et al.
Published: (2025)
by: Hong, Yuyang, et al.
Published: (2025)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
by: Luo, Jie, et al.
Published: (2025)
by: Luo, Jie, et al.
Published: (2025)
TCFormer: A 5M-Parameter Transformer with Density-Guided Aggregation for Weakly-Supervised Crowd Counting
by: Guo, Qiang, et al.
Published: (2025)
by: Guo, Qiang, et al.
Published: (2025)
PulseMind: A Multi-Modal Medical Model for Real-World Clinical Diagnosis
by: Xu, Jiao, et al.
Published: (2026)
by: Xu, Jiao, et al.
Published: (2026)
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
by: Chen, Pengzhen, et al.
Published: (2026)
by: Chen, Pengzhen, et al.
Published: (2026)
FluenceFormer: Transformer-Driven Multi-Beam Fluence Map Regression for Radiotherapy Planning
by: Mgboh, Ujunwa, et al.
Published: (2025)
by: Mgboh, Ujunwa, et al.
Published: (2025)
WaveFormer: A 3D Transformer with Wavelet-Driven Feature Representation for Efficient Medical Image Segmentation
by: Hasan, Md Mahfuz Al, et al.
Published: (2025)
by: Hasan, Md Mahfuz Al, et al.
Published: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
CRFT: Consistent-Recurrent Feature Flow Transformer for Cross-Modal Image Registration
by: Liu, Xuecong, et al.
Published: (2026)
by: Liu, Xuecong, et al.
Published: (2026)
Similar Items
-
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026) -
Bridge the Gap: From Weak to Full Supervision for Temporal Action Localization with PseudoFormer
by: Liu, Ziyi, et al.
Published: (2025) -
SentiFormer: Metadata Enhanced Transformer for Image Sentiment Analysis
by: Feng, Bin, et al.
Published: (2025) -
Toward Aligning Human and Robot Actions via Multi-Modal Demonstration Learning
by: Zahid, Azizul, et al.
Published: (2025) -
Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
by: Liu, Junlin, et al.
Published: (2024)