Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Hao, Wang, Hu, Zheng, Tiantian, Bhattarai, Prajjwal, Alhanai, Tuka |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
von: Eghbali, Niloufar, et al.
Veröffentlicht: (2025)
von: Eghbali, Niloufar, et al.
Veröffentlicht: (2025)
MoENAS: Mixture-of-Expert based Neural Architecture Search for jointly Accurate, Fair, and Robust Edge Deep Neural Networks
von: Mecharbat, Lotfi Abdelkrim, et al.
Veröffentlicht: (2025)
von: Mecharbat, Lotfi Abdelkrim, et al.
Veröffentlicht: (2025)
Feature Imitating Networks Enhance The Performance, Reliability And Speed Of Deep Learning On Biomedical Image Processing Tasks
von: Min, Shangyang, et al.
Veröffentlicht: (2023)
von: Min, Shangyang, et al.
Veröffentlicht: (2023)
Vision Transformers: From Semantic Segmentation to Dense Prediction
von: Zhang, Li, et al.
Veröffentlicht: (2022)
von: Zhang, Li, et al.
Veröffentlicht: (2022)
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
von: He, Yong, et al.
Veröffentlicht: (2025)
von: He, Yong, et al.
Veröffentlicht: (2025)
HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
von: Lu, Wenquan, et al.
Veröffentlicht: (2023)
Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation
von: Ortega, Joel Valdivia, et al.
Veröffentlicht: (2026)
von: Ortega, Joel Valdivia, et al.
Veröffentlicht: (2026)
Knowledge distillation through geometry-aware representational alignment
von: Bhattarai, Prajjwal, et al.
Veröffentlicht: (2025)
von: Bhattarai, Prajjwal, et al.
Veröffentlicht: (2025)
Generative Hierarchical Temporal Transformer for Hand Pose and Action Modeling
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
AdaptDiff: Cross-Modality Domain Adaptation via Weak Conditional Semantic Diffusion for Retinal Vessel Segmentation
von: Hu, Dewei, et al.
Veröffentlicht: (2024)
von: Hu, Dewei, et al.
Veröffentlicht: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
von: Fu, Bowen, et al.
Veröffentlicht: (2023)
von: Fu, Bowen, et al.
Veröffentlicht: (2023)
Diffusion-based RGB-D Semantic Segmentation with Deformable Attention Transformer
von: Bui, Minh, et al.
Veröffentlicht: (2024)
von: Bui, Minh, et al.
Veröffentlicht: (2024)
Ego-centric Predictive Model Conditioned on Hand Trajectories
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
von: Zhang, Binjie, et al.
Veröffentlicht: (2025)
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
von: Dong, Bowen, et al.
Veröffentlicht: (2024)
D3S2: Diffusion-Guided Dataset Distillation for Semantic Segmentation
von: Zheng, Wenjie, et al.
Veröffentlicht: (2026)
von: Zheng, Wenjie, et al.
Veröffentlicht: (2026)
A Gift from the Integration of Discriminative and Diffusion-based Generative Learning: Boundary Refinement Remote Sensing Semantic Segmentation
von: Wang, Hao, et al.
Veröffentlicht: (2025)
von: Wang, Hao, et al.
Veröffentlicht: (2025)
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer
von: Hu, Jinyi, et al.
Veröffentlicht: (2024)
von: Hu, Jinyi, et al.
Veröffentlicht: (2024)
Faster Diffusion Action Segmentation
von: Wang, Shuaibing, et al.
Veröffentlicht: (2024)
von: Wang, Shuaibing, et al.
Veröffentlicht: (2024)
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
von: Sayem, MD Khalequzzaman Chowdhury, et al.
Veröffentlicht: (2026)
von: Sayem, MD Khalequzzaman Chowdhury, et al.
Veröffentlicht: (2026)
SDiT: Semantic Region-Adaptive for Diffusion Transformers
von: Lin, Bowen, et al.
Veröffentlicht: (2026)
von: Lin, Bowen, et al.
Veröffentlicht: (2026)
PASS:Test-Time Prompting to Adapt Styles and Semantic Shapes in Medical Image Segmentation
von: Zhang, Chuyan, et al.
Veröffentlicht: (2024)
von: Zhang, Chuyan, et al.
Veröffentlicht: (2024)
EHPE: A Segmented Architecture for Enhanced Hand Pose Estimation
von: Zheng, Bolun, et al.
Veröffentlicht: (2025)
von: Zheng, Bolun, et al.
Veröffentlicht: (2025)
A Hidden Semantic Bottleneck in Conditional Embeddings of Diffusion Transformers
von: Pham, Trung X., et al.
Veröffentlicht: (2026)
von: Pham, Trung X., et al.
Veröffentlicht: (2026)
Heuristic Self-Paced Learning for Domain Adaptive Semantic Segmentation under Adverse Conditions
von: Wang, Shiqin, et al.
Veröffentlicht: (2026)
von: Wang, Shiqin, et al.
Veröffentlicht: (2026)
MemorySAM: Memorize Modalities and Semantics with Segment Anything Model 2 for Multi-modal Semantic Segmentation
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
von: Liao, Chenfei, et al.
Veröffentlicht: (2025)
Video Virtual Try-on with Conditional Diffusion Transformer Inpainter
von: Zou, Cheng, et al.
Veröffentlicht: (2025)
von: Zou, Cheng, et al.
Veröffentlicht: (2025)
HandBooster: Boosting 3D Hand-Mesh Reconstruction by Conditional Synthesis and Sampling of Hand-Object Interactions
von: Xu, Hao, et al.
Veröffentlicht: (2024)
von: Xu, Hao, et al.
Veröffentlicht: (2024)
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
von: Zhao, Zhuoran, et al.
Veröffentlicht: (2026)
von: Zhao, Zhuoran, et al.
Veröffentlicht: (2026)
Learning Spectral-Decomposed Tokens for Domain Generalized Semantic Segmentation
von: Yi, Jingjun, et al.
Veröffentlicht: (2024)
von: Yi, Jingjun, et al.
Veröffentlicht: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
Condition-Invariant Semantic Segmentation
von: Sakaridis, Christos, et al.
Veröffentlicht: (2023)
von: Sakaridis, Christos, et al.
Veröffentlicht: (2023)
Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
Multi-Granularity Hand Action Detection
von: Zhe, Ting, et al.
Veröffentlicht: (2023)
von: Zhe, Ting, et al.
Veröffentlicht: (2023)
Dual-Stream Alignment for Action Segmentation
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
von: Gammulle, Harshala, et al.
Veröffentlicht: (2025)
Efficient and Effective Weakly-Supervised Action Segmentation via Action-Transition-Aware Boundary Alignment
von: Xu, Angchi, et al.
Veröffentlicht: (2024)
von: Xu, Angchi, et al.
Veröffentlicht: (2024)
Marine Saliency Segmenter: Object-Focused Conditional Diffusion with Region-Level Semantic Knowledge Distillation
von: Chang, Laibin, et al.
Veröffentlicht: (2025)
von: Chang, Laibin, et al.
Veröffentlicht: (2025)
Unified Multimodal Coherent Field: Synchronous Semantic-Spatial-Vision Fusion for Brain Tumor Segmentation
von: Zhang, Mingda, et al.
Veröffentlicht: (2025)
von: Zhang, Mingda, et al.
Veröffentlicht: (2025)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
A Dual-Mode ViT-Conditioned Diffusion Framework with an Adaptive Conditioning Bridge for Breast Cancer Segmentation
von: Singh, Prateek, et al.
Veröffentlicht: (2025)
von: Singh, Prateek, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GLoG-CSUnet: Enhancing Vision Transformers with Adaptable Radiomic Features for Medical Image Segmentation
von: Eghbali, Niloufar, et al.
Veröffentlicht: (2025) -
MoENAS: Mixture-of-Expert based Neural Architecture Search for jointly Accurate, Fair, and Robust Edge Deep Neural Networks
von: Mecharbat, Lotfi Abdelkrim, et al.
Veröffentlicht: (2025) -
Feature Imitating Networks Enhance The Performance, Reliability And Speed Of Deep Learning On Biomedical Image Processing Tasks
von: Min, Shangyang, et al.
Veröffentlicht: (2023) -
Vision Transformers: From Semantic Segmentation to Dense Prediction
von: Zhang, Li, et al.
Veröffentlicht: (2022) -
PointDiffuse: A Dual-Conditional Diffusion Model for Enhanced Point Cloud Semantic Segmentation
von: He, Yong, et al.
Veröffentlicht: (2025)