STARFlow: Spatial Temporal Feature Re-embedding with Attentive Learning for Real-world Scene Flow
Fuente:
arXiv
Guardado en:
| Autores principales: | Lu, Zhiyang, Chen, Qinghan, Cheng, Ming |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SSRFlow: Semantic-aware Fusion with Spatial Temporal Re-embedding for Real-world Scene Flow
por: Lu, Zhiyang, et al.
Publicado: (2024)
por: Lu, Zhiyang, et al.
Publicado: (2024)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
por: Gu, Jiatao, et al.
Publicado: (2025)
por: Gu, Jiatao, et al.
Publicado: (2025)
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
por: Lu, Zhiyang, et al.
Publicado: (2026)
por: Lu, Zhiyang, et al.
Publicado: (2026)
ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction
por: Liang, Yanzhe, et al.
Publicado: (2026)
por: Liang, Yanzhe, et al.
Publicado: (2026)
ARBEx: Attentive Feature Extraction with Reliability Balancing for Robust Facial Expression Learning
por: Wasi, Azmine Toushik, et al.
Publicado: (2023)
por: Wasi, Azmine Toushik, et al.
Publicado: (2023)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
por: Li, Xin, et al.
Publicado: (2024)
por: Li, Xin, et al.
Publicado: (2024)
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation
por: Hu, Youbing, et al.
Publicado: (2025)
por: Hu, Youbing, et al.
Publicado: (2025)
SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
por: Ma, Ziping, et al.
Publicado: (2024)
por: Ma, Ziping, et al.
Publicado: (2024)
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition
por: Liu, Feng, et al.
Publicado: (2025)
por: Liu, Feng, et al.
Publicado: (2025)
MambaFlow: A Novel and Flow-guided State Space Model for Scene Flow Estimation
por: Luo, Jiehao, et al.
Publicado: (2025)
por: Luo, Jiehao, et al.
Publicado: (2025)
FedRSU: Federated Learning for Scene Flow Estimation on Roadside Units
por: Fang, Shaoheng, et al.
Publicado: (2024)
por: Fang, Shaoheng, et al.
Publicado: (2024)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
por: Li, Bing, et al.
Publicado: (2022)
por: Li, Bing, et al.
Publicado: (2022)
Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions
por: Lu, Zhiyang, et al.
Publicado: (2026)
por: Lu, Zhiyang, et al.
Publicado: (2026)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
por: Ropero, Fernando, et al.
Publicado: (2026)
por: Ropero, Fernando, et al.
Publicado: (2026)
Attentive Graph Enhanced Region Representation Learning
por: Chen, Weiliang, et al.
Publicado: (2023)
por: Chen, Weiliang, et al.
Publicado: (2023)
Text-guided Feature Disentanglement for Cross-modal Gait Recognition
por: Lu, Zhiyang, et al.
Publicado: (2026)
por: Lu, Zhiyang, et al.
Publicado: (2026)
Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation
por: Yang, Xiaochen, et al.
Publicado: (2026)
por: Yang, Xiaochen, et al.
Publicado: (2026)
DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
por: Wang, Zhechao, et al.
Publicado: (2026)
por: Wang, Zhechao, et al.
Publicado: (2026)
From Easy to Hard: Learning Curricular Shape-aware Features for Robust Panoptic Scene Graph Generation
por: Shi, Hanrong, et al.
Publicado: (2024)
por: Shi, Hanrong, et al.
Publicado: (2024)
Coarse-to-Fine Domain Incremental Learning with Attentive Distillation for Mining Footprint Segmentation in Multispectral Imagery
por: Handoyo, Alif Tri, et al.
Publicado: (2026)
por: Handoyo, Alif Tri, et al.
Publicado: (2026)
Image-Goal Navigation Using Refined Feature Guidance and Scene Graph Enhancement
por: Feng, Zhicheng, et al.
Publicado: (2025)
por: Feng, Zhicheng, et al.
Publicado: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
por: Jun, Youngjun, et al.
Publicado: (2026)
por: Jun, Youngjun, et al.
Publicado: (2026)
STSA: Spatial-Temporal Semantic Alignment for Visual Dubbing
por: Ding, Zijun, et al.
Publicado: (2025)
por: Ding, Zijun, et al.
Publicado: (2025)
Feature-Aware Noise Contrastive Learning for Unsupervised Red Panda Re-Identification
por: Zhang, Jincheng, et al.
Publicado: (2024)
por: Zhang, Jincheng, et al.
Publicado: (2024)
Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2 Feature Representations for Video Action Analysis
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
por: Kodathala, Sai Varun, et al.
Publicado: (2025)
Layer-Wise Feature Metric of Semantic-Pixel Matching for Few-Shot Learning
por: Tang, Hao, et al.
Publicado: (2024)
por: Tang, Hao, et al.
Publicado: (2024)
VoteFlow: Enforcing Local Rigidity in Self-Supervised Scene Flow
por: Lin, Yancong, et al.
Publicado: (2025)
por: Lin, Yancong, et al.
Publicado: (2025)
Scene Structure Guidance Network: Unfolding Graph Partitioning into Pixel-Wise Feature Learning
por: Shin, Jisu, et al.
Publicado: (2023)
por: Shin, Jisu, et al.
Publicado: (2023)
Towards Open-world Generalized Deepfake Detection: General Feature Extraction via Unsupervised Domain Adaptation
por: Guo, Midou, et al.
Publicado: (2025)
por: Guo, Midou, et al.
Publicado: (2025)
REMOTE: Real-time Ego-motion Tracking for Various Endoscopes via Multimodal Visual Feature Learning
por: Shao, Liangjing, et al.
Publicado: (2025)
por: Shao, Liangjing, et al.
Publicado: (2025)
TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text Editing
por: Bao, Yuchen, et al.
Publicado: (2025)
por: Bao, Yuchen, et al.
Publicado: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
por: Wang, Haibo, et al.
Publicado: (2024)
por: Wang, Haibo, et al.
Publicado: (2024)
Self-Attentive Spatio-Temporal Calibration for Precise Intermediate Layer Matching in ANN-to-SNN Distillation
por: Hong, Di, et al.
Publicado: (2025)
por: Hong, Di, et al.
Publicado: (2025)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
por: Back, Seunghyeok, et al.
Publicado: (2025)
por: Back, Seunghyeok, et al.
Publicado: (2025)
Learner Attentiveness and Engagement Analysis in Online Education Using Computer Vision
por: Gogawale, Sharva, et al.
Publicado: (2024)
por: Gogawale, Sharva, et al.
Publicado: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
por: Kim, Namho, et al.
Publicado: (2025)
por: Kim, Namho, et al.
Publicado: (2025)
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification
por: Tan, Xuan, et al.
Publicado: (2024)
por: Tan, Xuan, et al.
Publicado: (2024)
An Attentive Representative Sample Selection Strategy Combined with Balanced Batch Training for Skin Lesion Segmentation
por: Lloyd-Brown, Stephen, et al.
Publicado: (2025)
por: Lloyd-Brown, Stephen, et al.
Publicado: (2025)
SwinSF: Image Reconstruction from Spatial-Temporal Spike Streams
por: Jiang, Liangyan, et al.
Publicado: (2024)
por: Jiang, Liangyan, et al.
Publicado: (2024)
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
por: Yu, Huangyue, et al.
Publicado: (2025)
por: Yu, Huangyue, et al.
Publicado: (2025)
Ejemplares similares
-
SSRFlow: Semantic-aware Fusion with Spatial Temporal Re-embedding for Real-world Scene Flow
por: Lu, Zhiyang, et al.
Publicado: (2024) -
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
por: Gu, Jiatao, et al.
Publicado: (2025) -
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
por: Lu, Zhiyang, et al.
Publicado: (2026) -
ReFlow: Self-correction Motion Learning for Dynamic Scene Reconstruction
por: Liang, Yanzhe, et al.
Publicado: (2026) -
ARBEx: Attentive Feature Extraction with Reliability Balancing for Robust Facial Expression Learning
por: Wasi, Azmine Toushik, et al.
Publicado: (2023)