TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial Scenes
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Liangyu, Lu, Wanxuan, Yu, Hongfeng, Mao, Yongqiang, Bi, Hanbo, Liu, Chenglong, Sun, Xian, Fu, Kun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
di: Xu, Liangyu, et al.
Pubblicazione: (2024)
di: Xu, Liangyu, et al.
Pubblicazione: (2024)
SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multi-View Stereo Reconstruction in Remote Sensing
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024)
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024)
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
di: Diao, Wenhui, et al.
Pubblicazione: (2024)
di: Diao, Wenhui, et al.
Pubblicazione: (2024)
Twin Deformable Point Convolutions for Point Cloud Semantic Segmentation in Remote Sensing Scenes
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024)
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024)
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
di: Bi, Hanbo, et al.
Pubblicazione: (2024)
di: Bi, Hanbo, et al.
Pubblicazione: (2024)
Prompt-and-Transfer: Dynamic Class-aware Enhancement for Few-shot Segmentation
di: Bi, Hanbo, et al.
Pubblicazione: (2024)
di: Bi, Hanbo, et al.
Pubblicazione: (2024)
ReCon1M:A Large-scale Benchmark Dataset for Relation Comprehension in Remote Sensing Imagery
di: Sun, Xian, et al.
Pubblicazione: (2024)
di: Sun, Xian, et al.
Pubblicazione: (2024)
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
di: Jin, Yang, et al.
Pubblicazione: (2024)
di: Jin, Yang, et al.
Pubblicazione: (2024)
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
di: Hu, Huiyang, et al.
Pubblicazione: (2025)
di: Hu, Huiyang, et al.
Pubblicazione: (2025)
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
di: Bi, Hanbo, et al.
Pubblicazione: (2025)
Stable and High-Precision 3D Positioning via Tunable Composite-Dimensional Hong-Ou-Mandel Interference
di: Li, Yongqiang, et al.
Pubblicazione: (2025)
di: Li, Yongqiang, et al.
Pubblicazione: (2025)
PhysioFormer: Integrating Multimodal Physiological Signals and Symbolic Regression for Explainable Affective State Prediction
di: Wang, Zhifeng, et al.
Pubblicazione: (2024)
di: Wang, Zhifeng, et al.
Pubblicazione: (2024)
Target-aware Bidirectional Fusion Transformer for Aerial Object Tracking
di: Sun, Xinglong, et al.
Pubblicazione: (2025)
di: Sun, Xinglong, et al.
Pubblicazione: (2025)
ControlMTR: Control-Guided Motion Transformer with Scene-Compliant Intention Points for Feasible Motion Prediction
di: Sun, Jiawei, et al.
Pubblicazione: (2024)
di: Sun, Jiawei, et al.
Pubblicazione: (2024)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
di: Yang, Yuxiao, et al.
Pubblicazione: (2025)
di: Yang, Yuxiao, et al.
Pubblicazione: (2025)
Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes
di: Jiang, Lihan, et al.
Pubblicazione: (2024)
di: Jiang, Lihan, et al.
Pubblicazione: (2024)
EMatch: A Unified Framework for Event-based Optical Flow and Stereo Matching
di: Zhang, Pengjie, et al.
Pubblicazione: (2024)
di: Zhang, Pengjie, et al.
Pubblicazione: (2024)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
di: Zhou, Zhenghong, et al.
Pubblicazione: (2026)
di: Zhou, Zhenghong, et al.
Pubblicazione: (2026)
Scene-aware Human Motion Forecasting via Mutual Distance Prediction
di: Xing, Chaoyue, et al.
Pubblicazione: (2023)
di: Xing, Chaoyue, et al.
Pubblicazione: (2023)
Prototypical Transformer as Unified Motion Learners
di: Han, Cheng, et al.
Pubblicazione: (2024)
di: Han, Cheng, et al.
Pubblicazione: (2024)
JointMotion: Joint Self-Supervision for Joint Motion Prediction
di: Wagner, Royden, et al.
Pubblicazione: (2024)
di: Wagner, Royden, et al.
Pubblicazione: (2024)
On Evaluating GPT ‐4 for CHB Treatment Recommendations: Reproducibility, Vignette Design and Prompting
di: Siqi Sun, et al.
Pubblicazione: (2025)
di: Siqi Sun, et al.
Pubblicazione: (2025)
SceneGlue: Scene-Aware Transformer for Feature Matching without Scene-Level Annotation
di: Du, Songlin, et al.
Pubblicazione: (2026)
di: Du, Songlin, et al.
Pubblicazione: (2026)
Inter-object Discriminative Graph Modeling for Indoor Scene Recognition
di: Song, Chuanxin, et al.
Pubblicazione: (2023)
di: Song, Chuanxin, et al.
Pubblicazione: (2023)
LMFormer: Lane based Motion Prediction Transformer
di: Yadav, Harsh, et al.
Pubblicazione: (2025)
di: Yadav, Harsh, et al.
Pubblicazione: (2025)
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
di: Yu, Haiyang, et al.
Pubblicazione: (2024)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
di: Chen, Seng Nam, et al.
Pubblicazione: (2026)
di: Chen, Seng Nam, et al.
Pubblicazione: (2026)
JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation
di: Chen, Fangda, et al.
Pubblicazione: (2025)
di: Chen, Fangda, et al.
Pubblicazione: (2025)
4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
di: Guo, Mengqi, et al.
Pubblicazione: (2025)
di: Guo, Mengqi, et al.
Pubblicazione: (2025)
Flying Calligrapher: Contact-Aware Motion and Force Planning and Control for Aerial Manipulation
di: Guo, Xiaofeng, et al.
Pubblicazione: (2024)
di: Guo, Xiaofeng, et al.
Pubblicazione: (2024)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
di: Wang, Wenchuan, et al.
Pubblicazione: (2025)
di: Wang, Wenchuan, et al.
Pubblicazione: (2025)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
di: Mei, Jianbiao, et al.
Pubblicazione: (2024)
di: Mei, Jianbiao, et al.
Pubblicazione: (2024)
Semantic-embedded Similarity Prototype for Scene Recognition
di: Song, Chuanxin, et al.
Pubblicazione: (2023)
di: Song, Chuanxin, et al.
Pubblicazione: (2023)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
di: Diao, Lingxiao, et al.
Pubblicazione: (2025)
di: Diao, Lingxiao, et al.
Pubblicazione: (2025)
Efficient Motion-Aware Video MLLM
di: Zhao, Zijia, et al.
Pubblicazione: (2025)
di: Zhao, Zijia, et al.
Pubblicazione: (2025)
RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images
di: Hu, Guyue, et al.
Pubblicazione: (2026)
di: Hu, Guyue, et al.
Pubblicazione: (2026)
What Happens Next? Next Scene Prediction with a Unified Video Model
di: Li, Xinjie, et al.
Pubblicazione: (2025)
di: Li, Xinjie, et al.
Pubblicazione: (2025)
Motion-Aware Transformer for Multi-Object Tracking
di: Yang, Xu, et al.
Pubblicazione: (2025)
di: Yang, Xu, et al.
Pubblicazione: (2025)
AdaMorph: Unified Motion Retargeting via Embodiment-Aware Adaptive Transformers
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
di: Xu, Liangyu, et al.
Pubblicazione: (2024) -
SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multi-View Stereo Reconstruction in Remote Sensing
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024) -
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
di: Diao, Wenhui, et al.
Pubblicazione: (2024) -
Twin Deformable Point Convolutions for Point Cloud Semantic Segmentation in Remote Sensing Scenes
di: Mao, Yong-Qiang, et al.
Pubblicazione: (2024) -
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
di: Bi, Hanbo, et al.
Pubblicazione: (2024)