TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial Scenes
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xu, Liangyu, Lu, Wanxuan, Yu, Hongfeng, Mao, Yongqiang, Bi, Hanbo, Liu, Chenglong, Sun, Xian, Fu, Kun |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
par: Xu, Liangyu, et autres
Publié: (2024)
par: Xu, Liangyu, et autres
Publié: (2024)
SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multi-View Stereo Reconstruction in Remote Sensing
par: Mao, Yong-Qiang, et autres
Publié: (2024)
par: Mao, Yong-Qiang, et autres
Publié: (2024)
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
par: Diao, Wenhui, et autres
Publié: (2024)
par: Diao, Wenhui, et autres
Publié: (2024)
Twin Deformable Point Convolutions for Point Cloud Semantic Segmentation in Remote Sensing Scenes
par: Mao, Yong-Qiang, et autres
Publié: (2024)
par: Mao, Yong-Qiang, et autres
Publié: (2024)
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
par: Bi, Hanbo, et autres
Publié: (2024)
par: Bi, Hanbo, et autres
Publié: (2024)
Prompt-and-Transfer: Dynamic Class-aware Enhancement for Few-shot Segmentation
par: Bi, Hanbo, et autres
Publié: (2024)
par: Bi, Hanbo, et autres
Publié: (2024)
ReCon1M:A Large-scale Benchmark Dataset for Relation Comprehension in Remote Sensing Imagery
par: Sun, Xian, et autres
Publié: (2024)
par: Sun, Xian, et autres
Publié: (2024)
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
par: Bi, Hanbo, et autres
Publié: (2025)
par: Bi, Hanbo, et autres
Publié: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
par: Jin, Yang, et autres
Publié: (2024)
par: Jin, Yang, et autres
Publié: (2024)
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
par: Hu, Huiyang, et autres
Publié: (2025)
par: Hu, Huiyang, et autres
Publié: (2025)
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
par: Bi, Hanbo, et autres
Publié: (2025)
par: Bi, Hanbo, et autres
Publié: (2025)
Stable and High-Precision 3D Positioning via Tunable Composite-Dimensional Hong-Ou-Mandel Interference
par: Li, Yongqiang, et autres
Publié: (2025)
par: Li, Yongqiang, et autres
Publié: (2025)
PhysioFormer: Integrating Multimodal Physiological Signals and Symbolic Regression for Explainable Affective State Prediction
par: Wang, Zhifeng, et autres
Publié: (2024)
par: Wang, Zhifeng, et autres
Publié: (2024)
Target-aware Bidirectional Fusion Transformer for Aerial Object Tracking
par: Sun, Xinglong, et autres
Publié: (2025)
par: Sun, Xinglong, et autres
Publié: (2025)
ControlMTR: Control-Guided Motion Transformer with Scene-Compliant Intention Points for Feasible Motion Prediction
par: Sun, Jiawei, et autres
Publié: (2024)
par: Sun, Jiawei, et autres
Publié: (2024)
EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
par: Yang, Yuxiao, et autres
Publié: (2025)
par: Yang, Yuxiao, et autres
Publié: (2025)
Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes
par: Jiang, Lihan, et autres
Publié: (2024)
par: Jiang, Lihan, et autres
Publié: (2024)
EMatch: A Unified Framework for Event-based Optical Flow and Stereo Matching
par: Zhang, Pengjie, et autres
Publié: (2024)
par: Zhang, Pengjie, et autres
Publié: (2024)
Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion
par: Zhou, Zhenghong, et autres
Publié: (2026)
par: Zhou, Zhenghong, et autres
Publié: (2026)
Scene-aware Human Motion Forecasting via Mutual Distance Prediction
par: Xing, Chaoyue, et autres
Publié: (2023)
par: Xing, Chaoyue, et autres
Publié: (2023)
Prototypical Transformer as Unified Motion Learners
par: Han, Cheng, et autres
Publié: (2024)
par: Han, Cheng, et autres
Publié: (2024)
JointMotion: Joint Self-Supervision for Joint Motion Prediction
par: Wagner, Royden, et autres
Publié: (2024)
par: Wagner, Royden, et autres
Publié: (2024)
On Evaluating GPT ‐4 for CHB Treatment Recommendations: Reproducibility, Vignette Design and Prompting
par: Siqi Sun, et autres
Publié: (2025)
par: Siqi Sun, et autres
Publié: (2025)
SceneGlue: Scene-Aware Transformer for Feature Matching without Scene-Level Annotation
par: Du, Songlin, et autres
Publié: (2026)
par: Du, Songlin, et autres
Publié: (2026)
Inter-object Discriminative Graph Modeling for Indoor Scene Recognition
par: Song, Chuanxin, et autres
Publié: (2023)
par: Song, Chuanxin, et autres
Publié: (2023)
LMFormer: Lane based Motion Prediction Transformer
par: Yadav, Harsh, et autres
Publié: (2025)
par: Yadav, Harsh, et autres
Publié: (2025)
EAFormer: Scene Text Segmentation with Edge-Aware Transformers
par: Yu, Haiyang, et autres
Publié: (2024)
par: Yu, Haiyang, et autres
Publié: (2024)
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark
par: Chen, Seng Nam, et autres
Publié: (2026)
par: Chen, Seng Nam, et autres
Publié: (2026)
JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation
par: Chen, Fangda, et autres
Publié: (2025)
par: Chen, Fangda, et autres
Publié: (2025)
4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos
par: Guo, Mengqi, et autres
Publié: (2025)
par: Guo, Mengqi, et autres
Publié: (2025)
Flying Calligrapher: Contact-Aware Motion and Force Planning and Control for Aerial Manipulation
par: Guo, Xiaofeng, et autres
Publié: (2024)
par: Guo, Xiaofeng, et autres
Publié: (2024)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
par: Wang, Wenchuan, et autres
Publié: (2025)
par: Wang, Wenchuan, et autres
Publié: (2025)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
par: Mei, Jianbiao, et autres
Publié: (2024)
par: Mei, Jianbiao, et autres
Publié: (2024)
Semantic-embedded Similarity Prototype for Scene Recognition
par: Song, Chuanxin, et autres
Publié: (2023)
par: Song, Chuanxin, et autres
Publié: (2023)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
par: Diao, Lingxiao, et autres
Publié: (2025)
par: Diao, Lingxiao, et autres
Publié: (2025)
Efficient Motion-Aware Video MLLM
par: Zhao, Zijia, et autres
Publié: (2025)
par: Zhao, Zijia, et autres
Publié: (2025)
RefAerial: A Benchmark and Approach for Referring Detection in Aerial Images
par: Hu, Guyue, et autres
Publié: (2026)
par: Hu, Guyue, et autres
Publié: (2026)
What Happens Next? Next Scene Prediction with a Unified Video Model
par: Li, Xinjie, et autres
Publié: (2025)
par: Li, Xinjie, et autres
Publié: (2025)
Motion-Aware Transformer for Multi-Object Tracking
par: Yang, Xu, et autres
Publié: (2025)
par: Yang, Xu, et autres
Publié: (2025)
AdaMorph: Unified Motion Retargeting via Embodiment-Aware Adaptive Transformers
par: Zhang, Haoyu, et autres
Publié: (2026)
par: Zhang, Haoyu, et autres
Publié: (2026)
Documents similaires
-
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
par: Xu, Liangyu, et autres
Publié: (2024) -
SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multi-View Stereo Reconstruction in Remote Sensing
par: Mao, Yong-Qiang, et autres
Publié: (2024) -
RingMo-Aerial: An Aerial Remote Sensing Foundation Model With Affine Transformation Contrastive Learning
par: Diao, Wenhui, et autres
Publié: (2024) -
Twin Deformable Point Convolutions for Point Cloud Semantic Segmentation in Remote Sensing Scenes
par: Mao, Yong-Qiang, et autres
Publié: (2024) -
AgMTR: Agent Mining Transformer for Few-shot Segmentation in Remote Sensing
par: Bi, Hanbo, et autres
Publié: (2024)