AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Teng, Liu, Yihan, Chen, Jiongxu, Wang, Teng, Li, Jiaqi, Zhong, Bingzhuo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PiCo: Active Manifold Canonicalization for Robust Robotic Visual Anomaly Detection
by: Yan, Teng, et al.
Published: (2026)
by: Yan, Teng, et al.
Published: (2026)
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
by: Nguyen, Huy, et al.
Published: (2024)
by: Nguyen, Huy, et al.
Published: (2024)
Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification
by: Wen, Wen, et al.
Published: (2026)
by: Wen, Wen, et al.
Published: (2026)
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
by: Teng, Wenbin, et al.
Published: (2025)
by: Teng, Wenbin, et al.
Published: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
LongVPO: From Anchored Cues to Self-Reasoning for Long-Form Video Preference Optimization
by: Huang, Zhenpeng, et al.
Published: (2026)
by: Huang, Zhenpeng, et al.
Published: (2026)
Skyeyes: Ground Roaming using Aerial View Images
by: Gao, Zhiyuan, et al.
Published: (2024)
by: Gao, Zhiyuan, et al.
Published: (2024)
AG-VPReID 2025: Aerial-Ground Video-based Person Re-identification Challenge Results
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
ARSS: Taming Decoder-only Autoregressive Visual Generation for View Synthesis From Single View
by: Teng, Wenbin, et al.
Published: (2025)
by: Teng, Wenbin, et al.
Published: (2025)
DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation
by: Ouyang, Junyi, et al.
Published: (2026)
by: Ouyang, Junyi, et al.
Published: (2026)
Spatial-Temporal Decoupled Reference Conditioning for Identity-Preserving Text-to-Video Generation
by: Chen, Yuheng, et al.
Published: (2026)
by: Chen, Yuheng, et al.
Published: (2026)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
Long Video Understanding with Learnable Retrieval in Video-Language Models
by: Xu, Jiaqi, et al.
Published: (2023)
by: Xu, Jiaqi, et al.
Published: (2023)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation
by: Li, Yonglin, et al.
Published: (2023)
by: Li, Yonglin, et al.
Published: (2023)
ReTool-Video: Recursive Tool-Using Video Agents with Meta-Augmented Tool Grounding
by: Liu, Xiao, et al.
Published: (2026)
by: Liu, Xiao, et al.
Published: (2026)
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
by: Pu, Junfu, et al.
Published: (2025)
by: Pu, Junfu, et al.
Published: (2025)
View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
by: Li, Kaining, et al.
Published: (2025)
by: Li, Kaining, et al.
Published: (2025)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
by: Zou, Kai, et al.
Published: (2026)
by: Zou, Kai, et al.
Published: (2026)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Goodbye Drift: Anchored Tree Sampling for Long-Horizon Video-to-Video Generation
by: Bendel, Matthew, et al.
Published: (2026)
by: Bendel, Matthew, et al.
Published: (2026)
DeAR: Fine-Grained VLM Adaptation by Decomposing Attention Head Roles
by: Ma, Yiming, et al.
Published: (2026)
by: Ma, Yiming, et al.
Published: (2026)
AnchoredDream: Zero-Shot 360° Indoor Scene Generation from a Single View via Geometric Grounding
by: Yao, Runmao, et al.
Published: (2026)
by: Yao, Runmao, et al.
Published: (2026)
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation
by: Lin, Ci-Siang, et al.
Published: (2024)
by: Lin, Ci-Siang, et al.
Published: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
by: Jiang, Jianwen, et al.
Published: (2024)
by: Jiang, Jianwen, et al.
Published: (2024)
SD-ReID: View-aware Stable Diffusion for Aerial-Ground Person Re-Identification
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
by: Gu, Xin, et al.
Published: (2026)
by: Gu, Xin, et al.
Published: (2026)
Event-based Video Person Re-identification via Cross-Modality and Temporal Collaboration
by: Li, Renkai, et al.
Published: (2025)
by: Li, Renkai, et al.
Published: (2025)
View-decoupled Transformer for Person Re-identification under Aerial-ground Camera Network
by: Zhang, Quan, et al.
Published: (2024)
by: Zhang, Quan, et al.
Published: (2024)
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
by: Zhu, Lanyun, et al.
Published: (2025)
by: Zhu, Lanyun, et al.
Published: (2025)
AdaFV: Rethinking of Visual-Language alignment for VLM acceleration
by: Han, Jiayi, et al.
Published: (2025)
by: Han, Jiayi, et al.
Published: (2025)
Recurrence over Video Frames (RoVF) for the Re-identification of Meerkats
by: Rogers, Mitchell, et al.
Published: (2024)
by: Rogers, Mitchell, et al.
Published: (2024)
LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
by: Yan, Cilin, et al.
Published: (2025)
by: Yan, Cilin, et al.
Published: (2025)
AR4D: Autoregressive 4D Generation from Monocular Videos
by: Zhu, Hanxin, et al.
Published: (2025)
by: Zhu, Hanxin, et al.
Published: (2025)
REC-RL: Referring expression counting via Gaussian and range-based reward optimization
by: Liu, Hui, et al.
Published: (2026)
by: Liu, Hui, et al.
Published: (2026)
View while Moving: Efficient Video Recognition in Long-untrimmed Videos
by: Tian, Ye, et al.
Published: (2023)
by: Tian, Ye, et al.
Published: (2023)
Similar Items
-
PiCo: Active Manifold Canonicalization for Robust Robotic Visual Anomaly Detection
by: Yan, Teng, et al.
Published: (2026) -
AG-ReID.v2: Bridging Aerial and Ground Views for Person Re-identification
by: Nguyen, Huy, et al.
Published: (2024) -
Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification
by: Wen, Wen, et al.
Published: (2026) -
FVGen: Accelerating Novel-View Synthesis with Adversarial Video Diffusion Distillation
by: Teng, Wenbin, et al.
Published: (2025) -
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)