SORCE: Small Object Retrieval in Complex Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Chunxu, Xie, Chi, Chen, Xiaxu, Li, Wei, Zhu, Feng, Zhao, Rui, Wang, Limin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
von: Chen, Xiaxu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaxu, et al.
Veröffentlicht: (2025)
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
von: Liu, Chunxu, et al.
Veröffentlicht: (2025)
Sparse Global Matching for Video Frame Interpolation with Large Motion
von: Liu, Chunxu, et al.
Veröffentlicht: (2024)
von: Liu, Chunxu, et al.
Veröffentlicht: (2024)
History-Aware Transformation of ReID Features for Multiple Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2025)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2025)
FreeRet: MLLMs as Training-Free Retrievers
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)
Described Object Detection: Liberating Object Detection with Flexible Expressions
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
On the Robustness of Human-Object Interaction Detection against Distribution Shift
von: Xie, Chi, et al.
Veröffentlicht: (2025)
von: Xie, Chi, et al.
Veröffentlicht: (2025)
VFIMamba: Video Frame Interpolation with State Space Models
von: Zhang, Guozhen, et al.
Veröffentlicht: (2024)
von: Zhang, Guozhen, et al.
Veröffentlicht: (2024)
ADUGS-VINS: Generalized Visual-Inertial Odometry for Robust Navigation in Highly Dynamic and Complex Environments
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
von: Zhou, Rui, et al.
Veröffentlicht: (2024)
Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
von: Zhu, Yun, et al.
Veröffentlicht: (2026)
von: Zhu, Yun, et al.
Veröffentlicht: (2026)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
von: Zhu, Chenhui, et al.
Veröffentlicht: (2025)
AD-Det: Boosting Object Detection in UAV Images with Focused Small Objects and Balanced Tail Classes
von: Li, Zhenteng, et al.
Veröffentlicht: (2025)
von: Li, Zhenteng, et al.
Veröffentlicht: (2025)
Multiple Object Tracking as ID Prediction
von: Gao, Ruopeng, et al.
Veröffentlicht: (2024)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2024)
RemDet: Rethinking Efficient Model Design for UAV Object Detection
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
StageInteractor: Query-based Object Detector with Cross-stage Interaction
von: Teng, Yao, et al.
Veröffentlicht: (2023)
von: Teng, Yao, et al.
Veröffentlicht: (2023)
Multi-modal Reference Learning for Fine-grained Text-to-Image Retrieval
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
von: Ma, Zehong, et al.
Veröffentlicht: (2025)
CycleHOI: Improving Human-Object Interaction Detection with Cycle Consistency of Detection and Generation
von: Wang, Yisen, et al.
Veröffentlicht: (2024)
von: Wang, Yisen, et al.
Veröffentlicht: (2024)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
von: Wang, Tong, et al.
Veröffentlicht: (2025)
von: Wang, Tong, et al.
Veröffentlicht: (2025)
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
von: Green, Michael, et al.
Veröffentlicht: (2025)
von: Green, Michael, et al.
Veröffentlicht: (2025)
Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
von: Liang, Siyuan, et al.
Veröffentlicht: (2024)
EscapeCraft: A 3D Room Escape Environment for Benchmarking Complex Multimodal Reasoning Ability
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
von: Wang, Ziyue, et al.
Veröffentlicht: (2025)
Small Object Detection in Complex Backgrounds with Multi-Scale Attention and Global Relation Modeling
von: Tao, Wenguang, et al.
Veröffentlicht: (2026)
von: Tao, Wenguang, et al.
Veröffentlicht: (2026)
Effective Gaussian Management for High-fidelity Object Reconstruction
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
von: Liu, Lemin, et al.
Veröffentlicht: (2025)
von: Liu, Lemin, et al.
Veröffentlicht: (2025)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2023)
von: Zhao, Zhiyu, et al.
Veröffentlicht: (2023)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
von: Li, Xinhao, et al.
Veröffentlicht: (2023)
Retrieval-Augmented Egocentric Video Captioning
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
von: Xu, Jilan, et al.
Veröffentlicht: (2024)
SOEDiff: Efficient Distillation for Small Object Editing
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
Re-Aligning Language to Visual Objects with an Agentic Workflow
von: Chen, Yuming, et al.
Veröffentlicht: (2025)
von: Chen, Yuming, et al.
Veröffentlicht: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
Boltzmann Attention Sampling for Image Analysis with Small Objects
von: Zhao, Theodore, et al.
Veröffentlicht: (2025)
von: Zhao, Theodore, et al.
Veröffentlicht: (2025)
GLRT-Based Metric Learning for Remote Sensing Object Retrieval
von: Zhang, Linping, et al.
Veröffentlicht: (2024)
von: Zhang, Linping, et al.
Veröffentlicht: (2024)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Object-Centric Framework for Video Moment Retrieval
von: Li, Zongyao, et al.
Veröffentlicht: (2025)
von: Li, Zongyao, et al.
Veröffentlicht: (2025)
Small Object Detection Model with Spatial Laplacian Pyramid Attention and Multi-Scale Features Enhancement in Aerial Images
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
von: Wang, Zhichuan, et al.
Veröffentlicht: (2025)
von: Wang, Zhichuan, et al.
Veröffentlicht: (2025)
WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
von: Zhao, Haoren, et al.
Veröffentlicht: (2026)
Exploiting Scale-Variant Attention for Segmenting Small Medical Objects
von: Dai, Wei, et al.
Veröffentlicht: (2024)
von: Dai, Wei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
von: Chen, Xiaxu, et al.
Veröffentlicht: (2025) -
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
von: Liu, Chunxu, et al.
Veröffentlicht: (2025) -
Sparse Global Matching for Video Frame Interpolation with Large Motion
von: Liu, Chunxu, et al.
Veröffentlicht: (2024) -
History-Aware Transformation of ReID Features for Multiple Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2025) -
FreeRet: MLLMs as Training-Free Retrievers
von: Zhu, Yuhan, et al.
Veröffentlicht: (2025)