Multi-Point Positional Insertion Tuning for Small Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Goto, Kanoko, Karasawa, Takumi, Hirose, Takumi, Kawakami, Rei, Inoue, Nakamasa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025)
by: Goto, Kanoko, et al.
Published: (2025)
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
by: Inoue, Nakamasa, et al.
Published: (2025)
by: Inoue, Nakamasa, et al.
Published: (2025)
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
by: Tanaka, Daichi, et al.
Published: (2025)
by: Tanaka, Daichi, et al.
Published: (2025)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
by: Inoue, Nakamasa, et al.
Published: (2024)
by: Inoue, Nakamasa, et al.
Published: (2024)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
by: Ukai, Mahiro, et al.
Published: (2025)
by: Ukai, Mahiro, et al.
Published: (2025)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
by: Yajima, Masaru, et al.
Published: (2025)
by: Yajima, Masaru, et al.
Published: (2025)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
by: Shen, Ruoyue, et al.
Published: (2024)
by: Shen, Ruoyue, et al.
Published: (2024)
Boundary and Position Information Mining for Aerial Small Object Detection
by: Huang, Rongxin, et al.
Published: (2026)
by: Huang, Rongxin, et al.
Published: (2026)
MOD-CL: Multi-label Object Detection with Constrained Loss
by: Moriyama, Sota, et al.
Published: (2024)
by: Moriyama, Sota, et al.
Published: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
by: Ukai, Mahiro, et al.
Published: (2024)
by: Ukai, Mahiro, et al.
Published: (2024)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
by: Yokomizo, Hisayuki, et al.
Published: (2026)
by: Yokomizo, Hisayuki, et al.
Published: (2026)
Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster
by: Kamada, Shin, et al.
Published: (2025)
by: Kamada, Shin, et al.
Published: (2025)
Controllable Video Object Insertion via Multiview Priors
by: Qi, Xia, et al.
Published: (2026)
by: Qi, Xia, et al.
Published: (2026)
CityNav: A Large-Scale Dataset for Real-World Aerial Navigation
by: Lee, Jungdae, et al.
Published: (2024)
by: Lee, Jungdae, et al.
Published: (2024)
Generative Object Insertion in Gaussian Splatting with a Multi-View Diffusion Model
by: Zhong, Hongliang, et al.
Published: (2024)
by: Zhong, Hongliang, et al.
Published: (2024)
CrimEdit: Controllable Editing for Counterfactual Object Removal, Insertion, and Movement
by: Jeon, Boseong, et al.
Published: (2025)
by: Jeon, Boseong, et al.
Published: (2025)
Improving the Detection of Small Oriented Objects in Aerial Images
by: Doloriel, Chandler Timm C., et al.
Published: (2024)
by: Doloriel, Chandler Timm C., et al.
Published: (2024)
Photorealistic Object Insertion with Diffusion-Guided Inverse Rendering
by: Liang, Ruofan, et al.
Published: (2024)
by: Liang, Ruofan, et al.
Published: (2024)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
by: Yu, Jiangyong, et al.
Published: (2025)
by: Yu, Jiangyong, et al.
Published: (2025)
Beyond CNNs: Efficient Fine-Tuning of Multi-Modal LLMs for Object Detection on Low-Data Regimes
by: Elamon, Nirmal, et al.
Published: (2025)
by: Elamon, Nirmal, et al.
Published: (2025)
MKSNet: Advanced Small Object Detection in Remote Sensing Imagery with Multi-Kernel and Dual Attention Mechanisms
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
Small Object Detection for Indoor Assistance to the Blind using YOLO NAS Small and Super Gradients
by: BN, Rashmi, et al.
Published: (2024)
by: BN, Rashmi, et al.
Published: (2024)
SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
by: Taguchi, Shun, et al.
Published: (2025)
by: Taguchi, Shun, et al.
Published: (2025)
Just a Hint: Point-Supervised Camouflaged Object Detection
by: Chen, Huafeng, et al.
Published: (2024)
by: Chen, Huafeng, et al.
Published: (2024)
SimInsert: Seamless Video Object Insertion via Regional Sparse Attention Fusion
by: Chen, Xinyu, et al.
Published: (2026)
by: Chen, Xinyu, et al.
Published: (2026)
An Empirical Study of Methods for Small Object Detection from Satellite Imagery
by: Yuan, Xiaohui, et al.
Published: (2025)
by: Yuan, Xiaohui, et al.
Published: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
Geo-ORBIT: A Federated Digital Twin Framework for Scene-Adaptive Lane Geometry Detection
by: Tamaru, Rei, et al.
Published: (2025)
by: Tamaru, Rei, et al.
Published: (2025)
PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
by: Zhang, Peiyuan, et al.
Published: (2025)
by: Zhang, Peiyuan, et al.
Published: (2025)
Behavior-Grounded Lane Representation Learning for Multi-Task Traffic Digital Twins
by: Tamaru, Rei, et al.
Published: (2026)
by: Tamaru, Rei, et al.
Published: (2026)
DANet: Enhancing Small Object Detection through an Efficient Deformable Attention Network
by: Mia, Md Sohag, et al.
Published: (2023)
by: Mia, Md Sohag, et al.
Published: (2023)
InsertAnywhere: Bridging 4D Scene Geometry and Diffusion Models for Realistic Video Object Insertion
by: Jin, Hoiyeong, et al.
Published: (2025)
by: Jin, Hoiyeong, et al.
Published: (2025)
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
by: Ren, Botao, et al.
Published: (2024)
by: Ren, Botao, et al.
Published: (2024)
TSOM: Small Object Motion Detection Neural Network Inspired by Avian Visual Circuit
by: Hu, Pignge, et al.
Published: (2024)
by: Hu, Pignge, et al.
Published: (2024)
SO-DETR: Leveraging Dual-Domain Features and Knowledge Distillation for Small Object Detection
by: Zhang, Huaxiang, et al.
Published: (2025)
by: Zhang, Huaxiang, et al.
Published: (2025)
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
by: Yu, Yi, et al.
Published: (2025)
by: Yu, Yi, et al.
Published: (2025)
Mixed Precision PointPillars for Efficient 3D Object Detection with TensorRT
by: Fuengfusin, Ninnart, et al.
Published: (2026)
by: Fuengfusin, Ninnart, et al.
Published: (2026)
PEFT-DML: Parameter-Efficient Fine-Tuning Deep Metric Learning for Robust Multi-Modal 3D Object Detection in Autonomous Driving
by: Rezaei, Abdolazim, et al.
Published: (2025)
by: Rezaei, Abdolazim, et al.
Published: (2025)
Similar Items
-
Referring Expression Comprehension for Small Objects
by: Goto, Kanoko, et al.
Published: (2025) -
DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
by: Inoue, Nakamasa, et al.
Published: (2025) -
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
by: Zhou, Yifan, et al.
Published: (2025) -
Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
by: Tanaka, Daichi, et al.
Published: (2025) -
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
by: Inoue, Nakamasa, et al.
Published: (2024)