InstructDET: Diversifying Referring Object Detection with Generalized Instructions
Fuente:
arXiv
Salvato in:
| Autori principali: | Dang, Ronghao, Feng, Jiangyan, Zhang, Haodong, Ge, Chongjian, Song, Lin, Gong, Lijun, Liu, Chengju, Chen, Qijun, Zhu, Feng, Zhao, Rui, Song, Yibing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Re-Aligning Language to Visual Objects with an Agentic Workflow
di: Chen, Yuming, et al.
Pubblicazione: (2025)
di: Chen, Yuming, et al.
Pubblicazione: (2025)
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
di: Lin, Xiao, et al.
Pubblicazione: (2024)
di: Lin, Xiao, et al.
Pubblicazione: (2024)
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
di: Zhu, Minghao, et al.
Pubblicazione: (2023)
di: Zhu, Minghao, et al.
Pubblicazione: (2023)
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
di: Zhu, Minghao, et al.
Pubblicazione: (2024)
di: Zhu, Minghao, et al.
Pubblicazione: (2024)
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
Vision-and-Language Navigation via Causal Learning
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
di: Wang, Liuyi, et al.
Pubblicazione: (2024)
NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
di: He, Zongtao, et al.
Pubblicazione: (2025)
di: He, Zongtao, et al.
Pubblicazione: (2025)
A Dual Semantic-Aware Recurrent Global-Adaptive Network For Vision-and-Language Navigation
di: Wang, Liuyi, et al.
Pubblicazione: (2023)
di: Wang, Liuyi, et al.
Pubblicazione: (2023)
TransPose: 6D Object Pose Estimation with Geometry-Aware Transformer
di: Lin, Xiao, et al.
Pubblicazione: (2023)
di: Lin, Xiao, et al.
Pubblicazione: (2023)
CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge Distillation
di: Lin, Xiao, et al.
Pubblicazione: (2025)
di: Lin, Xiao, et al.
Pubblicazione: (2025)
NEU-DET-and-GC10-DET-with-YOLO-fromat
di: guan, li
Pubblicazione: (2025)
di: guan, li
Pubblicazione: (2025)
InstructVEdit: A Holistic Approach for Instructional Video Editing
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
MusicDET: Zero-Shot AI-Generated Music Detection
di: Han, Chaolei, et al.
Pubblicazione: (2026)
di: Han, Chaolei, et al.
Pubblicazione: (2026)
OA-DET3D: Embedding Object Awareness as a General Plug-in for Multi-Camera 3D Object Detection
di: Chu, Xiaomeng, et al.
Pubblicazione: (2023)
di: Chu, Xiaomeng, et al.
Pubblicazione: (2023)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
di: Zhou, Gengze, et al.
Pubblicazione: (2025)
di: Zhou, Gengze, et al.
Pubblicazione: (2025)
Learning to Instruct for Visual Instruction Tuning
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
di: Zhou, Zhihan, et al.
Pubblicazione: (2025)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
Instruct-Imagen: Image Generation with Multi-modal Instruction
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
di: Hu, Hexiang, et al.
Pubblicazione: (2024)
Knowledge Amalgamation for Object Detection with Transformers
di: Zhang, Haofei, et al.
Pubblicazione: (2022)
di: Zhang, Haofei, et al.
Pubblicazione: (2022)
On the Robustness of Object Detection Models on Aerial Images
di: He, Haodong, et al.
Pubblicazione: (2023)
di: He, Haodong, et al.
Pubblicazione: (2023)
InstructAttribute: Fine-grained Object Attributes editing with Instruction
di: Yin, Xingxi, et al.
Pubblicazione: (2025)
di: Yin, Xingxi, et al.
Pubblicazione: (2025)
SAM-LAD: Segment Anything Model Meets Zero-Shot Logic Anomaly Detection
di: Peng, Yun, et al.
Pubblicazione: (2024)
di: Peng, Yun, et al.
Pubblicazione: (2024)
Efficient Text-driven Motion Generation via Latent Consistency Training
di: Hu, Mengxian, et al.
Pubblicazione: (2024)
di: Hu, Mengxian, et al.
Pubblicazione: (2024)
Foley-Flow: Coordinated Video-to-Audio Generation with Masked Audio-Visual Alignment and Dynamic Conditional Flows
di: Mo, Shentong, et al.
Pubblicazione: (2026)
di: Mo, Shentong, et al.
Pubblicazione: (2026)
Secure-Instruct: An Automated Pipeline for Synthesizing Instruction-Tuning Datasets Using LLMs for Secure Code Generation
di: Li, Junjie, et al.
Pubblicazione: (2025)
di: Li, Junjie, et al.
Pubblicazione: (2025)
Instruct-ReID: A Multi-purpose Person Re-identification Task with Instructions
di: He, Weizhen, et al.
Pubblicazione: (2023)
di: He, Weizhen, et al.
Pubblicazione: (2023)
FinPos: A Position-Aware Trading Agent System for Real Financial Markets
di: Liu, Bijia, et al.
Pubblicazione: (2025)
di: Liu, Bijia, et al.
Pubblicazione: (2025)
FINRS: A Risk-Sensitive Trading Framework for Real Financial Markets
di: Liu, Bijia, et al.
Pubblicazione: (2025)
di: Liu, Bijia, et al.
Pubblicazione: (2025)
Realizing Text-Driven Motion Generation on NAO Robot: A Reinforcement Learning-Optimized Control Pipeline
di: Xu, Zihan, et al.
Pubblicazione: (2025)
di: Xu, Zihan, et al.
Pubblicazione: (2025)
HICO-DET-SG and V-COCO-SG: New Data Splits for Evaluating the Systematic Generalization Performance of Human-Object Interaction Detection Models
di: Takemoto, Kentaro, et al.
Pubblicazione: (2023)
di: Takemoto, Kentaro, et al.
Pubblicazione: (2023)
Adaptive Denoising-Enhanced LiDAR Odometry for Degeneration Resilience in Diverse Terrains
di: Ji, Mazeyu, et al.
Pubblicazione: (2023)
di: Ji, Mazeyu, et al.
Pubblicazione: (2023)
Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation
di: Xiao, Kaiyan, et al.
Pubblicazione: (2025)
di: Xiao, Kaiyan, et al.
Pubblicazione: (2025)
DEYO: DETR with YOLO for End-to-End Object Detection
di: Ouyang, Haodong
Pubblicazione: (2024)
di: Ouyang, Haodong
Pubblicazione: (2024)
Ada-Instruct: Adapting Instruction Generators for Complex Reasoning
di: Cui, Wanyun, et al.
Pubblicazione: (2023)
di: Cui, Wanyun, et al.
Pubblicazione: (2023)
Described Object Detection: Liberating Object Detection with Flexible Expressions
di: Xie, Chi, et al.
Pubblicazione: (2023)
di: Xie, Chi, et al.
Pubblicazione: (2023)
New Reference: Diversifying Service Delivery.
di: Garner, Imogen
Pubblicazione: (1999)
di: Garner, Imogen
Pubblicazione: (1999)
InstructUDrag: Joint Text Instructions and Object Dragging for Interactive Image Editing
di: Yu, Haoran, et al.
Pubblicazione: (2025)
di: Yu, Haoran, et al.
Pubblicazione: (2025)
MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark
di: Song, Anyang, et al.
Pubblicazione: (2026)
di: Song, Anyang, et al.
Pubblicazione: (2026)
AS400-DET: Detection using Deep Learning Model for IBM i (AS/400)
di: Tran, Thanh, et al.
Pubblicazione: (2025)
di: Tran, Thanh, et al.
Pubblicazione: (2025)
Scoring, Remember, and Reference: Catching Camouflaged Objects in Videos
di: Feng, Yuang, et al.
Pubblicazione: (2025)
di: Feng, Yuang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Re-Aligning Language to Visual Objects with an Agentic Workflow
di: Chen, Yuming, et al.
Pubblicazione: (2025) -
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
di: Lin, Xiao, et al.
Pubblicazione: (2024) -
Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation Learning
di: Zhu, Minghao, et al.
Pubblicazione: (2023) -
MoTE: Reconciling Generalization with Specialization for Visual-Language to Video Knowledge Transfer
di: Zhu, Minghao, et al.
Pubblicazione: (2024) -
Causality-based Cross-Modal Representation Learning for Vision-and-Language Navigation
di: Wang, Liuyi, et al.
Pubblicazione: (2024)