Generative Region-Language Pretraining for Open-Ended Object Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Lin, Chuang, Jiang, Yi, Qu, Lizhen, Yuan, Zehuan, Cai, Jianfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
di: Lin, Zhiwei, et al.
Pubblicazione: (2024)
Open-Det: An Efficient Learning Framework for Open-Ended Detection
di: Cao, Guiping, et al.
Pubblicazione: (2025)
di: Cao, Guiping, et al.
Pubblicazione: (2025)
Waver: Wave Your Way to Lifelike Video Generation
di: Zhang, Yifu, et al.
Pubblicazione: (2025)
di: Zhang, Yifu, et al.
Pubblicazione: (2025)
Recognize Any Regions
di: Yang, Haosen, et al.
Pubblicazione: (2023)
di: Yang, Haosen, et al.
Pubblicazione: (2023)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
di: Chen, Fangyi, et al.
Pubblicazione: (2024)
di: Chen, Fangyi, et al.
Pubblicazione: (2024)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
di: Wu, Junfeng, et al.
Pubblicazione: (2024)
di: Wu, Junfeng, et al.
Pubblicazione: (2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
di: Wang, Junke, et al.
Pubblicazione: (2024)
di: Wang, Junke, et al.
Pubblicazione: (2024)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
di: Lin, Chuang, et al.
Pubblicazione: (2024)
di: Lin, Chuang, et al.
Pubblicazione: (2024)
Region-centric Image-Language Pretraining for Open-Vocabulary Detection
di: Kim, Dahun, et al.
Pubblicazione: (2023)
di: Kim, Dahun, et al.
Pubblicazione: (2023)
Open World Object Detection: A Survey
di: Li, Yiming, et al.
Pubblicazione: (2024)
di: Li, Yiming, et al.
Pubblicazione: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
di: Tian, Keyu, et al.
Pubblicazione: (2024)
di: Tian, Keyu, et al.
Pubblicazione: (2024)
DifFUSER: Diffusion Model for Robust Multi-Sensor Fusion in 3D Object Detection and BEV Segmentation
di: Le, Duy-Tho, et al.
Pubblicazione: (2024)
di: Le, Duy-Tho, et al.
Pubblicazione: (2024)
MagicSeg: Open-World Segmentation Pretraining via Counterfactural Diffusion-Based Auto-Generation
di: Cai, Kaixin, et al.
Pubblicazione: (2026)
di: Cai, Kaixin, et al.
Pubblicazione: (2026)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
di: Sun, Peize, et al.
Pubblicazione: (2024)
di: Sun, Peize, et al.
Pubblicazione: (2024)
Bridging Coarse and Fine Recognition: A Hybrid Approach for Open-Ended Multi-Granularity Object Recognition in Interactive Educational Games
di: Yi, Hanling, et al.
Pubblicazione: (2026)
di: Yi, Hanling, et al.
Pubblicazione: (2026)
E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
di: Liu, Ye, et al.
Pubblicazione: (2024)
di: Liu, Ye, et al.
Pubblicazione: (2024)
Open-Vocabulary Object Detection via Neighboring Region Attention Alignment
di: Qiang, Sunyuan, et al.
Pubblicazione: (2024)
di: Qiang, Sunyuan, et al.
Pubblicazione: (2024)
MarvelOVD: Marrying Object Recognition and Vision-Language Models for Robust Open-Vocabulary Object Detection
di: Wang, Kuo, et al.
Pubblicazione: (2024)
di: Wang, Kuo, et al.
Pubblicazione: (2024)
VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
di: Lin, Zhiwei, et al.
Pubblicazione: (2025)
di: Lin, Zhiwei, et al.
Pubblicazione: (2025)
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments
di: Guruprasad, Pranav, et al.
Pubblicazione: (2025)
di: Guruprasad, Pranav, et al.
Pubblicazione: (2025)
Open-Vocabulary Object Detection via Language Hierarchy
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
di: Huang, Jiaxing, et al.
Pubblicazione: (2024)
DSAA: Dual-Stage Attribute Activation for Fine-grained Open Vocabulary Detection
di: Jiang, Donghong, et al.
Pubblicazione: (2026)
di: Jiang, Donghong, et al.
Pubblicazione: (2026)
Generative Refinement Networks for Visual Synthesis
di: Han, Jian, et al.
Pubblicazione: (2026)
di: Han, Jian, et al.
Pubblicazione: (2026)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
di: Ma, Chuofan, et al.
Pubblicazione: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection under Challenging Conditions
di: Cui, Luoping, et al.
Pubblicazione: (2025)
di: Cui, Luoping, et al.
Pubblicazione: (2025)
Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles
di: Liu, Jiawei, et al.
Pubblicazione: (2026)
di: Liu, Jiawei, et al.
Pubblicazione: (2026)
UniPCB: A Unified Vision-Language Benchmark for Open-Ended PCB Quality Inspection
di: Sun, Fuxiang, et al.
Pubblicazione: (2026)
di: Sun, Fuxiang, et al.
Pubblicazione: (2026)
Towards Open-Ended Visual Scientific Discovery with Sparse Autoencoders
di: Stevens, Samuel, et al.
Pubblicazione: (2025)
di: Stevens, Samuel, et al.
Pubblicazione: (2025)
Accelerate 3D Object Detection Models via Zero-Shot Attention Key Pruning
di: Xu, Lizhen, et al.
Pubblicazione: (2025)
di: Xu, Lizhen, et al.
Pubblicazione: (2025)
Towards Generalized Few-Shot Open-Set Object Detection
di: Su, Binyi, et al.
Pubblicazione: (2022)
di: Su, Binyi, et al.
Pubblicazione: (2022)
Task-Specific Zero-shot Quantization-Aware Training for Object Detection
di: Li, Changhao, et al.
Pubblicazione: (2025)
di: Li, Changhao, et al.
Pubblicazione: (2025)
Credible Teacher for Semi-Supervised Object Detection in Open Scene
di: Zhuang, Jingyu, et al.
Pubblicazione: (2024)
di: Zhuang, Jingyu, et al.
Pubblicazione: (2024)
InfinityStar: Unified Spacetime AutoRegressive Modeling for Visual Generation
di: Liu, Jinlai, et al.
Pubblicazione: (2025)
di: Liu, Jinlai, et al.
Pubblicazione: (2025)
SToLa: Self-Adaptive Touch-Language Framework with Tactile Commonsense Reasoning in Open-Ended Scenarios
di: Cheng, Ning, et al.
Pubblicazione: (2025)
di: Cheng, Ning, et al.
Pubblicazione: (2025)
LaMP: Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning
di: Li, Zhe, et al.
Pubblicazione: (2024)
di: Li, Zhe, et al.
Pubblicazione: (2024)
PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
di: Liu, Fengchun, et al.
Pubblicazione: (2025)
di: Liu, Fengchun, et al.
Pubblicazione: (2025)
LV-OSD: Language-Vision-Complementary Open-Set Object Detection
di: Zhang, Yupeng, et al.
Pubblicazione: (2026)
di: Zhang, Yupeng, et al.
Pubblicazione: (2026)
Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
di: Guo, Hao, et al.
Pubblicazione: (2025)
di: Guo, Hao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts
di: Lin, Zhiwei, et al.
Pubblicazione: (2024) -
Open-Det: An Efficient Learning Framework for Open-Ended Detection
di: Cao, Guiping, et al.
Pubblicazione: (2025) -
Waver: Wave Your Way to Lifelike Video Generation
di: Zhang, Yifu, et al.
Pubblicazione: (2025) -
Recognize Any Regions
di: Yang, Haosen, et al.
Pubblicazione: (2023) -
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
di: Chen, Xiuyuan, et al.
Pubblicazione: (2023)