InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Chengjian, Zhong, Yujie, Jie, Zequn, Xie, Weidi, Ma, Lin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
di: Zeng, Yingsen, et al.
Pubblicazione: (2024)
di: Zeng, Yingsen, et al.
Pubblicazione: (2024)
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
di: Huang, Zhijian, et al.
Pubblicazione: (2024)
di: Huang, Zhijian, et al.
Pubblicazione: (2024)
DisTime: Distribution-based Time Representation for Video Large Language Models
di: Zeng, Yingsen, et al.
Pubblicazione: (2025)
di: Zeng, Yingsen, et al.
Pubblicazione: (2025)
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
di: Xiao, Baihui, et al.
Pubblicazione: (2025)
di: Xiao, Baihui, et al.
Pubblicazione: (2025)
Matten: Video Generation with Mamba-Attention
di: Gao, Yu, et al.
Pubblicazione: (2024)
di: Gao, Yu, et al.
Pubblicazione: (2024)
MRStyle: A Unified Framework for Color Style Transfer with Multi-Modality Reference
di: Huang, Jiancheng, et al.
Pubblicazione: (2024)
di: Huang, Jiancheng, et al.
Pubblicazione: (2024)
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis
di: Chen, Lei, et al.
Pubblicazione: (2024)
di: Chen, Lei, et al.
Pubblicazione: (2024)
OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
RFSR: Improving ISR Diffusion Models via Reward Feedback Learning
di: Sun, Xiaopeng, et al.
Pubblicazione: (2024)
di: Sun, Xiaopeng, et al.
Pubblicazione: (2024)
Appearance-Based Refinement for Object-Centric Motion Segmentation
di: Xie, Junyu, et al.
Pubblicazione: (2023)
di: Xie, Junyu, et al.
Pubblicazione: (2023)
InstructVEdit: A Holistic Approach for Instructional Video Editing
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
di: Chen, Shaoxiang, et al.
Pubblicazione: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
di: Lan, Xiaohan, et al.
Pubblicazione: (2024)
Aerial Monocular 3D Object Detection
di: Hu, Yue, et al.
Pubblicazione: (2022)
di: Hu, Yue, et al.
Pubblicazione: (2022)
AP-CAP: Advancing High-Quality Data Synthesis for Animal Pose Estimation via a Controllable Image Generation Pipeline
di: Wang, Lei, et al.
Pubblicazione: (2025)
di: Wang, Lei, et al.
Pubblicazione: (2025)
CONQUER: Context-Aware Representation with Query Enhancement for Text-Based Person Search
di: Xie, Zequn
Pubblicazione: (2026)
di: Xie, Zequn
Pubblicazione: (2026)
Moving Object Segmentation: All You Need Is SAM (and Flow)
di: Xie, Junyu, et al.
Pubblicazione: (2024)
di: Xie, Junyu, et al.
Pubblicazione: (2024)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
di: Xie, Junyu, et al.
Pubblicazione: (2026)
di: Xie, Junyu, et al.
Pubblicazione: (2026)
Synthetic-to-Real Camouflaged Object Detection
di: Luo, Zhihao, et al.
Pubblicazione: (2025)
di: Luo, Zhihao, et al.
Pubblicazione: (2025)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
di: Liu, Chang, et al.
Pubblicazione: (2023)
di: Liu, Chang, et al.
Pubblicazione: (2023)
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
di: Liang, Zhanhao, et al.
Pubblicazione: (2026)
di: Liang, Zhanhao, et al.
Pubblicazione: (2026)
X-SAM: From Segment Anything to Any Segmentation
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
Track-On2: Enhancing Online Point Tracking with Memory
di: Aydemir, Görkay, et al.
Pubblicazione: (2025)
di: Aydemir, Görkay, et al.
Pubblicazione: (2025)
Every Dataset Counts: Scaling up Monocular 3D Object Detection with Joint Datasets Training
di: Ma, Fulong, et al.
Pubblicazione: (2023)
di: Ma, Fulong, et al.
Pubblicazione: (2023)
Enhancing Robustness of Vision-Language Models through Orthogonality Learning and Self-Regularization
di: Li, Jinlong, et al.
Pubblicazione: (2024)
di: Li, Jinlong, et al.
Pubblicazione: (2024)
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
di: Zhu, Yixuan, et al.
Pubblicazione: (2025)
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
di: Meng, Yanxu, et al.
Pubblicazione: (2025)
di: Meng, Yanxu, et al.
Pubblicazione: (2025)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
di: Ma, Lichen, et al.
Pubblicazione: (2024)
di: Ma, Lichen, et al.
Pubblicazione: (2024)
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
di: You, Junqi, et al.
Pubblicazione: (2025)
di: You, Junqi, et al.
Pubblicazione: (2025)
EchoSight: Advancing Visual-Language Models with Wiki Knowledge
di: Yan, Yibin, et al.
Pubblicazione: (2024)
di: Yan, Yibin, et al.
Pubblicazione: (2024)
Grounded Question-Answering in Long Egocentric Videos
di: Di, Shangzhe, et al.
Pubblicazione: (2023)
di: Di, Shangzhe, et al.
Pubblicazione: (2023)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
di: Xu, Jilan, et al.
Pubblicazione: (2025)
di: Xu, Jilan, et al.
Pubblicazione: (2025)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
di: Ge, Mingji, et al.
Pubblicazione: (2026)
di: Ge, Mingji, et al.
Pubblicazione: (2026)
Mr. DETR++: Instructive Multi-Route Training for Detection Transformers with Mixture-of-Experts
di: Zhang, Chang-Bin, et al.
Pubblicazione: (2024)
di: Zhang, Chang-Bin, et al.
Pubblicazione: (2024)
Salient Object Detection in Traffic Scene through the TSOD10K Dataset
di: Qiu, Yu, et al.
Pubblicazione: (2025)
di: Qiu, Yu, et al.
Pubblicazione: (2025)
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation
di: Shi, Yudi, et al.
Pubblicazione: (2024)
di: Shi, Yudi, et al.
Pubblicazione: (2024)
AeroGen: Enhancing Remote Sensing Object Detection with Diffusion-Driven Data Generation
di: Tang, Datao, et al.
Pubblicazione: (2024)
di: Tang, Datao, et al.
Pubblicazione: (2024)
InstaScene: Towards Complete 3D Instance Decomposition and Reconstruction from Cluttered Scenes
di: Yang, Zesong, et al.
Pubblicazione: (2025)
di: Yang, Zesong, et al.
Pubblicazione: (2025)
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
di: Leonardi, Rosario, et al.
Pubblicazione: (2026)
di: Leonardi, Rosario, et al.
Pubblicazione: (2026)
InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
di: Hou, Xianbao, et al.
Pubblicazione: (2025)
di: Hou, Xianbao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection
di: Zeng, Yingsen, et al.
Pubblicazione: (2024) -
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
di: Huang, Zhijian, et al.
Pubblicazione: (2024) -
DisTime: Distribution-based Time Representation for Video Large Language Models
di: Zeng, Yingsen, et al.
Pubblicazione: (2025) -
RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
di: Xiao, Baihui, et al.
Pubblicazione: (2025) -
Matten: Video Generation with Mamba-Attention
di: Gao, Yu, et al.
Pubblicazione: (2024)