Generating Enhanced Negatives for Training Language-Based Object Detectors
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Shiyu, Zhao, Long, G, Vijay Kumar B., Suh, Yumin, Metaxas, Dimitris N., Chandraker, Manmohan, Schulter, Samuel |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Self-Training for Open-Vocabulary Object Detection
by: Zhao, Shiyu, et al.
Published: (2023)
by: Zhao, Shiyu, et al.
Published: (2023)
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024)
by: Aich, Abhishek, et al.
Published: (2024)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
Instantaneous Perception of Moving Objects in 3D
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
by: Liang, Mingfu, et al.
Published: (2024)
by: Liang, Mingfu, et al.
Published: (2024)
LLM-Assist: Enhancing Closed-Loop Planning with Language-Based Reasoning
by: Sharan, S P, et al.
Published: (2023)
by: Sharan, S P, et al.
Published: (2023)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
by: Ko, Dohwan, et al.
Published: (2025)
by: Ko, Dohwan, et al.
Published: (2025)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
by: Zhangli, Qilong, et al.
Published: (2024)
by: Zhangli, Qilong, et al.
Published: (2024)
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
by: Yao, Manyi, et al.
Published: (2024)
by: Yao, Manyi, et al.
Published: (2024)
Tuned Contrastive Learning
by: Animesh, Chaitanya, et al.
Published: (2023)
by: Animesh, Chaitanya, et al.
Published: (2023)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
by: Li, Haodong, et al.
Published: (2026)
by: Li, Haodong, et al.
Published: (2026)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
Locally Orderless Images for Optimization in Differentiable Rendering
by: Mehta, Ishit, et al.
Published: (2025)
by: Mehta, Ishit, et al.
Published: (2025)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
by: Aich, Abhishek, et al.
Published: (2026)
by: Aich, Abhishek, et al.
Published: (2026)
UDA-Bench: Revisiting Common Assumptions in Unsupervised Domain Adaptation Using a Standardized Framework
by: Kalluri, Tarun, et al.
Published: (2024)
by: Kalluri, Tarun, et al.
Published: (2024)
PhyCo: Learning Controllable Physical Priors for Generative Motion
by: Narayanan, Sriram, et al.
Published: (2026)
by: Narayanan, Sriram, et al.
Published: (2026)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
by: Ghosh, Anurag, et al.
Published: (2026)
by: Ghosh, Anurag, et al.
Published: (2026)
Test-Time Spectrum-Aware Latent Steering for Zero-Shot Generalization in Vision-Language Models
by: Dafnis, Konstantinos M., et al.
Published: (2025)
by: Dafnis, Konstantinos M., et al.
Published: (2025)
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
by: Dao, Quan, et al.
Published: (2026)
by: Dao, Quan, et al.
Published: (2026)
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
by: He, Yun, et al.
Published: (2025)
by: He, Yun, et al.
Published: (2025)
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
by: Jain, Seemandhar, et al.
Published: (2026)
by: Jain, Seemandhar, et al.
Published: (2026)
InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
by: Xu, Sirui, et al.
Published: (2026)
by: Xu, Sirui, et al.
Published: (2026)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2023)
Stable Signer: Hierarchical Sign Language Generative Model
by: Fang, Sen, et al.
Published: (2025)
by: Fang, Sen, et al.
Published: (2025)
GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detection
by: Malić, Dušan, et al.
Published: (2025)
by: Malić, Dušan, et al.
Published: (2025)
LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model Adaptation
by: Jin, Can, et al.
Published: (2025)
by: Jin, Can, et al.
Published: (2025)
Steering Rectified Flow Models in the Vector Field for Controlled Image Generation
by: Patel, Maitreya, et al.
Published: (2024)
by: Patel, Maitreya, et al.
Published: (2024)
Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
by: Zhao, Shiyu, et al.
Published: (2024)
by: Zhao, Shiyu, et al.
Published: (2024)
LidaRF: Delving into Lidar for Neural Radiance Field on Street Scenes
by: Sun, Shanlin, et al.
Published: (2024)
by: Sun, Shanlin, et al.
Published: (2024)
HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles
by: Wang, Yifan, et al.
Published: (2026)
by: Wang, Yifan, et al.
Published: (2026)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
by: Lin, Chuang, et al.
Published: (2024)
by: Lin, Chuang, et al.
Published: (2024)
MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition
by: Zhao, Mingyu, et al.
Published: (2025)
by: Zhao, Mingyu, et al.
Published: (2025)
Temporal Dynamics Enhancer for Directly Trained Spiking Object Detectors
by: Luo, Fan, et al.
Published: (2025)
by: Luo, Fan, et al.
Published: (2025)
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
Score-Guided Diffusion for 3D Human Recovery
by: Stathopoulos, Anastasis, et al.
Published: (2024)
by: Stathopoulos, Anastasis, et al.
Published: (2024)
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
by: Yao, Manyi, et al.
Published: (2025)
by: Yao, Manyi, et al.
Published: (2025)
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
by: Fang, Sen, et al.
Published: (2026)
by: Fang, Sen, et al.
Published: (2026)
RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment
by: Gu, Difei, et al.
Published: (2025)
by: Gu, Difei, et al.
Published: (2025)
Similar Items
-
Taming Self-Training for Open-Vocabulary Object Detection
by: Zhao, Shiyu, et al.
Published: (2023) -
Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
by: Aich, Abhishek, et al.
Published: (2024) -
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
by: Khan, Zaid, et al.
Published: (2024) -
Instantaneous Perception of Moving Objects in 3D
by: Liu, Di, et al.
Published: (2024) -
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
by: Liang, Mingfu, et al.
Published: (2024)