Data Factory with Minimal Human Effort Using VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Jiaojiao, Zhong, Jiaxing, Xie, Qian, Zhou, Yuzhou, Trigoni, Niki, Markham, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024)
by: Wang, Jialu, et al.
Published: (2024)
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
by: Shin, Sangyun, et al.
Published: (2023)
by: Shin, Sangyun, et al.
Published: (2023)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)
by: Ma, Chenyang, et al.
Published: (2024)
ZeST: Zero-Shot Material Transfer from a Single Image
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
DynPoint: Dynamic Neural Point For View Synthesis
by: Zhou, Kaichen, et al.
Published: (2023)
by: Zhou, Kaichen, et al.
Published: (2023)
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
by: Vankadari, Madhu, et al.
Published: (2024)
by: Vankadari, Madhu, et al.
Published: (2024)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
by: He, Yuhang, et al.
Published: (2024)
by: He, Yuhang, et al.
Published: (2024)
Learning Continuous 3D Words for Text-to-Image Generation
by: Cheng, Ta-Ying, et al.
Published: (2024)
by: Cheng, Ta-Ying, et al.
Published: (2024)
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
by: Zhou, Kaichen, et al.
Published: (2020)
by: Zhou, Kaichen, et al.
Published: (2020)
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
by: Deng, Qianyi, et al.
Published: (2024)
by: Deng, Qianyi, et al.
Published: (2024)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
by: Yeh, Chun-Hsiao, et al.
Published: (2024)
Graph Convolutional Long Short-Term Memory Attention Network for Post-Stroke Compensatory Movement Detection Based on Skeleton Data
by: Fan, Jiaxing, et al.
Published: (2025)
by: Fan, Jiaxing, et al.
Published: (2025)
LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
by: Jin, Sheng, et al.
Published: (2024)
by: Jin, Sheng, et al.
Published: (2024)
Constructing Concept-based Models to Mitigate Spurious Correlations with Minimal Human Effort
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
by: Toubal, Imad Eddine, et al.
Published: (2024)
by: Toubal, Imad Eddine, et al.
Published: (2024)
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
by: Chatzichristodoulou, Georgios, et al.
Published: (2025)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
by: Duan, Zhizhao, et al.
Published: (2024)
by: Duan, Zhizhao, et al.
Published: (2024)
Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
by: Xu, Mingjie, et al.
Published: (2025)
by: Xu, Mingjie, et al.
Published: (2025)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Understanding the Cross-Domain Capabilities of Video-Based Few-Shot Action Recognition Models
by: Markham, Georgia, et al.
Published: (2024)
by: Markham, Georgia, et al.
Published: (2024)
S2D: Sparse to Dense Lifting for 3D Reconstruction with Minimal Inputs
by: Ji, Yuzhou, et al.
Published: (2026)
by: Ji, Yuzhou, et al.
Published: (2026)
Should VLMs be Pre-trained with Image Data?
by: Keh, Sedrick, et al.
Published: (2025)
by: Keh, Sedrick, et al.
Published: (2025)
Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation
by: Qian, Haotian, et al.
Published: (2024)
by: Qian, Haotian, et al.
Published: (2024)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
by: Yang, Yuchen, et al.
Published: (2026)
by: Yang, Yuchen, et al.
Published: (2026)
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
by: Chinchure, Aditya, et al.
Published: (2025)
by: Chinchure, Aditya, et al.
Published: (2025)
CountGD: Multi-Modal Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2024)
by: Amini-Naieni, Niki, et al.
Published: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
SynCellFactory: Generative Data Augmentation for Cell Tracking
by: Sturm, Moritz, et al.
Published: (2024)
by: Sturm, Moritz, et al.
Published: (2024)
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
by: Xie, Shaoyuan, et al.
Published: (2025)
by: Xie, Shaoyuan, et al.
Published: (2025)
A Hitchhikers Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning
by: Foteinopoulou, Niki Maria, et al.
Published: (2024)
by: Foteinopoulou, Niki Maria, et al.
Published: (2024)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents
by: Tan, Xin, et al.
Published: (2025)
by: Tan, Xin, et al.
Published: (2025)
Fact-checking based fake news detection: a review
by: Yang, Yuzhou, et al.
Published: (2024)
by: Yang, Yuzhou, et al.
Published: (2024)
EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion
by: Siy, Joshua, et al.
Published: (2026)
by: Siy, Joshua, et al.
Published: (2026)
Similar Items
-
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
by: Zhou, Kaichen, et al.
Published: (2023) -
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
by: Wang, Jialu, et al.
Published: (2024) -
MambaLoc: Efficient Camera Localisation via State Space Model
by: Wang, Jialu, et al.
Published: (2024) -
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
by: Shin, Sangyun, et al.
Published: (2023) -
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
by: Ma, Chenyang, et al.
Published: (2024)