Data Factory with Minimal Human Effort Using VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ye, Jiaojiao, Zhong, Jiaxing, Xie, Qian, Zhou, Yuzhou, Trigoni, Niki, Markham, Andrew |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
von: Zhou, Kaichen, et al.
Veröffentlicht: (2023)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2023)
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
von: Wang, Jialu, et al.
Veröffentlicht: (2024)
von: Wang, Jialu, et al.
Veröffentlicht: (2024)
MambaLoc: Efficient Camera Localisation via State Space Model
von: Wang, Jialu, et al.
Veröffentlicht: (2024)
von: Wang, Jialu, et al.
Veröffentlicht: (2024)
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
von: Shin, Sangyun, et al.
Veröffentlicht: (2023)
von: Shin, Sangyun, et al.
Veröffentlicht: (2023)
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)
ZeST: Zero-Shot Material Transfer from a Single Image
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
DynPoint: Dynamic Neural Point For View Synthesis
von: Zhou, Kaichen, et al.
Veröffentlicht: (2023)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2023)
Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models
von: Vankadari, Madhu, et al.
Veröffentlicht: (2024)
von: Vankadari, Madhu, et al.
Veröffentlicht: (2024)
SoundLoc3D: Invisible 3D Sound Source Localization and Classification Using a Multimodal RGB-D Acoustic Camera
von: He, Yuhang, et al.
Veröffentlicht: (2024)
von: He, Yuhang, et al.
Veröffentlicht: (2024)
Learning Continuous 3D Words for Text-to-Image Generation
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
von: Cheng, Ta-Ying, et al.
Veröffentlicht: (2024)
VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization
von: Zhou, Kaichen, et al.
Veröffentlicht: (2020)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2020)
Towards Multi-Modal Animal Pose Estimation: A Survey and In-Depth Analysis
von: Deng, Qianyi, et al.
Veröffentlicht: (2024)
von: Deng, Qianyi, et al.
Veröffentlicht: (2024)
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2024)
von: Yeh, Chun-Hsiao, et al.
Veröffentlicht: (2024)
Graph Convolutional Long Short-Term Memory Attention Network for Post-Stroke Compensatory Movement Detection Based on Skeleton Data
von: Fan, Jiaxing, et al.
Veröffentlicht: (2025)
von: Fan, Jiaxing, et al.
Veröffentlicht: (2025)
LLMs Meet VLMs: Boost Open Vocabulary Object Detection with Fine-grained Descriptors
von: Jin, Sheng, et al.
Veröffentlicht: (2024)
von: Jin, Sheng, et al.
Veröffentlicht: (2024)
Constructing Concept-based Models to Mitigate Spurious Correlations with Minimal Human Effort
von: Kim, Jeeyung, et al.
Veröffentlicht: (2024)
von: Kim, Jeeyung, et al.
Veröffentlicht: (2024)
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
von: Toubal, Imad Eddine, et al.
Veröffentlicht: (2024)
von: Toubal, Imad Eddine, et al.
Veröffentlicht: (2024)
CountGD++: Generalized Prompting for Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
von: Chatzichristodoulou, Georgios, et al.
Veröffentlicht: (2025)
von: Chatzichristodoulou, Georgios, et al.
Veröffentlicht: (2025)
CityLLaVA: Efficient Fine-Tuning for VLMs in City Scenario
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
von: Duan, Zhizhao, et al.
Veröffentlicht: (2024)
Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
VisualActBench: Can VLMs See and Act like a Human?
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
Understanding the Cross-Domain Capabilities of Video-Based Few-Shot Action Recognition Models
von: Markham, Georgia, et al.
Veröffentlicht: (2024)
von: Markham, Georgia, et al.
Veröffentlicht: (2024)
S2D: Sparse to Dense Lifting for 3D Reconstruction with Minimal Inputs
von: Ji, Yuzhou, et al.
Veröffentlicht: (2026)
von: Ji, Yuzhou, et al.
Veröffentlicht: (2026)
Should VLMs be Pre-trained with Image Data?
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
von: Keh, Sedrick, et al.
Veröffentlicht: (2025)
Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation
von: Qian, Haotian, et al.
Veröffentlicht: (2024)
von: Qian, Haotian, et al.
Veröffentlicht: (2024)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
von: Yang, Yuchen, et al.
Veröffentlicht: (2026)
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
von: Dong, Sixun, et al.
Veröffentlicht: (2025)
von: Dong, Sixun, et al.
Veröffentlicht: (2025)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
CountGD: Multi-Modal Open-World Counting
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
SynCellFactory: Generative Data Augmentation for Cell Tracking
von: Sturm, Moritz, et al.
Veröffentlicht: (2024)
von: Sturm, Moritz, et al.
Veröffentlicht: (2024)
Are VLMs Ready for Autonomous Driving? An Empirical Study from the Reliability, Data, and Metric Perspectives
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2025)
von: Xie, Shaoyuan, et al.
Veröffentlicht: (2025)
A Hitchhikers Guide to Fine-Grained Face Forgery Detection Using Common Sense Reasoning
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2024)
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2024)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
Open-World Object Counting in Videos
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
von: Amini-Naieni, Niki, et al.
Veröffentlicht: (2025)
FMLGS: Fast Multilevel Language Embedded Gaussians for Part-level Interactive Agents
von: Tan, Xin, et al.
Veröffentlicht: (2025)
von: Tan, Xin, et al.
Veröffentlicht: (2025)
Fact-checking based fake news detection: a review
von: Yang, Yuzhou, et al.
Veröffentlicht: (2024)
von: Yang, Yuzhou, et al.
Veröffentlicht: (2024)
EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion
von: Siy, Joshua, et al.
Veröffentlicht: (2026)
von: Siy, Joshua, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes
von: Zhou, Kaichen, et al.
Veröffentlicht: (2023) -
WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization
von: Wang, Jialu, et al.
Veröffentlicht: (2024) -
MambaLoc: Efficient Camera Localisation via State Space Model
von: Wang, Jialu, et al.
Veröffentlicht: (2024) -
Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation
von: Shin, Sangyun, et al.
Veröffentlicht: (2023) -
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
von: Ma, Chenyang, et al.
Veröffentlicht: (2024)