PolyGen: Fully Synthetic Vision-Language Training via Multi-Generator Ensembles
Fuente:
arXiv
Saved in:
| Main Authors: | Brusini, Leonardo, Sbrolli, Cristian, Lomurno, Eugenio, Yamasaki, Toshihiko, Matteucci, Matteo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026)
by: Sbrolli, Cristian, et al.
Published: (2026)
Your Image Generator Is Your New Private Dataset
by: Resmini, Nicolo, et al.
Published: (2025)
by: Resmini, Nicolo, et al.
Published: (2025)
Neuro-Symbolic Scene Graph Conditioning for Synthetic Image Dataset Generation
by: Savazzi, Giacomo, et al.
Published: (2025)
by: Savazzi, Giacomo, et al.
Published: (2025)
Federated Knowledge Recycling: Privacy-Preserving Synthetic Data Sharing
by: Lomurno, Eugenio, et al.
Published: (2024)
by: Lomurno, Eugenio, et al.
Published: (2024)
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
by: Sbrolli, Cristian, et al.
Published: (2024)
by: Sbrolli, Cristian, et al.
Published: (2024)
Synthetic Image Learning: Preserving Performance and Preventing Membership Inference Attacks
by: Lomurno, Eugenio, et al.
Published: (2024)
by: Lomurno, Eugenio, et al.
Published: (2024)
SCENEFORGE: Enhancing 3D-text alignment with Structured Scene Compositions
by: Sbrolli, Cristian, et al.
Published: (2025)
by: Sbrolli, Cristian, et al.
Published: (2025)
Can Shape-Infused Joint Embeddings Improve Image-Conditioned 3D Diffusion?
by: Sbrolli, Cristian, et al.
Published: (2024)
by: Sbrolli, Cristian, et al.
Published: (2024)
Stable Diffusion Dataset Generation for Downstream Classification Tasks
by: Lomurno, Eugenio, et al.
Published: (2024)
by: Lomurno, Eugenio, et al.
Published: (2024)
POMONAG: Pareto-Optimal Many-Objective Neural Architecture Generator
by: Lomurno, Eugenio, et al.
Published: (2024)
by: Lomurno, Eugenio, et al.
Published: (2024)
Unified Vector Floorplan Generation via Markup Representation
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
by: Tanji, Naoto, et al.
Published: (2025)
by: Tanji, Naoto, et al.
Published: (2025)
A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision
by: Samele, Stefano, et al.
Published: (2026)
by: Samele, Stefano, et al.
Published: (2026)
CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
by: Kuroki, Michihiro, et al.
Published: (2025)
by: Kuroki, Michihiro, et al.
Published: (2025)
A Lightweight Neural Architecture Search Model for Medical Image Classification
by: Xie, Lunchen, et al.
Published: (2024)
by: Xie, Lunchen, et al.
Published: (2024)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
by: Zhao, Zaiying, et al.
Published: (2025)
by: Zhao, Zaiying, et al.
Published: (2025)
From Obstacles to Resources: Semi-supervised Learning Faces Synthetic Data Contamination
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
Language-guided Detection and Mitigation of Unknown Dataset Bias
by: Zhao, Zaiying, et al.
Published: (2024)
by: Zhao, Zaiying, et al.
Published: (2024)
BSED: Baseline Shapley-Based Explainable Detector
by: Kuroki, Michihiro, et al.
Published: (2023)
by: Kuroki, Michihiro, et al.
Published: (2023)
Face2Diffusion for Fast and Editable Face Personalization
by: Shiohara, Kaede, et al.
Published: (2024)
by: Shiohara, Kaede, et al.
Published: (2024)
Adversarial Training from Mean Field Perspective
by: Kumano, Soichiro, et al.
Published: (2025)
by: Kumano, Soichiro, et al.
Published: (2025)
ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
by: Okumura, Ryota, et al.
Published: (2025)
by: Okumura, Ryota, et al.
Published: (2025)
Dealing with Synthetic Data Contamination in Online Continual Learning
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
by: Wang, Zerun, et al.
Published: (2024)
by: Wang, Zerun, et al.
Published: (2024)
Act, Think or Abstain: Complexity-Aware Adaptive Inference for Vision-Language-Action Models
by: Izzo, Riccardo Andrea, et al.
Published: (2026)
by: Izzo, Riccardo Andrea, et al.
Published: (2026)
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
by: Xiao, Ling, et al.
Published: (2026)
by: Xiao, Ling, et al.
Published: (2026)
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
by: Xiao, Ling, et al.
Published: (2022)
by: Xiao, Ling, et al.
Published: (2022)
Language-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
by: Sugihara, Tomoya, et al.
Published: (2024)
by: Sugihara, Tomoya, et al.
Published: (2024)
ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors
by: Shiohara, Kaede, et al.
Published: (2026)
by: Shiohara, Kaede, et al.
Published: (2026)
Mirai: Autoregressive Visual Generation Needs Foresight
by: Yu, Yonghao, et al.
Published: (2026)
by: Yu, Yonghao, et al.
Published: (2026)
Improving Robustness of Vision-Language-Action Models by Restoring Corrupted Visual Inputs
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
by: Orjuela, Daniel Yezid Guarnizo, et al.
Published: (2026)
More than the Sum of Its Parts: Ensembling Backbone Networks for Few-Shot Segmentation
by: Catalano, Nico, et al.
Published: (2024)
by: Catalano, Nico, et al.
Published: (2024)
Reward Incremental Learning in Text-to-Image Generation
by: Wang, Maorong, et al.
Published: (2024)
by: Wang, Maorong, et al.
Published: (2024)
Theoretical Understanding of Learning from Adversarial Perturbations
by: Kumano, Soichiro, et al.
Published: (2024)
by: Kumano, Soichiro, et al.
Published: (2024)
Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
by: Kumano, Soichiro, et al.
Published: (2025)
by: Kumano, Soichiro, et al.
Published: (2025)
Wide Two-Layer Networks can Learn from Adversarial Perturbations
by: Kumano, Soichiro, et al.
Published: (2024)
by: Kumano, Soichiro, et al.
Published: (2024)
ZO-DARTS++: An Efficient and Size-Variable Zeroth-Order Neural Architecture Search Algorithm
by: Xie, Lunchen, et al.
Published: (2025)
by: Xie, Lunchen, et al.
Published: (2025)
Joint Fusion and Encoding: Advancing Multimodal Retrieval from the Ground Up
by: Huang, Lang, et al.
Published: (2025)
by: Huang, Lang, et al.
Published: (2025)
VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
by: Zhang, Jipeng, et al.
Published: (2025)
by: Zhang, Jipeng, et al.
Published: (2025)
InstaGen: Enhancing Object Detection by Training on Synthetic Dataset
by: Feng, Chengjian, et al.
Published: (2024)
by: Feng, Chengjian, et al.
Published: (2024)
Similar Items
-
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
by: Sbrolli, Cristian, et al.
Published: (2026) -
Your Image Generator Is Your New Private Dataset
by: Resmini, Nicolo, et al.
Published: (2025) -
Neuro-Symbolic Scene Graph Conditioning for Synthetic Image Dataset Generation
by: Savazzi, Giacomo, et al.
Published: (2025) -
Federated Knowledge Recycling: Privacy-Preserving Synthetic Data Sharing
by: Lomurno, Eugenio, et al.
Published: (2024) -
No Captions, No Problem: Captionless 3D-CLIP Alignment with Hard Negatives via CLIP Knowledge and LLMs
by: Sbrolli, Cristian, et al.
Published: (2024)