Controlled Training Data Generation with Diffusion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Yeo, Teresa, Atanov, Andrei, Benoit, Harold, Alekseev, Aleksandr, Ray, Ruchira, Akhoondi, Pooya Esmaeil, Zamir, Amir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
por: Atanov, Andrei, et al.
Publicado: (2024)
por: Atanov, Andrei, et al.
Publicado: (2024)
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
por: Salehi, Sogand, et al.
Publicado: (2024)
por: Salehi, Sogand, et al.
Publicado: (2024)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
por: Ramachandran, Rahul, et al.
Publicado: (2025)
por: Ramachandran, Rahul, et al.
Publicado: (2025)
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
por: Yang, Mingyue, et al.
Publicado: (2025)
por: Yang, Mingyue, et al.
Publicado: (2025)
AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection
por: Alamgeer, Sana, et al.
Publicado: (2025)
por: Alamgeer, Sana, et al.
Publicado: (2025)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
por: Atanov, Andrei, et al.
Publicado: (2026)
por: Atanov, Andrei, et al.
Publicado: (2026)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
por: Fayyazsanavi, Pooya, et al.
Publicado: (2024)
When Diffusion Breaks Constraints: Sequential Autoregressive Generation with RL and MCTS
por: Zhao, Zirui, et al.
Publicado: (2025)
por: Zhao, Zirui, et al.
Publicado: (2025)
Scalable Vision Language Model Training via High Quality Data Curation
por: Dong, Hongyuan, et al.
Publicado: (2025)
por: Dong, Hongyuan, et al.
Publicado: (2025)
Mitigating Exaggerated Safety in Large Language Models
por: Ray, Ruchira, et al.
Publicado: (2024)
por: Ray, Ruchira, et al.
Publicado: (2024)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
por: Cheng, Sheng, et al.
Publicado: (2024)
por: Cheng, Sheng, et al.
Publicado: (2024)
Train a Unified Multimodal Data Quality Classifier with Synthetic Data
por: Wang, Weizhi, et al.
Publicado: (2025)
por: Wang, Weizhi, et al.
Publicado: (2025)
Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
por: Lin, Han, et al.
Publicado: (2025)
por: Lin, Han, et al.
Publicado: (2025)
Effective Training Data Synthesis for Improving MLLM Chart Understanding
por: Yang, Yuwei, et al.
Publicado: (2025)
por: Yang, Yuwei, et al.
Publicado: (2025)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
por: Whitehead, Spencer, et al.
Publicado: (2024)
por: Whitehead, Spencer, et al.
Publicado: (2024)
Decoder-Only LLMs are Better Controllers for Diffusion Models
por: Dong, Ziyi, et al.
Publicado: (2025)
por: Dong, Ziyi, et al.
Publicado: (2025)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
por: Sun, Zhenhong, et al.
Publicado: (2024)
por: Sun, Zhenhong, et al.
Publicado: (2024)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
por: Wang, Chenglin, et al.
Publicado: (2026)
por: Wang, Chenglin, et al.
Publicado: (2026)
SliceWorld: A Predictive and Controllable World-State Model for CT Report Generation
por: Tian, Yuanhe, et al.
Publicado: (2026)
por: Tian, Yuanhe, et al.
Publicado: (2026)
Diffusion-RSCC: Diffusion Probabilistic Model for Change Captioning in Remote Sensing Images
por: Yu, Xiaofei, et al.
Publicado: (2024)
por: Yu, Xiaofei, et al.
Publicado: (2024)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
por: Simoncini, Walter, et al.
Publicado: (2024)
por: Simoncini, Walter, et al.
Publicado: (2024)
Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training
por: Zheng, Ruobing, et al.
Publicado: (2026)
por: Zheng, Ruobing, et al.
Publicado: (2026)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
por: Wang, Junyang, et al.
Publicado: (2026)
por: Wang, Junyang, et al.
Publicado: (2026)
Towards Efficient and Robust VQA-NLE Data Generation with Large Vision-Language Models
por: Irawan, Patrick Amadeus, et al.
Publicado: (2024)
por: Irawan, Patrick Amadeus, et al.
Publicado: (2024)
Can Prompt Modifiers Control Bias? A Comparative Analysis of Text-to-Image Generative Models
por: Shin, Philip Wootaek, et al.
Publicado: (2024)
por: Shin, Philip Wootaek, et al.
Publicado: (2024)
Taming the Tri-Space Tension: ARC-Guided Hallucination Modeling and Control for Text-to-Image Generation
por: Yang, Jianjiang, et al.
Publicado: (2025)
por: Yang, Jianjiang, et al.
Publicado: (2025)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
por: Zhao, Qinyu, et al.
Publicado: (2025)
por: Zhao, Qinyu, et al.
Publicado: (2025)
Deciphering Oracle Bone Language with Diffusion Models
por: Guan, Haisu, et al.
Publicado: (2024)
por: Guan, Haisu, et al.
Publicado: (2024)
Scaling Concept With Text-Guided Diffusion Models
por: Huang, Chao, et al.
Publicado: (2024)
por: Huang, Chao, et al.
Publicado: (2024)
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
por: Yeo, Wei Jie, et al.
Publicado: (2025)
por: Yeo, Wei Jie, et al.
Publicado: (2025)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
por: Castro, Santiago, et al.
Publicado: (2024)
por: Castro, Santiago, et al.
Publicado: (2024)
Survey of Video Diffusion Models: Foundations, Implementations, and Applications
por: Wang, Yimu, et al.
Publicado: (2025)
por: Wang, Yimu, et al.
Publicado: (2025)
Why Instruction-Based Unlearning Fails in Diffusion Models?
por: Zhang, Zeliang, et al.
Publicado: (2026)
por: Zhang, Zeliang, et al.
Publicado: (2026)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
por: Su, Xin, et al.
Publicado: (2024)
por: Su, Xin, et al.
Publicado: (2024)
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
por: Luo, Run, et al.
Publicado: (2024)
por: Luo, Run, et al.
Publicado: (2024)
Mojito: Motion Trajectory and Intensity Control for Video Generation
por: He, Xuehai, et al.
Publicado: (2024)
por: He, Xuehai, et al.
Publicado: (2024)
MedVisionLlama: Leveraging Pre-Trained Large Language Model Layers to Enhance Medical Image Segmentation
por: Kumar, Gurucharan Marthi Krishna, et al.
Publicado: (2024)
por: Kumar, Gurucharan Marthi Krishna, et al.
Publicado: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
por: Huang, Jia-Hong, et al.
Publicado: (2024)
por: Huang, Jia-Hong, et al.
Publicado: (2024)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
por: Wu, Xueqing, et al.
Publicado: (2026)
por: Wu, Xueqing, et al.
Publicado: (2026)
Enhancing Large Vision Language Models with Self-Training on Image Comprehension
por: Deng, Yihe, et al.
Publicado: (2024)
por: Deng, Yihe, et al.
Publicado: (2024)
Ejemplares similares
-
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
por: Atanov, Andrei, et al.
Publicado: (2024) -
ViPer: Visual Personalization of Generative Models via Individual Preference Learning
por: Salehi, Sogand, et al.
Publicado: (2024) -
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
por: Ramachandran, Rahul, et al.
Publicado: (2025) -
CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation
por: Yang, Mingyue, et al.
Publicado: (2025) -
AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection
por: Alamgeer, Sana, et al.
Publicado: (2025)