DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Baresi, Luciano, Hu, Davide Yi Xian, Mas'udi, Muhammad Irfan, Quattrocchi, Giovanni |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
por: Baresi, Luciano, et al.
Publicado: (2024)
por: Baresi, Luciano, et al.
Publicado: (2024)
Can LLMs Generate User Stories and Assess Their Quality?
por: Quattrocchi, Giovanni, et al.
Publicado: (2025)
por: Quattrocchi, Giovanni, et al.
Publicado: (2025)
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
por: Zhao, Guanghui, et al.
Publicado: (2026)
por: Zhao, Guanghui, et al.
Publicado: (2026)
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
por: Allison, Christopher John, et al.
Publicado: (2024)
por: Allison, Christopher John, et al.
Publicado: (2024)
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
por: Jin, Duosi, et al.
Publicado: (2026)
por: Jin, Duosi, et al.
Publicado: (2026)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
por: Oketunji, Abiodun Finbarrs
Publicado: (2024)
Autark: A Serverless Toolkit for Prototyping Urban Visual Analytics Systems
por: Alexandre, Lucas, et al.
Publicado: (2026)
por: Alexandre, Lucas, et al.
Publicado: (2026)
TARIPlay: A Test Framework for AR Applications based on Interactive Area Tracking in Playback Videos
por: Mousavi, Seyed Amir, et al.
Publicado: (2026)
por: Mousavi, Seyed Amir, et al.
Publicado: (2026)
Code Arcades: 3d Visualization of Classes, Dependencies and Software Metrics
por: Savidis, Anthony, et al.
Publicado: (2025)
por: Savidis, Anthony, et al.
Publicado: (2025)
The Vibe-Check Protocol: Quantifying Cognitive Offloading in AI Programming
por: Aiersilan, Aizierjiang
Publicado: (2026)
por: Aiersilan, Aizierjiang
Publicado: (2026)
Investigating Traffic Accident Detection Using Multimodal Large Language Models
por: Skender, Ilhan, et al.
Publicado: (2025)
por: Skender, Ilhan, et al.
Publicado: (2025)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
por: Liang, Shanchao, et al.
Publicado: (2024)
por: Liang, Shanchao, et al.
Publicado: (2024)
A Plausibility Study of Using Augmented Reality in the Ventriculoperitoneal Shunt Operations
por: Dorji, Tandin, et al.
Publicado: (2024)
por: Dorji, Tandin, et al.
Publicado: (2024)
Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
por: Baresi, Luciano, et al.
Publicado: (2026)
por: Baresi, Luciano, et al.
Publicado: (2026)
SMooGPT: Stylized Motion Generation using Large Language Models
por: Zhong, Lei, et al.
Publicado: (2025)
por: Zhong, Lei, et al.
Publicado: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
por: Yin, Zixin, et al.
Publicado: (2025)
por: Yin, Zixin, et al.
Publicado: (2025)
TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
por: Lu, You, et al.
Publicado: (2025)
por: Lu, You, et al.
Publicado: (2025)
A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
por: Khalid, Waleed, et al.
Publicado: (2025)
por: Khalid, Waleed, et al.
Publicado: (2025)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
por: Wu, Ronghuan, et al.
Publicado: (2024)
por: Wu, Ronghuan, et al.
Publicado: (2024)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
por: Li, Kaixin, et al.
Publicado: (2024)
por: Li, Kaixin, et al.
Publicado: (2024)
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
por: Wang, Kaixuan, et al.
Publicado: (2026)
por: Wang, Kaixuan, et al.
Publicado: (2026)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
por: Meng, Fanqing, et al.
Publicado: (2026)
por: Meng, Fanqing, et al.
Publicado: (2026)
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
por: Zhao, Qingcheng, et al.
Publicado: (2024)
por: Zhao, Qingcheng, et al.
Publicado: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
por: Para, Wamiq Reyaz, et al.
Publicado: (2024)
por: Para, Wamiq Reyaz, et al.
Publicado: (2024)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
por: Fang, Shuangkang, et al.
Publicado: (2025)
por: Fang, Shuangkang, et al.
Publicado: (2025)
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
por: Tang, Jiapeng, et al.
Publicado: (2024)
por: Tang, Jiapeng, et al.
Publicado: (2024)
Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation
por: Ng, Kam Woh, et al.
Publicado: (2025)
por: Ng, Kam Woh, et al.
Publicado: (2025)
Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs
por: Scaramuzza, Filippo, et al.
Publicado: (2025)
por: Scaramuzza, Filippo, et al.
Publicado: (2025)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
por: Deng, Youming, et al.
Publicado: (2025)
por: Deng, Youming, et al.
Publicado: (2025)
Shap-MeD
por: Laverde, Nicolás, et al.
Publicado: (2025)
por: Laverde, Nicolás, et al.
Publicado: (2025)
ROMAN: Reward-Orchestrated Multi-Head Attention Network for Autonomous Driving System Testing
por: Chi, Jianlei, et al.
Publicado: (2026)
por: Chi, Jianlei, et al.
Publicado: (2026)
GUing: A Mobile GUI Search Engine using a Vision-Language Model
por: Wei, Jialiang, et al.
Publicado: (2024)
por: Wei, Jialiang, et al.
Publicado: (2024)
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
por: Li, Xinyang, et al.
Publicado: (2025)
por: Li, Xinyang, et al.
Publicado: (2025)
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
por: Wei, Guangshun, et al.
Publicado: (2024)
por: Wei, Guangshun, et al.
Publicado: (2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
por: Kuang, Zhengfei, et al.
Publicado: (2024)
por: Kuang, Zhengfei, et al.
Publicado: (2024)
Large-Scale Multi-Character Interaction Synthesis
por: Chang, Ziyi, et al.
Publicado: (2025)
por: Chang, Ziyi, et al.
Publicado: (2025)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
por: Zhong, Jianping, et al.
Publicado: (2026)
por: Zhong, Jianping, et al.
Publicado: (2026)
Multisensory extended reality applications offer benefits for volumetric biomedical image analysis in research and medicine
por: Krieger, Kathrin, et al.
Publicado: (2023)
por: Krieger, Kathrin, et al.
Publicado: (2023)
DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer
por: Huang, Nisha, et al.
Publicado: (2026)
por: Huang, Nisha, et al.
Publicado: (2026)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
por: Chen, Hansheng, et al.
Publicado: (2024)
por: Chen, Hansheng, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
por: Baresi, Luciano, et al.
Publicado: (2024) -
Can LLMs Generate User Stories and Assess Their Quality?
por: Quattrocchi, Giovanni, et al.
Publicado: (2025) -
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
por: Zhao, Guanghui, et al.
Publicado: (2026) -
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
por: Allison, Christopher John, et al.
Publicado: (2024) -
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
por: Jin, Duosi, et al.
Publicado: (2026)