DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Baresi, Luciano, Hu, Davide Yi Xian, Mas'udi, Muhammad Irfan, Quattrocchi, Giovanni |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
by: Baresi, Luciano, et al.
Published: (2024)
by: Baresi, Luciano, et al.
Published: (2024)
Can LLMs Generate User Stories and Assess Their Quality?
by: Quattrocchi, Giovanni, et al.
Published: (2025)
by: Quattrocchi, Giovanni, et al.
Published: (2025)
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
by: Zhao, Guanghui, et al.
Published: (2026)
by: Zhao, Guanghui, et al.
Published: (2026)
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
by: Allison, Christopher John, et al.
Published: (2024)
by: Allison, Christopher John, et al.
Published: (2024)
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
by: Jin, Duosi, et al.
Published: (2026)
by: Jin, Duosi, et al.
Published: (2026)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Autark: A Serverless Toolkit for Prototyping Urban Visual Analytics Systems
by: Alexandre, Lucas, et al.
Published: (2026)
by: Alexandre, Lucas, et al.
Published: (2026)
TARIPlay: A Test Framework for AR Applications based on Interactive Area Tracking in Playback Videos
by: Mousavi, Seyed Amir, et al.
Published: (2026)
by: Mousavi, Seyed Amir, et al.
Published: (2026)
Code Arcades: 3d Visualization of Classes, Dependencies and Software Metrics
by: Savidis, Anthony, et al.
Published: (2025)
by: Savidis, Anthony, et al.
Published: (2025)
The Vibe-Check Protocol: Quantifying Cognitive Offloading in AI Programming
by: Aiersilan, Aizierjiang
Published: (2026)
by: Aiersilan, Aizierjiang
Published: (2026)
Investigating Traffic Accident Detection Using Multimodal Large Language Models
by: Skender, Ilhan, et al.
Published: (2025)
by: Skender, Ilhan, et al.
Published: (2025)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
by: Liang, Shanchao, et al.
Published: (2024)
by: Liang, Shanchao, et al.
Published: (2024)
A Plausibility Study of Using Augmented Reality in the Ventriculoperitoneal Shunt Operations
by: Dorji, Tandin, et al.
Published: (2024)
by: Dorji, Tandin, et al.
Published: (2024)
Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
by: Baresi, Luciano, et al.
Published: (2026)
by: Baresi, Luciano, et al.
Published: (2026)
SMooGPT: Stylized Motion Generation using Large Language Models
by: Zhong, Lei, et al.
Published: (2025)
by: Zhong, Lei, et al.
Published: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
by: Yin, Zixin, et al.
Published: (2025)
by: Yin, Zixin, et al.
Published: (2025)
TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
by: Lu, You, et al.
Published: (2025)
by: Lu, You, et al.
Published: (2025)
A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
by: Khalid, Waleed, et al.
Published: (2025)
by: Khalid, Waleed, et al.
Published: (2025)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
by: Wu, Ronghuan, et al.
Published: (2024)
by: Wu, Ronghuan, et al.
Published: (2024)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
by: Li, Kaixin, et al.
Published: (2024)
by: Li, Kaixin, et al.
Published: (2024)
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
by: Wang, Kaixuan, et al.
Published: (2026)
by: Wang, Kaixuan, et al.
Published: (2026)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
by: Meng, Fanqing, et al.
Published: (2026)
by: Meng, Fanqing, et al.
Published: (2026)
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
by: Zhao, Qingcheng, et al.
Published: (2024)
by: Zhao, Qingcheng, et al.
Published: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
by: Para, Wamiq Reyaz, et al.
Published: (2024)
by: Para, Wamiq Reyaz, et al.
Published: (2024)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
by: Fang, Shuangkang, et al.
Published: (2025)
by: Fang, Shuangkang, et al.
Published: (2025)
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
by: Tang, Jiapeng, et al.
Published: (2024)
by: Tang, Jiapeng, et al.
Published: (2024)
Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation
by: Ng, Kam Woh, et al.
Published: (2025)
by: Ng, Kam Woh, et al.
Published: (2025)
Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs
by: Scaramuzza, Filippo, et al.
Published: (2025)
by: Scaramuzza, Filippo, et al.
Published: (2025)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
by: Deng, Youming, et al.
Published: (2025)
by: Deng, Youming, et al.
Published: (2025)
Shap-MeD
by: Laverde, Nicolás, et al.
Published: (2025)
by: Laverde, Nicolás, et al.
Published: (2025)
ROMAN: Reward-Orchestrated Multi-Head Attention Network for Autonomous Driving System Testing
by: Chi, Jianlei, et al.
Published: (2026)
by: Chi, Jianlei, et al.
Published: (2026)
GUing: A Mobile GUI Search Engine using a Vision-Language Model
by: Wei, Jialiang, et al.
Published: (2024)
by: Wei, Jialiang, et al.
Published: (2024)
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
by: Li, Xinyang, et al.
Published: (2025)
by: Li, Xinyang, et al.
Published: (2025)
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
by: Wei, Guangshun, et al.
Published: (2024)
by: Wei, Guangshun, et al.
Published: (2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
by: Kuang, Zhengfei, et al.
Published: (2024)
by: Kuang, Zhengfei, et al.
Published: (2024)
Large-Scale Multi-Character Interaction Synthesis
by: Chang, Ziyi, et al.
Published: (2025)
by: Chang, Ziyi, et al.
Published: (2025)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
by: Zhong, Jianping, et al.
Published: (2026)
by: Zhong, Jianping, et al.
Published: (2026)
Multisensory extended reality applications offer benefits for volumetric biomedical image analysis in research and medicine
by: Krieger, Kathrin, et al.
Published: (2023)
by: Krieger, Kathrin, et al.
Published: (2023)
DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer
by: Huang, Nisha, et al.
Published: (2026)
by: Huang, Nisha, et al.
Published: (2026)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
by: Chen, Hansheng, et al.
Published: (2024)
by: Chen, Hansheng, et al.
Published: (2024)
Similar Items
-
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
by: Baresi, Luciano, et al.
Published: (2024) -
Can LLMs Generate User Stories and Assess Their Quality?
by: Quattrocchi, Giovanni, et al.
Published: (2025) -
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
by: Zhao, Guanghui, et al.
Published: (2026) -
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
by: Allison, Christopher John, et al.
Published: (2024) -
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
by: Jin, Duosi, et al.
Published: (2026)