DILLEMA: Diffusion and Large Language Models for Multi-Modal Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Baresi, Luciano, Hu, Davide Yi Xian, Mas'udi, Muhammad Irfan, Quattrocchi, Giovanni |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
von: Baresi, Luciano, et al.
Veröffentlicht: (2024)
von: Baresi, Luciano, et al.
Veröffentlicht: (2024)
Can LLMs Generate User Stories and Assess Their Quality?
von: Quattrocchi, Giovanni, et al.
Veröffentlicht: (2025)
von: Quattrocchi, Giovanni, et al.
Veröffentlicht: (2025)
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
von: Zhao, Guanghui, et al.
Veröffentlicht: (2026)
von: Zhao, Guanghui, et al.
Veröffentlicht: (2026)
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
von: Allison, Christopher John, et al.
Veröffentlicht: (2024)
von: Allison, Christopher John, et al.
Veröffentlicht: (2024)
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
von: Jin, Duosi, et al.
Veröffentlicht: (2026)
von: Jin, Duosi, et al.
Veröffentlicht: (2026)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
von: Oketunji, Abiodun Finbarrs
Veröffentlicht: (2024)
Autark: A Serverless Toolkit for Prototyping Urban Visual Analytics Systems
von: Alexandre, Lucas, et al.
Veröffentlicht: (2026)
von: Alexandre, Lucas, et al.
Veröffentlicht: (2026)
TARIPlay: A Test Framework for AR Applications based on Interactive Area Tracking in Playback Videos
von: Mousavi, Seyed Amir, et al.
Veröffentlicht: (2026)
von: Mousavi, Seyed Amir, et al.
Veröffentlicht: (2026)
Code Arcades: 3d Visualization of Classes, Dependencies and Software Metrics
von: Savidis, Anthony, et al.
Veröffentlicht: (2025)
von: Savidis, Anthony, et al.
Veröffentlicht: (2025)
The Vibe-Check Protocol: Quantifying Cognitive Offloading in AI Programming
von: Aiersilan, Aizierjiang
Veröffentlicht: (2026)
von: Aiersilan, Aizierjiang
Veröffentlicht: (2026)
Investigating Traffic Accident Detection Using Multimodal Large Language Models
von: Skender, Ilhan, et al.
Veröffentlicht: (2025)
von: Skender, Ilhan, et al.
Veröffentlicht: (2025)
WAFFLE: Finetuning Multi-Modal Models for Automated Front-End Development
von: Liang, Shanchao, et al.
Veröffentlicht: (2024)
von: Liang, Shanchao, et al.
Veröffentlicht: (2024)
A Plausibility Study of Using Augmented Reality in the Ventriculoperitoneal Shunt Operations
von: Dorji, Tandin, et al.
Veröffentlicht: (2024)
von: Dorji, Tandin, et al.
Veröffentlicht: (2024)
Can LLMs Solve Science or Just Write Code? Evaluating Quantum Solver Generation
von: Baresi, Luciano, et al.
Veröffentlicht: (2026)
von: Baresi, Luciano, et al.
Veröffentlicht: (2026)
SMooGPT: Stylized Motion Generation using Large Language Models
von: Zhong, Lei, et al.
Veröffentlicht: (2025)
von: Zhong, Lei, et al.
Veröffentlicht: (2025)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
von: Yin, Zixin, et al.
Veröffentlicht: (2025)
von: Yin, Zixin, et al.
Veröffentlicht: (2025)
TigAug: Data Augmentation for Testing Traffic Light Detection in Autonomous Driving Systems
von: Lu, You, et al.
Veröffentlicht: (2025)
von: Lu, You, et al.
Veröffentlicht: (2025)
A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
von: Khalid, Waleed, et al.
Veröffentlicht: (2025)
von: Khalid, Waleed, et al.
Veröffentlicht: (2025)
Chat2SVG: Vector Graphics Generation with Large Language Models and Image Diffusion Models
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
von: Wu, Ronghuan, et al.
Veröffentlicht: (2024)
MMCode: Benchmarking Multimodal Large Language Models for Code Generation with Visually Rich Programming Problems
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
von: Li, Kaixin, et al.
Veröffentlicht: (2024)
ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K
von: Wang, Kaixuan, et al.
Veröffentlicht: (2026)
von: Wang, Kaixuan, et al.
Veröffentlicht: (2026)
ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents
von: Meng, Fanqing, et al.
Veröffentlicht: (2026)
von: Meng, Fanqing, et al.
Veröffentlicht: (2026)
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
von: Zhao, Qingcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Qingcheng, et al.
Veröffentlicht: (2024)
AvatarMMC: 3D Head Avatar Generation and Editing with Multi-Modal Conditioning
von: Para, Wamiq Reyaz, et al.
Veröffentlicht: (2024)
von: Para, Wamiq Reyaz, et al.
Veröffentlicht: (2024)
MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
von: Fang, Shuangkang, et al.
Veröffentlicht: (2025)
von: Fang, Shuangkang, et al.
Veröffentlicht: (2025)
GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view Diffusion
von: Tang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Tang, Jiapeng, et al.
Veröffentlicht: (2024)
Chirpy3D: Part-Aware Multi-View Diffusion for Creative Fine-Grained Object Generation
von: Ng, Kam Woh, et al.
Veröffentlicht: (2025)
von: Ng, Kam Woh, et al.
Veröffentlicht: (2025)
Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs
von: Scaramuzza, Filippo, et al.
Veröffentlicht: (2025)
von: Scaramuzza, Filippo, et al.
Veröffentlicht: (2025)
Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction
von: Deng, Youming, et al.
Veröffentlicht: (2025)
von: Deng, Youming, et al.
Veröffentlicht: (2025)
Shap-MeD
von: Laverde, Nicolás, et al.
Veröffentlicht: (2025)
von: Laverde, Nicolás, et al.
Veröffentlicht: (2025)
ROMAN: Reward-Orchestrated Multi-Head Attention Network for Autonomous Driving System Testing
von: Chi, Jianlei, et al.
Veröffentlicht: (2026)
von: Chi, Jianlei, et al.
Veröffentlicht: (2026)
GUing: A Mobile GUI Search Engine using a Vision-Language Model
von: Wei, Jialiang, et al.
Veröffentlicht: (2024)
von: Wei, Jialiang, et al.
Veröffentlicht: (2024)
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
von: Li, Xinyang, et al.
Veröffentlicht: (2025)
von: Li, Xinyang, et al.
Veröffentlicht: (2025)
PCDreamer: Point Cloud Completion Through Multi-view Diffusion Priors
von: Wei, Guangshun, et al.
Veröffentlicht: (2024)
von: Wei, Guangshun, et al.
Veröffentlicht: (2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2024)
Large-Scale Multi-Character Interaction Synthesis
von: Chang, Ziyi, et al.
Veröffentlicht: (2025)
von: Chang, Ziyi, et al.
Veröffentlicht: (2025)
Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression
von: Zhong, Jianping, et al.
Veröffentlicht: (2026)
von: Zhong, Jianping, et al.
Veröffentlicht: (2026)
Multisensory extended reality applications offer benefits for volumetric biomedical image analysis in research and medicine
von: Krieger, Kathrin, et al.
Veröffentlicht: (2023)
von: Krieger, Kathrin, et al.
Veröffentlicht: (2023)
DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer
von: Huang, Nisha, et al.
Veröffentlicht: (2026)
von: Huang, Nisha, et al.
Veröffentlicht: (2026)
Generic 3D Diffusion Adapter Using Controlled Multi-View Editing
von: Chen, Hansheng, et al.
Veröffentlicht: (2024)
von: Chen, Hansheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient Domain Augmentation for Autonomous Driving Testing Using Diffusion Models
von: Baresi, Luciano, et al.
Veröffentlicht: (2024) -
Can LLMs Generate User Stories and Assess Their Quality?
von: Quattrocchi, Giovanni, et al.
Veröffentlicht: (2025) -
Toward Reliable Scientific Visualization Pipeline Construction with Structure-Aware Retrieval-Augmented LLMs
von: Zhao, Guanghui, et al.
Veröffentlicht: (2026) -
FIRE-3DV: Framework-Independent Rendering Engine for 3D Graphics using Vulkan
von: Allison, Christopher John, et al.
Veröffentlicht: (2024) -
VARTS: A Tool for the Visualization and Analysis of Representative Time Series Data
von: Jin, Duosi, et al.
Veröffentlicht: (2026)