L-MAGIC: Language Model Assisted Generation of Images with Coherence
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cai, Zhipeng, Mueller, Matthias, Birkl, Reiner, Wofk, Diana, Tseng, Shao-Yen, Cheng, JunDa, Stan, Gabriela Ben-Melech, Lal, Vasudev, Paulitsch, Michael |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FastRM: An efficient and automatic explainability framework for multimodal generative models
par: Stan, Gabriela Ben-Melech, et autres
Publié: (2024)
par: Stan, Gabriela Ben-Melech, et autres
Publié: (2024)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
par: Aflalo, Estelle, et autres
Publié: (2024)
par: Aflalo, Estelle, et autres
Publié: (2024)
Learning from Reasoning Failures via Synthetic Data Generation
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025)
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024)
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
par: Chen, Yujin, et autres
Publié: (2025)
par: Chen, Yujin, et autres
Publié: (2025)
KernelFoundry: Hardware-aware evolutionary GPU kernel optimization
par: Wiedemann, Nina, et autres
Publié: (2026)
par: Wiedemann, Nina, et autres
Publié: (2026)
Mesh2NeRF: Direct Mesh Supervision for Neural Radiance Field Representation and Generation
par: Chen, Yujin, et autres
Publié: (2024)
par: Chen, Yujin, et autres
Publié: (2024)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
par: Hinck, Musashi, et autres
Publié: (2024)
par: Hinck, Musashi, et autres
Publié: (2024)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
par: Chatterjee, Agneet, et autres
Publié: (2024)
par: Chatterjee, Agneet, et autres
Publié: (2024)
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
par: Ratzlaff, Neale, et autres
Publié: (2024)
par: Ratzlaff, Neale, et autres
Publié: (2024)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
par: Luo, Man, et autres
Publié: (2025)
par: Luo, Man, et autres
Publié: (2025)
Probing the Representational Power of Sparse Autoencoders in Vision Models
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
par: Ratzlaff, Neale, et autres
Publié: (2024)
par: Ratzlaff, Neale, et autres
Publié: (2024)
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
par: Olson, Matthew Lyle, et autres
Publié: (2025)
par: Olson, Matthew Lyle, et autres
Publié: (2025)
Steering Large Language Models to Evaluate and Amplify Creativity
par: Olson, Matthew Lyle, et autres
Publié: (2024)
par: Olson, Matthew Lyle, et autres
Publié: (2024)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
par: Madasu, Avinash, et autres
Publié: (2023)
par: Madasu, Avinash, et autres
Publié: (2023)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
par: Olson, Matthew Lyle, et autres
Publié: (2026)
par: Olson, Matthew Lyle, et autres
Publié: (2026)
RoMeO: Robust Metric Visual Odometry
par: Cheng, Junda, et autres
Publié: (2024)
par: Cheng, Junda, et autres
Publié: (2024)
Adaptive Fusion of Single-View and Multi-View Depth for Autonomous Driving
par: Cheng, JunDa, et autres
Publié: (2024)
par: Cheng, JunDa, et autres
Publié: (2024)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
par: Spoczynski, Marcin, et autres
Publié: (2026)
par: Spoczynski, Marcin, et autres
Publié: (2026)
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
par: Rosenman, Shachar, et autres
Publié: (2023)
par: Rosenman, Shachar, et autres
Publié: (2023)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
par: Madasu, Avinash, et autres
Publié: (2025)
par: Madasu, Avinash, et autres
Publié: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
par: Madasu, Avinash, et autres
Publié: (2025)
par: Madasu, Avinash, et autres
Publié: (2025)
A simbologia da devoção: o retrato da fé demonstrado pelos ex-votos e a relação com a Igreja midiatizada
par: Ana Maria de Souza Melech
Publié: (2015)
par: Ana Maria de Souza Melech
Publié: (2015)
Rapid Salient Object Detection with Difference Convolutional Neural Networks
par: Su, Zhuo, et autres
Publié: (2025)
par: Su, Zhuo, et autres
Publié: (2025)
Coherent Single‐Atom Dipole–Dipole Coupling Mediates Holistic Regulation of K+ Migration for Superior Energy Storage and Dendrite‐Free Metal Deposition
par: Yen‐Yang Tseng, et autres
Publié: (2025)
par: Yen‐Yang Tseng, et autres
Publié: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
par: Madasu, Avinash, et autres
Publié: (2024)
par: Madasu, Avinash, et autres
Publié: (2024)
Fractional powers of the backward heat operator and Carleman type inequalities
par: Stan, Diana
Publié: (2025)
par: Stan, Diana
Publié: (2025)
ConDo: Continual Domain Expansion for Absolute Pose Regression
par: Li, Zijun, et autres
Publié: (2024)
par: Li, Zijun, et autres
Publié: (2024)
Recovering Unobserved Network Links from Aggregated Relational Data: Discussions on Bayesian Latent Surface Modeling and Penalized Regression
par: Tseng, Yen-hsuan
Publié: (2025)
par: Tseng, Yen-hsuan
Publié: (2025)
Inverse Gaussian Distribution, Introduction and Applications:Comprehensive Analysis of Power Plant Performance: A Study of Combined Cycle and Nuclear Power Plant
par: Tseng, Yen-hsuan
Publié: (2025)
par: Tseng, Yen-hsuan
Publié: (2025)
MAGIC and AGN
par: Aimo Sillanpää
Publié: (2008)
par: Aimo Sillanpää
Publié: (2008)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
par: Yu, Sungduk, et autres
Publié: (2025)
par: Yu, Sungduk, et autres
Publié: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
par: Yu, Sungduk, et autres
Publié: (2024)
par: Yu, Sungduk, et autres
Publié: (2024)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
par: Ratzlaff, Neale, et autres
Publié: (2024)
par: Ratzlaff, Neale, et autres
Publié: (2024)
The EarlyBird Gets the WORM: Heuristically Accelerating EarlyBird Convergence
par: Vasudev, Adithya
Publié: (2024)
par: Vasudev, Adithya
Publié: (2024)
Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks
par: Gohil, Vasudev
Publié: (2025)
par: Gohil, Vasudev
Publié: (2025)
The Self-Healing Effect on Bacteria-Enriched Steel Fiber-Reinforced SCC
par: Vasudev Raman
Publié: (2022)
par: Vasudev Raman
Publié: (2022)
Why do LLaVA Vision-Language Models Reply to Images in English?
par: Hinck, Musashi, et autres
Publié: (2024)
par: Hinck, Musashi, et autres
Publié: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
par: Rohekar, Raanan Y., et autres
Publié: (2024)
par: Rohekar, Raanan Y., et autres
Publié: (2024)
Documents similaires
-
FastRM: An efficient and automatic explainability framework for multimodal generative models
par: Stan, Gabriela Ben-Melech, et autres
Publié: (2024) -
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
par: Aflalo, Estelle, et autres
Publié: (2024) -
Learning from Reasoning Failures via Synthetic Data Generation
par: Stan, Gabriela Ben Melech, et autres
Publié: (2025) -
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
par: Stan, Gabriela Ben Melech, et autres
Publié: (2024) -
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
par: Chen, Yujin, et autres
Publié: (2025)