Deep Image-to-Recipe Translation
Fuente:
arXiv
Guardado en:
| Autores principales: | Ma, Jiangqin, Mawji, Bilal, Williams, Franz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A U-Net and Transformer Pipeline for Multilingual Image Translation
por: Sahay, Siddharth, et al.
Publicado: (2025)
por: Sahay, Siddharth, et al.
Publicado: (2025)
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
por: Bromonschenkel, Gabriel, et al.
Publicado: (2026)
por: Bromonschenkel, Gabriel, et al.
Publicado: (2026)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
por: Wang, Jinyin, et al.
Publicado: (2024)
por: Wang, Jinyin, et al.
Publicado: (2024)
Scaling Sign Language Translation
por: Zhang, Biao, et al.
Publicado: (2024)
por: Zhang, Biao, et al.
Publicado: (2024)
KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models
por: Mohbat, Fnu, et al.
Publicado: (2025)
por: Mohbat, Fnu, et al.
Publicado: (2025)
Translation-Enhanced Multilingual Text-to-Image Generation
por: Li, Yaoyiran, et al.
Publicado: (2023)
por: Li, Yaoyiran, et al.
Publicado: (2023)
Real-time Bangla Sign Language Translator
por: Pranto, Rotan Hawlader, et al.
Publicado: (2024)
por: Pranto, Rotan Hawlader, et al.
Publicado: (2024)
Crossing Language Borders: A Pipeline for Indonesian Manhwa Translation
por: Narasimhan, Nithyasri, et al.
Publicado: (2025)
por: Narasimhan, Nithyasri, et al.
Publicado: (2025)
Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
por: Hakim, Zaber Ibn Abdul, et al.
Publicado: (2023)
por: Hakim, Zaber Ibn Abdul, et al.
Publicado: (2023)
DeepCoT: Deep Continual Transformers for Real-Time Inference on Data Streams
por: Picón, Ginés Carreto, et al.
Publicado: (2025)
por: Picón, Ginés Carreto, et al.
Publicado: (2025)
Adaptative Context Normalization: A Boost for Deep Learning in Image Processing
por: Faye, Bilal, et al.
Publicado: (2024)
por: Faye, Bilal, et al.
Publicado: (2024)
Improving Deep Generative Models on Many-To-One Image-to-Image Translation
por: Saxena, Sagar, et al.
Publicado: (2024)
por: Saxena, Sagar, et al.
Publicado: (2024)
MOFI: Learning Image Representations from Noisy Entity Annotated Images
por: Wu, Wentao, et al.
Publicado: (2023)
por: Wu, Wentao, et al.
Publicado: (2023)
Deep Augmentation: Dropout as Augmentation for Self-Supervised Learning
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2023)
por: Brüel-Gabrielsson, Rickard, et al.
Publicado: (2023)
MAGIC: Near-Optimal Data Attribution for Deep Learning
por: Ilyas, Andrew, et al.
Publicado: (2025)
por: Ilyas, Andrew, et al.
Publicado: (2025)
ImageChain: Advancing Sequential Image-to-Text Reasoning in Multimodal Large Language Models
por: Villegas, Danae Sánchez, et al.
Publicado: (2025)
por: Villegas, Danae Sánchez, et al.
Publicado: (2025)
IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models
por: Lei, Jiayi, et al.
Publicado: (2025)
por: Lei, Jiayi, et al.
Publicado: (2025)
Dual-Process Image Generation
por: Luo, Grace, et al.
Publicado: (2025)
por: Luo, Grace, et al.
Publicado: (2025)
Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
por: Gafni, Tomer, et al.
Publicado: (2025)
por: Gafni, Tomer, et al.
Publicado: (2025)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
por: Simoncini, Walter, et al.
Publicado: (2024)
por: Simoncini, Walter, et al.
Publicado: (2024)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
por: Du, Yao, et al.
Publicado: (2026)
por: Du, Yao, et al.
Publicado: (2026)
Centered Masking for Language-Image Pre-Training
por: Liang, Mingliang, et al.
Publicado: (2024)
por: Liang, Mingliang, et al.
Publicado: (2024)
Evaluating Numerical Reasoning in Text-to-Image Models
por: Kajić, Ivana, et al.
Publicado: (2024)
por: Kajić, Ivana, et al.
Publicado: (2024)
A Survey of Deep Learning for Geometry Problem Solving
por: Ma, Jianzhe, et al.
Publicado: (2025)
por: Ma, Jianzhe, et al.
Publicado: (2025)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
por: Han, Tianyang, et al.
Publicado: (2024)
por: Han, Tianyang, et al.
Publicado: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
por: Li, Hao, et al.
Publicado: (2024)
por: Li, Hao, et al.
Publicado: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
por: Yu, Eric Yang, et al.
Publicado: (2024)
por: Yu, Eric Yang, et al.
Publicado: (2024)
Embedding Geometries of Contrastive Language-Image Pre-Training
por: Chou, Jason Chuan-Chih, et al.
Publicado: (2024)
por: Chou, Jason Chuan-Chih, et al.
Publicado: (2024)
Alt-Text with Context: Improving Accessibility for Images on Twitter
por: Srivatsan, Nikita, et al.
Publicado: (2023)
por: Srivatsan, Nikita, et al.
Publicado: (2023)
Linear Alignment of Vision-language Models for Image Captioning
por: Paischer, Fabian, et al.
Publicado: (2023)
por: Paischer, Fabian, et al.
Publicado: (2023)
LatentExplainer: Explaining Latent Representations in Deep Generative Models with Multimodal Large Language Models
por: Zhu, Mengdan, et al.
Publicado: (2024)
por: Zhu, Mengdan, et al.
Publicado: (2024)
Residual-based Language Models are Free Boosters for Biomedical Imaging
por: Lai, Zhixin, et al.
Publicado: (2024)
por: Lai, Zhixin, et al.
Publicado: (2024)
Text-to-Image Cross-Modal Generation: A Systematic Review
por: Żelaszczyk, Maciej, et al.
Publicado: (2024)
por: Żelaszczyk, Maciej, et al.
Publicado: (2024)
JPEG-LM: LLMs as Image Generators with Canonical Codec Representations
por: Han, Xiaochuang, et al.
Publicado: (2024)
por: Han, Xiaochuang, et al.
Publicado: (2024)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
por: Yin, Shukang, et al.
Publicado: (2024)
por: Yin, Shukang, et al.
Publicado: (2024)
MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification
por: Fieback, Laura, et al.
Publicado: (2024)
por: Fieback, Laura, et al.
Publicado: (2024)
BiasConnect: Investigating Bias Interactions in Text-to-Image Models
por: Shukla, Pushkar, et al.
Publicado: (2025)
por: Shukla, Pushkar, et al.
Publicado: (2025)
Improving MLLM Historical Record Extraction with Test-Time Image
por: Archibald, Taylor, et al.
Publicado: (2025)
por: Archibald, Taylor, et al.
Publicado: (2025)
Till the Layers Collapse: Compressing a Deep Neural Network through the Lenses of Batch Normalization Layers
por: Liao, Zhu, et al.
Publicado: (2024)
por: Liao, Zhu, et al.
Publicado: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
por: Singla, Vasu, et al.
Publicado: (2024)
por: Singla, Vasu, et al.
Publicado: (2024)
Ejemplares similares
-
A U-Net and Transformer Pipeline for Multilingual Image Translation
por: Sahay, Siddharth, et al.
Publicado: (2025) -
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
por: Bromonschenkel, Gabriel, et al.
Publicado: (2026) -
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
por: Wang, Jinyin, et al.
Publicado: (2024) -
Scaling Sign Language Translation
por: Zhang, Biao, et al.
Publicado: (2024) -
KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models
por: Mohbat, Fnu, et al.
Publicado: (2025)