Retrieval Augmented Recipe Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Guoshan, Yin, Hailong, Zhu, Bin, Chen, Jingjing, Ngo, Chong-Wah, Jiang, Yu-Gang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
by: Yin, Yuehao, et al.
Published: (2023)
by: Yin, Yuehao, et al.
Published: (2023)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
by: Jiao, Pengkun, et al.
Published: (2025)
by: Jiao, Pengkun, et al.
Published: (2025)
From Canteen Food to Daily Meals: Generalizing Food Recognition to More Practical Scenarios
by: Liu, Guoshan, et al.
Published: (2024)
by: Liu, Guoshan, et al.
Published: (2024)
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
by: Liu, Guoshan, et al.
Published: (2026)
by: Liu, Guoshan, et al.
Published: (2026)
Benchmarking Gaslighting Negation Attacks Against Reasoning Models
by: Zhu, Bin, et al.
Published: (2025)
by: Zhu, Bin, et al.
Published: (2025)
RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
by: Qi, Huiyan, et al.
Published: (2025)
by: Qi, Huiyan, et al.
Published: (2025)
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Efficient Test-Time Retrieval Augmented Generation
by: Yin, Hailong, et al.
Published: (2025)
by: Yin, Hailong, et al.
Published: (2025)
CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
by: Wang, Yuan, et al.
Published: (2025)
by: Wang, Yuan, et al.
Published: (2025)
RecipeGen: A Benchmark for Real-World Recipe Image Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
by: Hao, Yanbin, et al.
Published: (2024)
by: Hao, Yanbin, et al.
Published: (2024)
Spatial Retrieval Augmented Autonomous Driving
by: Jia, Xiaosong, et al.
Published: (2025)
by: Jia, Xiaosong, et al.
Published: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
by: Loo, Gowen, et al.
Published: (2025)
by: Loo, Gowen, et al.
Published: (2025)
Retrieval-Augmented Generation for AI-Generated Content: A Survey
by: Zhao, Penghao, et al.
Published: (2024)
by: Zhao, Penghao, et al.
Published: (2024)
OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
by: Wu, Xiongwei, et al.
Published: (2024)
by: Wu, Xiongwei, et al.
Published: (2024)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
by: Zhang, Ruoxuan, et al.
Published: (2025)
by: Zhang, Ruoxuan, et al.
Published: (2025)
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
by: Tang, Ziyao, et al.
Published: (2026)
by: Tang, Ziyao, et al.
Published: (2026)
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
by: Chen, Qi, et al.
Published: (2026)
by: Chen, Qi, et al.
Published: (2026)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
by: Li, Yian, et al.
Published: (2026)
by: Li, Yian, et al.
Published: (2026)
VCE: Safe Autoregressive Image Generation via Visual Contrast Exploitation
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
by: Zhou, Shijie, et al.
Published: (2026)
by: Zhou, Shijie, et al.
Published: (2026)
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Domain Expansion and Boundary Growth for Open-Set Single-Source Domain Generalization
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
by: Chen, I-Hsiang, et al.
Published: (2026)
by: Chen, I-Hsiang, et al.
Published: (2026)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
by: Cao, Linhan, et al.
Published: (2026)
by: Cao, Linhan, et al.
Published: (2026)
From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
by: Yu, Yue, et al.
Published: (2026)
by: Yu, Yue, et al.
Published: (2026)
MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
by: Zhu, Chenhui, et al.
Published: (2025)
by: Zhu, Chenhui, et al.
Published: (2025)
Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
by: Gong, Chao, et al.
Published: (2024)
by: Gong, Chao, et al.
Published: (2024)
EAGLE: Towards Efficient Arbitrary Referring Visual Prompts Comprehension for Multimodal Large Language Models
by: Zhang, Jiacheng, et al.
Published: (2024)
by: Zhang, Jiacheng, et al.
Published: (2024)
Class Agnostic Instance-level Descriptor for Visual Instance Search
by: Sun, Qi-Ying, et al.
Published: (2025)
by: Sun, Qi-Ying, et al.
Published: (2025)
Similar Items
-
FoodLMM: A Versatile Food Assistant using Large Multi-modal Model
by: Yin, Yuehao, et al.
Published: (2023) -
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning
by: Jiao, Pengkun, et al.
Published: (2024) -
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
by: Jiao, Pengkun, et al.
Published: (2025)