OVFoodSeg: Elevating Open-Vocabulary Food Image Segmentation via Image-Informed Textual Representation
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Xiongwei, Yu, Sicheng, Lim, Ee-Peng, Ngo, Chong-Wah |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024)
by: Ye, Chengyang, et al.
Published: (2024)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
by: Satar, Burak, et al.
Published: (2025)
by: Satar, Burak, et al.
Published: (2025)
Robust Relevance Feedback for Interactive Known-Item Video Search
by: Ma, Zhixin, et al.
Published: (2025)
by: Ma, Zhixin, et al.
Published: (2025)
MSPT: A Lightweight Face Image Quality Assessment Method with Multi-stage Progressive Training
by: Xiao, Xiongwei, et al.
Published: (2025)
by: Xiao, Xiongwei, et al.
Published: (2025)
Efficient Prompt Tuning for Hierarchical Ingredient Recognition
by: Gui, Yinxuan, et al.
Published: (2025)
by: Gui, Yinxuan, et al.
Published: (2025)
SIA-OVD: Shape-Invariant Adapter for Bridging the Image-Region Gap in Open-Vocabulary Detection
by: Wang, Zishuo, et al.
Published: (2024)
by: Wang, Zishuo, et al.
Published: (2024)
Class Agnostic Instance-level Descriptor for Visual Instance Search
by: Sun, Qi-Ying, et al.
Published: (2025)
by: Sun, Qi-Ying, et al.
Published: (2025)
Self-supervised Photographic Image Layout Representation Learning
by: Zhao, Zhaoran, et al.
Published: (2024)
by: Zhao, Zhaoran, et al.
Published: (2024)
Bringing Textual Prompt to AI-Generated Image Quality Assessment
by: Qu, Bowen, et al.
Published: (2024)
by: Qu, Bowen, et al.
Published: (2024)
Navigating Weight Prediction with Diet Diary
by: Gui, Yinxuan, et al.
Published: (2024)
by: Gui, Yinxuan, et al.
Published: (2024)
TOL: Textual Localization with OpenStreetMap
by: Liao, Youqi, et al.
Published: (2026)
by: Liao, Youqi, et al.
Published: (2026)
LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
SegTalker: Segmentation-based Talking Face Generation with Mask-guided Local Editing
by: Xiong, Lingyu, et al.
Published: (2024)
by: Xiong, Lingyu, et al.
Published: (2024)
Hierarchical Textual Knowledge for Enhanced Image Clustering
by: Zhong, Yijie, et al.
Published: (2026)
by: Zhong, Yijie, et al.
Published: (2026)
Towards Multimodal Emotional Support Conversation Systems
by: Chu, Yuqi, et al.
Published: (2024)
by: Chu, Yuqi, et al.
Published: (2024)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
by: Lin, Haoqiang, et al.
Published: (2025)
by: Lin, Haoqiang, et al.
Published: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Towards Open-Vocabulary Video Semantic Segmentation
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Towards Open-Vocabulary Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
by: Qi, Huiyan, et al.
Published: (2025)
by: Qi, Huiyan, et al.
Published: (2025)
A Simple Baseline with Single-encoder for Referring Image Segmentation
by: Yu, Seonghoon, et al.
Published: (2024)
by: Yu, Seonghoon, et al.
Published: (2024)
Open-Vocabulary Audio-Visual Semantic Segmentation
by: Guo, Ruohao, et al.
Published: (2024)
by: Guo, Ruohao, et al.
Published: (2024)
Multimodal LLM-based Query Paraphrasing for Video Search
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
by: Watanabe, Mitsuki, et al.
Published: (2025)
by: Watanabe, Mitsuki, et al.
Published: (2025)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
Robust Latent Representation Tuning for Image-text Classification
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
by: Yan, Yichen, et al.
Published: (2024)
by: Yan, Yichen, et al.
Published: (2024)
Leveraging Automatic Personalised Nutrition: Food Image Recognition Benchmark and Dataset based on Nutrition Taxonomy
by: Romero-Tapiador, Sergio, et al.
Published: (2022)
by: Romero-Tapiador, Sergio, et al.
Published: (2022)
TA-V2A: Textually Assisted Video-to-Audio Generation
by: You, Yuhuan, et al.
Published: (2025)
by: You, Yuhuan, et al.
Published: (2025)
Referring Flexible Image Restoration
by: Guan, Runwei, et al.
Published: (2024)
by: Guan, Runwei, et al.
Published: (2024)
AcoustEmo: Open-Vocabulary Emotion Reasoning via Utterance-Aware Acoustic Q-Former
by: Zhang, Liyun, et al.
Published: (2026)
by: Zhang, Liyun, et al.
Published: (2026)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Generating Attribute-Aware Human Motions from Textual Prompt
by: Wang, Xinghan, et al.
Published: (2025)
by: Wang, Xinghan, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
by: Chen, Shengkai, et al.
Published: (2025)
by: Chen, Shengkai, et al.
Published: (2025)
GSCodec Studio: A Modular Framework for Gaussian Splat Compression
by: Li, Sicheng, et al.
Published: (2025)
by: Li, Sicheng, et al.
Published: (2025)
Similar Items
-
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025) -
Interpretable Embedding for Ad-hoc Video Search
by: Wu, Jiaxin, et al.
Published: (2024) -
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
by: Ye, Chengyang, et al.
Published: (2024) -
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
by: Yang, Haibo, et al.
Published: (2024)