A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Toyooka, Mashiro, Aizawa, Kiyoharu, Yamakata, Yoko |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)
FineFake: A Knowledge-Enriched Dataset for Fine-Grained Multi-Domain Fake News Detection
von: Zhou, Ziyi, et al.
Veröffentlicht: (2024)
von: Zhou, Ziyi, et al.
Veröffentlicht: (2024)
SemEval-2024 Task 3: Multimodal Emotion Cause Analysis in Conversations
von: Wang, Fanfan, et al.
Veröffentlicht: (2024)
von: Wang, Fanfan, et al.
Veröffentlicht: (2024)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
A Survey on Image-text Multimodal Models
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
von: Guo, Ruifeng, et al.
Veröffentlicht: (2023)
A Survey of Generative Categories and Techniques in Multimodal Generative Models
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
von: Han, Longzhen, et al.
Veröffentlicht: (2025)
A Multimodal Framework for Explainable Evaluation of Soft Skills in Educational Environments
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
von: He, Zheqi, et al.
Veröffentlicht: (2024)
von: He, Zheqi, et al.
Veröffentlicht: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
von: Lu, Jinghui, et al.
Veröffentlicht: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Traj-MLLM: Can Multimodal Large Language Models Reform Trajectory Data Mining?
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
von: Liu, Shuo, et al.
Veröffentlicht: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
Cardiverse: Harnessing LLMs for Novel Card Game Prototyping
von: Li, Danrui, et al.
Veröffentlicht: (2025)
von: Li, Danrui, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Generation for Electrocardiogram-Language Models
von: Song, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Song, Xiaoyu, et al.
Veröffentlicht: (2025)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
von: Wu, Zichen, et al.
Veröffentlicht: (2025)
MultimodalHugs: Enabling Sign Language Processing in Hugging Face
von: Sant, Gerard, et al.
Veröffentlicht: (2025)
von: Sant, Gerard, et al.
Veröffentlicht: (2025)
Memory-Centric Embodied Question Answering
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
von: Zhai, Mingliang, et al.
Veröffentlicht: (2025)
PTA: Enhancing Multimodal Sentiment Analysis through Pipelined Prediction and Translation-based Alignment
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
von: Song, Shezheng, et al.
Veröffentlicht: (2024)
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
von: Ren, Yiming, et al.
Veröffentlicht: (2026)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
Knowledge-Guided Dynamic Modality Attention Fusion Framework for Multimodal Sentiment Analysis
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
von: Feng, Xinyu, et al.
Veröffentlicht: (2024)
"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
von: Bohus, Dan, et al.
Veröffentlicht: (2024)
von: Bohus, Dan, et al.
Veröffentlicht: (2024)
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Affective Analysis with Learned Live Comment Features
von: Deng, Zhaoyuan, et al.
Veröffentlicht: (2024)
von: Deng, Zhaoyuan, et al.
Veröffentlicht: (2024)
History-Guided Iterative Visual Reasoning with Self-Correction
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
von: Yang, Xinglong, et al.
Veröffentlicht: (2026)
Temporal-Spatial Decouple before Act: Disentangled Representation Learning for Multimodal Sentiment Analysis
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023)
von: Liu, Hui, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
FoodMLLM-JP: Leveraging Multimodal Large Language Models for Japanese Recipe Generation
von: Imajuku, Yuki, et al.
Veröffentlicht: (2024) -
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025) -
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025) -
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2026) -
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
von: Miyai, Atsuyuki, et al.
Veröffentlicht: (2025)