Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vinod, Gautham, Coburn, Bruce, Raghavan, Siddeshwar, He, Jiangpeng, Zhu, Fengqing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Food Portion Estimation via 3D Object Scaling
von: Vinod, Gautham, et al.
Veröffentlicht: (2024)
von: Vinod, Gautham, et al.
Veröffentlicht: (2024)
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
von: Ma, Jinge, et al.
Veröffentlicht: (2024)
von: Ma, Jinge, et al.
Veröffentlicht: (2024)
Implicit-Scale 3D Reconstruction for Multi-Food Volume Estimation from Monocular Images
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
Food Portion Estimation: From Pixels to Calories
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Online Class-Incremental Learning For Real-World Food Image Classification
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2023)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2023)
Can You Hear, Localize, and Segment Continually? An Exemplar-Free Continual Learning Benchmark for Audio-Visual Segmentation
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2026)
DELTA: Decoupling Long-Tailed Online Continual Learning
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2024)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2024)
PANDA -- Patch And Distribution-Aware Augmentation for Long-Tailed Exemplar-Free Continual Learning
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
MetaFood3D: 3D Food Dataset with Nutrition Values
von: Chen, Yuhao, et al.
Veröffentlicht: (2024)
von: Chen, Yuhao, et al.
Veröffentlicht: (2024)
Automatic Recognition of Food Ingestion Environment from the AIM-2 Wearable Sensor
von: Huang, Yuning, et al.
Veröffentlicht: (2024)
von: Huang, Yuning, et al.
Veröffentlicht: (2024)
FMiFood: Multi-modal Contrastive Learning for Food Image Classification
von: Pan, Xinyue, et al.
Veröffentlicht: (2024)
von: Pan, Xinyue, et al.
Veröffentlicht: (2024)
FoodLogAthl-218: Constructing a Real-World Food Image Dataset Using Dietary Management Applications
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
von: Watanabe, Mitsuki, et al.
Veröffentlicht: (2025)
GGAvatar: Reconstructing Garment-Separated 3D Gaussian Splatting Avatars from Monocular Video
von: Chen, Jingxuan
Veröffentlicht: (2024)
von: Chen, Jingxuan
Veröffentlicht: (2024)
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
Food Image Generation on Multi-Noun Categories
von: Pan, Xinyue, et al.
Veröffentlicht: (2025)
von: Pan, Xinyue, et al.
Veröffentlicht: (2025)
Robust3D-CIL: Robust Class-Incremental Learning for 3D Perception
von: Ma, Jinge, et al.
Veröffentlicht: (2025)
von: Ma, Jinge, et al.
Veröffentlicht: (2025)
SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
von: Xie, Shuzhao, et al.
Veröffentlicht: (2024)
von: Xie, Shuzhao, et al.
Veröffentlicht: (2024)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
von: Yan, Yichen, et al.
Veröffentlicht: (2024)
von: Yan, Yichen, et al.
Veröffentlicht: (2024)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion
von: Liu, Yu, et al.
Veröffentlicht: (2024)
von: Liu, Yu, et al.
Veröffentlicht: (2024)
Gradient Reweighting: Towards Imbalanced Class-Incremental Learning
von: He, Jiangpeng, et al.
Veröffentlicht: (2024)
von: He, Jiangpeng, et al.
Veröffentlicht: (2024)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
SMPLer: Taming Transformers for Monocular 3D Human Shape and Pose Estimation
von: Xu, Xiangyu, et al.
Veröffentlicht: (2024)
von: Xu, Xiangyu, et al.
Veröffentlicht: (2024)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
Marker-Based Extrinsic Calibration Method for Accurate Multi-Camera 3D Reconstruction
von: Garcia-D'Urso, Nahuel, et al.
Veröffentlicht: (2025)
von: Garcia-D'Urso, Nahuel, et al.
Veröffentlicht: (2025)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
von: Guo, Zile, et al.
Veröffentlicht: (2026)
von: Guo, Zile, et al.
Veröffentlicht: (2026)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata
von: Coburn, Bruce, et al.
Veröffentlicht: (2025)
von: Coburn, Bruce, et al.
Veröffentlicht: (2025)
3D Gaussian Editing with A Single Image
von: Luo, Guan, et al.
Veröffentlicht: (2024)
von: Luo, Guan, et al.
Veröffentlicht: (2024)
MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
von: Zhao, Haochen, et al.
Veröffentlicht: (2025)
von: Zhao, Haochen, et al.
Veröffentlicht: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
Leveraging Automatic Personalised Nutrition: Food Image Recognition Benchmark and Dataset based on Nutrition Taxonomy
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
von: Romero-Tapiador, Sergio, et al.
Veröffentlicht: (2022)
REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
One Size, Many Fits: Aligning Diverse Group-Wise Click Preferences in Large-Scale Advertising Image Generation
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Food Portion Estimation via 3D Object Scaling
von: Vinod, Gautham, et al.
Veröffentlicht: (2024) -
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026) -
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026) -
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
von: Ma, Jinge, et al.
Veröffentlicht: (2024) -
Implicit-Scale 3D Reconstruction for Multi-Food Volume Estimation from Monocular Images
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)