Food Portion Estimation via 3D Object Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vinod, Gautham, He, Jiangpeng, Shao, Zeman, Zhu, Fengqing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Food Portion Estimation: From Pixels to Calories
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
von: Ma, Jinge, et al.
Veröffentlicht: (2024)
von: Ma, Jinge, et al.
Veröffentlicht: (2024)
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
von: Zhang, Xinjie, et al.
Veröffentlicht: (2024)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
von: Yang, Jinhai, et al.
Veröffentlicht: (2024)
Learning to Classify New Foods Incrementally Via Compressed Exemplars
von: Yang, Justin, et al.
Veröffentlicht: (2024)
von: Yang, Justin, et al.
Veröffentlicht: (2024)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
von: Liu, Mufan, et al.
Veröffentlicht: (2024)
Towards Real-world Video Face Restoration: A New Benchmark
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
von: Chen, Ziyan, et al.
Veröffentlicht: (2024)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
von: Liu, Tianyi, et al.
Veröffentlicht: (2025)
Change Detection Between Optical Remote Sensing Imagery and Map Data via Segment Anything Model (SAM)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
von: Chen, Hongruixuan, et al.
Veröffentlicht: (2024)
PANDA -- Patch And Distribution-Aware Augmentation for Long-Tailed Exemplar-Free Continual Learning
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
von: Raghavan, Siddeshwar, et al.
Veröffentlicht: (2025)
Machine Perception-Driven Image Compression: A Layered Generative Approach
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
von: Zhang, Yuefeng, et al.
Veröffentlicht: (2023)
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
von: Wang, Qizhou, et al.
Veröffentlicht: (2026)
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
von: Zi, Xing, et al.
Veröffentlicht: (2025)
von: Zi, Xing, et al.
Veröffentlicht: (2025)
PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
von: Mishra, Aradhana, et al.
Veröffentlicht: (2025)
RAISE: Realness Assessment for Image Synthesis and Evaluation
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
von: Mukherjee, Aniruddha, et al.
Veröffentlicht: (2025)
From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
von: Adhikarla, Eashan, et al.
Veröffentlicht: (2024)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
von: Rakesh, Vineet Kumar, et al.
Veröffentlicht: (2026)
Movie Trailer Genre Classification Using Multimodal Pretrained Features
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
von: Sulun, Serkan, et al.
Veröffentlicht: (2024)
MIND: A Noise-Adaptive Denoising Framework for Medical Images Integrating Multi-Scale Transformer
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
von: Zhu, Jun, et al.
Veröffentlicht: (2025)
von: Zhu, Jun, et al.
Veröffentlicht: (2025)
Scaling Up Single Image Dehazing Algorithm by Cross-Data Vision Alignment for Richer Representation Learning and Beyond
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
von: Shi, Yukai, et al.
Veröffentlicht: (2024)
Tackle CSM in JPEG Steganalysis with Data Adaptation
von: Abecidan, Rony, et al.
Veröffentlicht: (2026)
von: Abecidan, Rony, et al.
Veröffentlicht: (2026)
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
von: Roy, Prasun, et al.
Veröffentlicht: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
von: Xu, Chuanzhi, et al.
Veröffentlicht: (2026)
Exploring Event-based Human Pose Estimation with 3D Event Representations
von: Yin, Xiaoting, et al.
Veröffentlicht: (2023)
von: Yin, Xiaoting, et al.
Veröffentlicht: (2023)
Object-Attribute-Relation Representation Based Video Semantic Communication
von: Du, Qiyuan, et al.
Veröffentlicht: (2024)
von: Du, Qiyuan, et al.
Veröffentlicht: (2024)
V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache Retrieval
von: Kim, Donghyuk, et al.
Veröffentlicht: (2025)
von: Kim, Donghyuk, et al.
Veröffentlicht: (2025)
R$^3$D: Regional-guided Residual Radar Diffusion
von: Li, Hao, et al.
Veröffentlicht: (2026)
von: Li, Hao, et al.
Veröffentlicht: (2026)
MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model
von: Zeng, Kang, et al.
Veröffentlicht: (2024)
von: Zeng, Kang, et al.
Veröffentlicht: (2024)
Beyond Alignment: Blind Video Face Restoration via Parsing-Guided Temporal-Coherent Transformer
von: Xu, Kepeng, et al.
Veröffentlicht: (2024)
von: Xu, Kepeng, et al.
Veröffentlicht: (2024)
ConcealGS: Concealing Invisible Copyright Information in 3D Gaussian Splatting
von: Yang, Yifeng, et al.
Veröffentlicht: (2025)
von: Yang, Yifeng, et al.
Veröffentlicht: (2025)
Hyperspectral Image Fusion with Spectral-Band and Fusion-Scale Agnosticism
von: Liang, Yu-Jie, et al.
Veröffentlicht: (2026)
von: Liang, Yu-Jie, et al.
Veröffentlicht: (2026)
Sphere-GAN: a GAN-based Approach for Saliency Estimation in 360° Videos
von: Wahba, Mahmoud Z. A., et al.
Veröffentlicht: (2025)
von: Wahba, Mahmoud Z. A., et al.
Veröffentlicht: (2025)
Adaptive 3D Gaussian Splatting Video Streaming: Visual Saliency-Aware Tiling and Meta-Learning-Based Bitrate Adaptation
von: Gong, Han, et al.
Veröffentlicht: (2025)
von: Gong, Han, et al.
Veröffentlicht: (2025)
Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
Joint End-to-End Image Compression and Denoising: Leveraging Contrastive Learning and Multi-Scale Self-ONNs
von: Xie, Yuxin, et al.
Veröffentlicht: (2024)
von: Xie, Yuxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Food Portion Estimation: From Pixels to Calories
von: Vinod, Gautham, et al.
Veröffentlicht: (2026) -
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
von: Ma, Jinge, et al.
Veröffentlicht: (2024) -
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
von: Vinod, Gautham, et al.
Veröffentlicht: (2026) -
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
von: Vinod, Gautham, et al.
Veröffentlicht: (2026) -
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
von: Vinod, Gautham, et al.
Veröffentlicht: (2026)