Food Portion Estimation: From Pixels to Calories
Fuente:
arXiv
Guardado en:
| Autores principales: | Vinod, Gautham, Zhu, Fengqing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Food Portion Estimation via 3D Object Scaling
por: Vinod, Gautham, et al.
Publicado: (2024)
por: Vinod, Gautham, et al.
Publicado: (2024)
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
por: Vinod, Gautham, et al.
Publicado: (2026)
por: Vinod, Gautham, et al.
Publicado: (2026)
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
por: Vinod, Gautham, et al.
Publicado: (2026)
por: Vinod, Gautham, et al.
Publicado: (2026)
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
por: Ma, Jinge, et al.
Publicado: (2024)
por: Ma, Jinge, et al.
Publicado: (2024)
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
por: Vinod, Gautham, et al.
Publicado: (2026)
por: Vinod, Gautham, et al.
Publicado: (2026)
PixelBoost: Leveraging Brownian Motion for Realistic-Image Super-Resolution
por: Mishra, Aradhana, et al.
Publicado: (2025)
por: Mishra, Aradhana, et al.
Publicado: (2025)
Optimal Transcoding Resolution Prediction for Efficient Per-Title Bitrate Ladder Estimation
por: Yang, Jinhai, et al.
Publicado: (2024)
por: Yang, Jinhai, et al.
Publicado: (2024)
From Darkness to Detail: Frequency-Aware SSMs for Low-Light Vision
por: Adhikarla, Eashan, et al.
Publicado: (2024)
por: Adhikarla, Eashan, et al.
Publicado: (2024)
SurFITR: A Dataset for Surveillance Image Forgery Detection and Localisation
por: Wang, Qizhou, et al.
Publicado: (2026)
por: Wang, Qizhou, et al.
Publicado: (2026)
VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head Reconstruction
por: Rakesh, Vineet Kumar, et al.
Publicado: (2026)
por: Rakesh, Vineet Kumar, et al.
Publicado: (2026)
Towards Real-world Video Face Restoration: A New Benchmark
por: Chen, Ziyan, et al.
Publicado: (2024)
por: Chen, Ziyan, et al.
Publicado: (2024)
EVAN: Evolutional Video Streaming Adaptation via Neural Representation
por: Liu, Mufan, et al.
Publicado: (2024)
por: Liu, Mufan, et al.
Publicado: (2024)
Machine Perception-Driven Image Compression: A Layered Generative Approach
por: Zhang, Yuefeng, et al.
Publicado: (2023)
por: Zhang, Yuefeng, et al.
Publicado: (2023)
Change Detection Between Optical Remote Sensing Imagery and Map Data via Segment Anything Model (SAM)
por: Chen, Hongruixuan, et al.
Publicado: (2024)
por: Chen, Hongruixuan, et al.
Publicado: (2024)
Visual and Text Prompt Segmentation: A Novel Multi-Model Framework for Remote Sensing
por: Zi, Xing, et al.
Publicado: (2025)
por: Zi, Xing, et al.
Publicado: (2025)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
por: Liu, Tianyi, et al.
Publicado: (2025)
por: Liu, Tianyi, et al.
Publicado: (2025)
RAISE: Realness Assessment for Image Synthesis and Evaluation
por: Mukherjee, Aniruddha, et al.
Publicado: (2025)
por: Mukherjee, Aniruddha, et al.
Publicado: (2025)
CAMSIC: Content-aware Masked Image Modeling Transformer for Stereo Image Compression
por: Zhang, Xinjie, et al.
Publicado: (2024)
por: Zhang, Xinjie, et al.
Publicado: (2024)
Movie Trailer Genre Classification Using Multimodal Pretrained Features
por: Sulun, Serkan, et al.
Publicado: (2024)
por: Sulun, Serkan, et al.
Publicado: (2024)
GaussianImage: 1000 FPS Image Representation and Compression by 2D Gaussian Splatting
por: Zhang, Xinjie, et al.
Publicado: (2024)
por: Zhang, Xinjie, et al.
Publicado: (2024)
Event Camera Demosaicing via Swin Transformer and Pixel-focus Loss
por: Lu, Yunfan, et al.
Publicado: (2024)
por: Lu, Yunfan, et al.
Publicado: (2024)
Context and Pixel Aware Large Language Model for Video Quality Assessment
por: Wen, Wen, et al.
Publicado: (2025)
por: Wen, Wen, et al.
Publicado: (2025)
DrawVideo: Generating Long Video from Storyboard Keyframe Sketches
por: Xu, Chuanzhi, et al.
Publicado: (2026)
por: Xu, Chuanzhi, et al.
Publicado: (2026)
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
por: Roy, Prasun, et al.
Publicado: (2025)
por: Roy, Prasun, et al.
Publicado: (2025)
Tackle CSM in JPEG Steganalysis with Data Adaptation
por: Abecidan, Rony, et al.
Publicado: (2026)
por: Abecidan, Rony, et al.
Publicado: (2026)
Sphere-GAN: a GAN-based Approach for Saliency Estimation in 360° Videos
por: Wahba, Mahmoud Z. A., et al.
Publicado: (2025)
por: Wahba, Mahmoud Z. A., et al.
Publicado: (2025)
Image Quality Assessment: From Human to Machine Preference
por: Li, Chunyi, et al.
Publicado: (2025)
por: Li, Chunyi, et al.
Publicado: (2025)
From "What" to "How": Constrained Reasoning for Autoregressive Image Generation
por: Yan, Ruxue, et al.
Publicado: (2026)
por: Yan, Ruxue, et al.
Publicado: (2026)
Exploiting Frequency Correlation for Hyperspectral Image Reconstruction
por: Yan, Muge, et al.
Publicado: (2024)
por: Yan, Muge, et al.
Publicado: (2024)
Inter-Frame Compression for Dynamic Point Cloud Geometry Coding
por: Akhtar, Anique, et al.
Publicado: (2022)
por: Akhtar, Anique, et al.
Publicado: (2022)
MSNeRV: Neural Video Representation with Multi-Scale Feature Fusion
por: Zhu, Jun, et al.
Publicado: (2025)
por: Zhu, Jun, et al.
Publicado: (2025)
LinMU: Multimodal Understanding Made Linear
por: Wang, Hongjie, et al.
Publicado: (2026)
por: Wang, Hongjie, et al.
Publicado: (2026)
RAPNet: A Receptive-Field Adaptive Convolutional Neural Network for Pansharpening
por: Tang, Tao, et al.
Publicado: (2025)
por: Tang, Tao, et al.
Publicado: (2025)
Semantic-Aware Adaptive Video Streaming Using Latent Diffusion Models for Wireless Networks
por: Yan, Zijiang, et al.
Publicado: (2025)
por: Yan, Zijiang, et al.
Publicado: (2025)
Mamba-360: Survey of State Space Models as Transformer Alternative for Long Sequence Modelling: Methods, Applications, and Challenges
por: Patro, Badri Narayana, et al.
Publicado: (2024)
por: Patro, Badri Narayana, et al.
Publicado: (2024)
VEMOCLAP: A video emotion classification web application
por: Sulun, Serkan, et al.
Publicado: (2024)
por: Sulun, Serkan, et al.
Publicado: (2024)
Health AI Developer Foundations
por: Kiraly, Atilla P., et al.
Publicado: (2024)
por: Kiraly, Atilla P., et al.
Publicado: (2024)
Improving Multi-label Recognition using Class Co-Occurrence Probabilities
por: Rawlekar, Samyak, et al.
Publicado: (2024)
por: Rawlekar, Samyak, et al.
Publicado: (2024)
FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait
por: Ki, Taekyung, et al.
Publicado: (2024)
por: Ki, Taekyung, et al.
Publicado: (2024)
MIND: A Noise-Adaptive Denoising Framework for Medical Images Integrating Multi-Scale Transformer
por: Tang, Tao, et al.
Publicado: (2025)
por: Tang, Tao, et al.
Publicado: (2025)
Ejemplares similares
-
Food Portion Estimation via 3D Object Scaling
por: Vinod, Gautham, et al.
Publicado: (2024) -
DietDelta: A Vision-Language Approach for Dietary Assessment via Before-and-After Images
por: Vinod, Gautham, et al.
Publicado: (2026) -
Not Your Stereo-Typical Estimator: Combining Vision and Language for Volume Perception
por: Vinod, Gautham, et al.
Publicado: (2026) -
MFP3D: Monocular Food Portion Estimation Leveraging 3D Point Clouds
por: Ma, Jinge, et al.
Publicado: (2024) -
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
por: Vinod, Gautham, et al.
Publicado: (2026)