CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Dongyu, Yao, Keling, Zhou, Junhong, Zhang, Yinghao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
by: Marmoret, Axel, et al.
Published: (2025)
by: Marmoret, Axel, et al.
Published: (2025)
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
by: Gad, Eyad, et al.
Published: (2025)
by: Gad, Eyad, et al.
Published: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
by: Baydemir, Poyraz
Published: (2025)
by: Baydemir, Poyraz
Published: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025)
by: Viveiros, André G., et al.
Published: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)
by: Svystun, Serhii, et al.
Published: (2025)
AC-LoRA: Auto Component LoRA for Personalized Artistic Style Image Generation
by: Cui, Zhipu, et al.
Published: (2025)
by: Cui, Zhipu, et al.
Published: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
by: Komurcu, Kursat, et al.
Published: (2026)
by: Komurcu, Kursat, et al.
Published: (2026)
IDOL: Instant Photorealistic 3D Human Creation from a Single Image
by: Zhuang, Yiyu, et al.
Published: (2024)
by: Zhuang, Yiyu, et al.
Published: (2024)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
by: Shekar, Pavan C, et al.
Published: (2025)
by: Shekar, Pavan C, et al.
Published: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
by: Acuaviva, Pablo, et al.
Published: (2025)
by: Acuaviva, Pablo, et al.
Published: (2025)
OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
by: Zhao, Weiyi, et al.
Published: (2025)
by: Zhao, Weiyi, et al.
Published: (2025)
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines
by: Wimalasiri, Chathura
Published: (2026)
by: Wimalasiri, Chathura
Published: (2026)
RG-TTA: Regime-Guided Meta-Control for Test-Time Adaptation in Streaming Time Series
by: Kumar, Indar, et al.
Published: (2026)
by: Kumar, Indar, et al.
Published: (2026)
A Landmark-Aware Visual Navigation Dataset
by: Johnson, Faith, et al.
Published: (2024)
by: Johnson, Faith, et al.
Published: (2024)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
by: Kushal, Koushik Ahmed, et al.
Published: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
by: Siddiqui, Yousuf Ahmed, et al.
Published: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
Enhancing Maritime Object Detection in Real-Time with RT-DETR and Data Augmentation
by: Nemati, Nader
Published: (2025)
by: Nemati, Nader
Published: (2025)
Structured Basis Function Networks: Loss-Centric Multi-Hypothesis Ensembles with Controllable Diversity
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
by: Dominguez, Alejandro Rodriguez, et al.
Published: (2025)
MRI Brain Tumor Detection with Computer Vision
by: Krolik, Jack, et al.
Published: (2025)
by: Krolik, Jack, et al.
Published: (2025)
Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
by: Tu, Songjun, et al.
Published: (2025)
by: Tu, Songjun, et al.
Published: (2025)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
by: Zhang, Haoran, et al.
Published: (2026)
by: Zhang, Haoran, et al.
Published: (2026)
Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping
by: Lai, Guannan, et al.
Published: (2025)
by: Lai, Guannan, et al.
Published: (2025)
Geometric-Stochastic Multimodal Deep Learning for Predictive Modeling of SUDEP and Stroke Vulnerability
by: Girish, Preksha, et al.
Published: (2025)
by: Girish, Preksha, et al.
Published: (2025)
Multi-Scale Graph Learning for Anti-Sparse Downscaling
by: Fan, Yingda, et al.
Published: (2025)
by: Fan, Yingda, et al.
Published: (2025)
To Whom are You Talking? A Deep Learning Model to Endow Social Robots with Addressee Estimation Skills
by: Mazzola, Carlo, et al.
Published: (2023)
by: Mazzola, Carlo, et al.
Published: (2023)
VALE: A Multimodal Visual and Language Explanation Framework for Image Classifiers using eXplainable AI and Language Models
by: Natarajan, Purushothaman, et al.
Published: (2024)
by: Natarajan, Purushothaman, et al.
Published: (2024)
Semantically Guided Adversarial Testing of Vision Models Using Language Models
by: Filus, Katarzyna, et al.
Published: (2025)
by: Filus, Katarzyna, et al.
Published: (2025)
GLL: A Differentiable Graph Learning Layer for Neural Networks
by: Brown, Jason, et al.
Published: (2024)
by: Brown, Jason, et al.
Published: (2024)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
by: Kambhatla, Akhila, et al.
Published: (2025)
by: Kambhatla, Akhila, et al.
Published: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
by: Dong, Haohua, et al.
Published: (2025)
by: Dong, Haohua, et al.
Published: (2025)
Parameter-efficient fine-tuning (PEFT) of Vision Foundation Models for Atypical Mitotic Figure Classification
by: Ramchandani, Lavish, et al.
Published: (2025)
by: Ramchandani, Lavish, et al.
Published: (2025)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
by: Alanazi, Ahmed, et al.
Published: (2025)
by: Alanazi, Ahmed, et al.
Published: (2025)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
by: Kalkanli, Beyza, et al.
Published: (2026)
by: Kalkanli, Beyza, et al.
Published: (2026)
A Hybrid Multimodal Deep Learning Framework for Intelligent Fashion Recommendation
by: Kalashi, Kamand, et al.
Published: (2025)
by: Kalashi, Kamand, et al.
Published: (2025)
Think, Act, Learn: A Framework for Autonomous Robotic Agents using Closed-Loop Large Language Models
by: Menon, Anjali R., et al.
Published: (2025)
by: Menon, Anjali R., et al.
Published: (2025)
DRL: Discriminative Representation Learning with Parallel Adapters for Class Incremental Learning
by: Zhan, Jiawei, et al.
Published: (2025)
by: Zhan, Jiawei, et al.
Published: (2025)
TexTile: A Differentiable Metric for Texture Tileability
by: Rodriguez-Pardo, Carlos, et al.
Published: (2024)
by: Rodriguez-Pardo, Carlos, et al.
Published: (2024)
Similar Items
-
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
by: Marmoret, Axel, et al.
Published: (2025) -
Advancing Brain Tumor Segmentation via Attention-based 3D U-Net Architecture and Digital Image Processing
by: Gad, Eyad, et al.
Published: (2025) -
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
by: Baydemir, Poyraz
Published: (2025) -
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
by: Viveiros, André G., et al.
Published: (2025) -
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
by: Svystun, Serhii, et al.
Published: (2025)