From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Suyash, Li, Qiang, Patil, Srikanth, Girdhar, Anubhav |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Are Two LLMs Better Than One? A Student-Teacher Dual-Head LLMs Architecture for Pharmaceutical Content Optimization
by: Mishra, Suyash, et al.
Published: (2026)
by: Mishra, Suyash, et al.
Published: (2026)
Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision
by: Borodin, Kirill, et al.
Published: (2026)
by: Borodin, Kirill, et al.
Published: (2026)
Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
by: Zheng, Shunjie-Fabian, et al.
Published: (2025)
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)
by: Faraz, Ali, et al.
Published: (2025)
Do VLMs Need Vision Transformers? Evaluating State Space Models as Vision Encoders
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
by: Kuo, Shang-Jui Ray, et al.
Published: (2026)
Understanding Task Transfer in Vision-Language Models
by: Sachdeva, Bhuvan, et al.
Published: (2025)
by: Sachdeva, Bhuvan, et al.
Published: (2025)
ClipGrader: Leveraging Vision-Language Models for Robust Label Quality Assessment in Object Detection
by: Lu, Hong, et al.
Published: (2025)
by: Lu, Hong, et al.
Published: (2025)
Focus Your Attention: Towards Data-Intuitive Lightweight Vision Transformers
by: Gaurav, Suyash, et al.
Published: (2025)
by: Gaurav, Suyash, et al.
Published: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
Probing Visual Language Priors in VLMs
by: Luo, Tiange, et al.
Published: (2024)
by: Luo, Tiange, et al.
Published: (2024)
Shaking Up VLMs: Comparing Transformers and Structured State Space Models for Vision & Language Modeling
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Towards Understanding How Knowledge Evolves in Large Vision-Language Models
by: Wang, Sudong, et al.
Published: (2025)
by: Wang, Sudong, et al.
Published: (2025)
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
by: Kataria, Anubhav, et al.
Published: (2025)
by: Kataria, Anubhav, et al.
Published: (2025)
Yo'LLaVA: Your Personalized Language and Vision Assistant
by: Nguyen, Thao, et al.
Published: (2024)
by: Nguyen, Thao, et al.
Published: (2024)
Feature Clipping for Uncertainty Calibration
by: Tao, Linwei, et al.
Published: (2024)
by: Tao, Linwei, et al.
Published: (2024)
Understanding and Rectifying Safety Perception Distortion in VLMs
by: Zou, Xiaohan, et al.
Published: (2025)
by: Zou, Xiaohan, et al.
Published: (2025)
Personalized Vision via Visual In-Context Learning
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
by: Le, Anjie, et al.
Published: (2025)
by: Le, Anjie, et al.
Published: (2025)
TechING: Towards Real World Technical Image Understanding via VLMs
by: Nadeem, Tafazzul, et al.
Published: (2026)
by: Nadeem, Tafazzul, et al.
Published: (2026)
Composition Vision-Language Understanding via Segment and Depth Anything Model
by: Huo, Mingxiao, et al.
Published: (2024)
by: Huo, Mingxiao, et al.
Published: (2024)
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
How Much Information Can a Vision Token Hold? A Scaling Law for Recognition Limits in VLMs
by: Zhuang, Shuxin, et al.
Published: (2026)
by: Zhuang, Shuxin, et al.
Published: (2026)
Dynamic Scene Understanding from Vision-Language Representations
by: Pruss, Shahaf, et al.
Published: (2025)
by: Pruss, Shahaf, et al.
Published: (2025)
Spectrum Extraction and Clipping for Implicitly Linear Layers
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
by: Boroojeny, Ali Ebrahimpour, et al.
Published: (2024)
Small Vision-Language Models are Smart Compressors for Long Video Understanding
by: Fei, Junjie, et al.
Published: (2026)
by: Fei, Junjie, et al.
Published: (2026)
Do Vision--Language Models Understand 3D Scenes or Just Catalogue Objects?
by: Maheshwari, Animesh, et al.
Published: (2026)
by: Maheshwari, Animesh, et al.
Published: (2026)
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
AnimateLCM: Computation-Efficient Personalized Style Video Generation without Personalized Video Data
by: Wang, Fu-Yun, et al.
Published: (2024)
by: Wang, Fu-Yun, et al.
Published: (2024)
Continual Learning with Vision-Language Models via Semantic-Geometry Preservation
by: He, Chiyuan, et al.
Published: (2026)
by: He, Chiyuan, et al.
Published: (2026)
Online Video Understanding: OVBench and VideoChat-Online
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
MCAT: Visual Query-Based Localization of Standard Anatomical Clips in Fetal Ultrasound Videos Using Multi-Tier Class-Aware Token Transformer
by: Mishra, Divyanshu, et al.
Published: (2025)
by: Mishra, Divyanshu, et al.
Published: (2025)
MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data
by: Li, Zongxia, et al.
Published: (2026)
by: Li, Zongxia, et al.
Published: (2026)
HourVideo: 1-Hour Video-Language Understanding
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
by: Chandrasegaran, Keshigeyan, et al.
Published: (2024)
EchoPrime: A Multi-Video View-Informed Vision-Language Model for Comprehensive Echocardiography Interpretation
by: Vukadinovic, Milos, et al.
Published: (2024)
by: Vukadinovic, Milos, et al.
Published: (2024)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
by: Ghiasvand, Sajjad, et al.
Published: (2025)
by: Ghiasvand, Sajjad, et al.
Published: (2025)
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
by: Li, Chunxiao, et al.
Published: (2026)
by: Li, Chunxiao, et al.
Published: (2026)
Similar Items
-
Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
by: Mishra, Suyash, et al.
Published: (2026) -
Are Two LLMs Better Than One? A Student-Teacher Dual-Head LLMs Architecture for Pharmaceutical Content Optimization
by: Mishra, Suyash, et al.
Published: (2026) -
Benchmarking Compact VLMs for Clip-Level Surveillance Anomaly Detection Under Weak Supervision
by: Borodin, Kirill, et al.
Published: (2026) -
Breast Cancer VLMs: Clinically Practical Vision-Language Train-Inference Models
by: Zheng, Shunjie-Fabian, et al.
Published: (2025) -
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
by: Faraz, Ali, et al.
Published: (2025)