Can Local Vision-Language Models improve Activity Recognition over Vision Transformers? -- Case Study on Newborn Resuscitation
Fuente:
arXiv
Guardado en:
| Autores principales: | Guerriero, Enrico, Engan, Kjersti, Meinich-Bache, Øyvind |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
por: Su, Yuetong, et al.
Publicado: (2025)
por: Su, Yuetong, et al.
Publicado: (2025)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
por: Chen, Jingkun, et al.
Publicado: (2025)
por: Chen, Jingkun, et al.
Publicado: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025)
por: Viveiros, André G., et al.
Publicado: (2025)
Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
por: Ma, Siteng, et al.
Publicado: (2025)
por: Ma, Siteng, et al.
Publicado: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
por: Zinnen, Mathias, et al.
Publicado: (2025)
por: Zinnen, Mathias, et al.
Publicado: (2025)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
por: Gautam, Sushant, et al.
Publicado: (2025)
por: Gautam, Sushant, et al.
Publicado: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
por: Agarwal, Amit, et al.
Publicado: (2025)
por: Agarwal, Amit, et al.
Publicado: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
por: Gautam, Sushant, et al.
Publicado: (2025)
por: Gautam, Sushant, et al.
Publicado: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
por: Farzulla, Murad
Publicado: (2026)
por: Farzulla, Murad
Publicado: (2026)
MATEX: Multi-scale Attention and Text-guided Explainability of Medical Vision-Language Models
por: Imran, Muhammad, et al.
Publicado: (2026)
por: Imran, Muhammad, et al.
Publicado: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
por: Fu, Tianyu, et al.
Publicado: (2024)
por: Fu, Tianyu, et al.
Publicado: (2024)
VDPP: Video Depth Post-Processing for Speed and Scalability
por: Yoon, Daewon, et al.
Publicado: (2026)
por: Yoon, Daewon, et al.
Publicado: (2026)
Automating Timed Up and Go Phase Segmentation and Gait Analysis via the tugturn Markerless 3D Pipeline
por: Chinaglia, Abel Gonçalves, et al.
Publicado: (2026)
por: Chinaglia, Abel Gonçalves, et al.
Publicado: (2026)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
por: Qesaraku, Bjorna, et al.
Publicado: (2025)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
por: Yang, Baoyao, et al.
Publicado: (2025)
por: Yang, Baoyao, et al.
Publicado: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
por: Jin, Xiaofeng, et al.
Publicado: (2025)
por: Jin, Xiaofeng, et al.
Publicado: (2025)
DSER: Spectral Epipolar Representation for Efficient Light Field Depth Estimation
por: Mohammad, Noor Islam S., et al.
Publicado: (2025)
por: Mohammad, Noor Islam S., et al.
Publicado: (2025)
Hierarchical Spatial Algorithms for High-Resolution Image Quantization and Feature Extraction
por: Mohammad, Noor Islam S.
Publicado: (2025)
por: Mohammad, Noor Islam S.
Publicado: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
por: Yim, Wen-wai, et al.
Publicado: (2025)
por: Yim, Wen-wai, et al.
Publicado: (2025)
3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model
por: Ko, Hyun-kyu, et al.
Publicado: (2026)
por: Ko, Hyun-kyu, et al.
Publicado: (2026)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
por: Dumpala, Sri Harsha, et al.
Publicado: (2024)
UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking
por: Perez, Joan, et al.
Publicado: (2026)
por: Perez, Joan, et al.
Publicado: (2026)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
por: Rashid, Muhammad, et al.
Publicado: (2026)
por: Rashid, Muhammad, et al.
Publicado: (2026)
A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
por: Nazeri, Ali, et al.
Publicado: (2025)
por: Nazeri, Ali, et al.
Publicado: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
por: Zhang, Junbin, et al.
Publicado: (2022)
por: Zhang, Junbin, et al.
Publicado: (2022)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
por: Patel, Hitesh Laxmichand, et al.
Publicado: (2025)
Technical Report: Automated Optical Inspection of Surgical Instruments
por: Shafqat, Zunaira, et al.
Publicado: (2026)
por: Shafqat, Zunaira, et al.
Publicado: (2026)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
por: Chahine, Makram, et al.
Publicado: (2024)
por: Chahine, Makram, et al.
Publicado: (2024)
Leonardo vindicated: Pythagorean trees for minimal reconstruction of the natural branching structures
por: Ruta, Dymitr, et al.
Publicado: (2024)
por: Ruta, Dymitr, et al.
Publicado: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
Do Generative Metrics Predict YOLO Performance? An Evaluation Across Models, Augmentation Ratios, and Dataset Complexity
por: Marian, Vasile, et al.
Publicado: (2026)
por: Marian, Vasile, et al.
Publicado: (2026)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
por: Lyu, Wenbo, et al.
Publicado: (2025)
por: Lyu, Wenbo, et al.
Publicado: (2025)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
por: Jung, Seoik, et al.
Publicado: (2025)
por: Jung, Seoik, et al.
Publicado: (2025)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
por: Gkountouras, John, et al.
Publicado: (2025)
por: Gkountouras, John, et al.
Publicado: (2025)
Gr-IoU: Ground-Intersection over Union for Robust Multi-Object Tracking with 3D Geometric Constraints
por: Toida, Keisuke, et al.
Publicado: (2024)
por: Toida, Keisuke, et al.
Publicado: (2024)
A Survey on Vision-Language-Action Models for Embodied AI
por: Ma, Yueen, et al.
Publicado: (2024)
por: Ma, Yueen, et al.
Publicado: (2024)
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning
por: Patel, Urjitkumar, et al.
Publicado: (2025)
por: Patel, Urjitkumar, et al.
Publicado: (2025)
A large-scale, physically-based synthetic dataset for satellite pose estimation
por: Velkei, Szabolcs, et al.
Publicado: (2025)
por: Velkei, Szabolcs, et al.
Publicado: (2025)
Ejemplares similares
-
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
por: Su, Yuetong, et al.
Publicado: (2025) -
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
por: Yanambakkam, Hemanth Teja, et al.
Publicado: (2025) -
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
por: Chen, Jingkun, et al.
Publicado: (2025) -
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
por: Viveiros, André G., et al.
Publicado: (2025) -
Deep Learning Approaches for Medical Imaging Under Varying Degrees of Label Availability: A Comprehensive Survey
por: Ma, Siteng, et al.
Publicado: (2025)