LLaVAction: evaluating and training multi-modal large language models for action understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Qi, Haozhe, Ye, Shaokai, Mathis, Alexander, Mathis, Mackenzie W. |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SuperAnimal pretrained pose estimation models for behavioral analysis
par: Ye, Shaokai, et autres
Publié: (2022)
par: Ye, Shaokai, et autres
Publié: (2022)
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
par: Qi, Haozhe, et autres
Publié: (2024)
par: Qi, Haozhe, et autres
Publié: (2024)
Adversarially Robust Out-of-Distribution Detection Using Lyapunov-Stabilized Embeddings
par: Mirzaei, Hossein, et autres
Publié: (2024)
par: Mirzaei, Hossein, et autres
Publié: (2024)
PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation
par: Yu, Xiaohang, et autres
Publié: (2026)
par: Yu, Xiaohang, et autres
Publié: (2026)
FMPose3D: monocular 3D pose estimation via flow matching
par: Wang, Ti, et autres
Publié: (2026)
par: Wang, Ti, et autres
Publié: (2026)
Segment anything model for head and neck tumor segmentation with CT, PET and MRI multi-modality images
par: Ren, Jintao, et autres
Publié: (2024)
par: Ren, Jintao, et autres
Publié: (2024)
MammAlps: A multi-view video behavior monitoring dataset of wild mammals in the Swiss Alps
par: Gabeff, Valentin, et autres
Publié: (2025)
par: Gabeff, Valentin, et autres
Publié: (2025)
AdaptToken: Entropy-based Adaptive Token Selection for MLLM Long Video Understanding
par: Qi, Haozhe, et autres
Publié: (2026)
par: Qi, Haozhe, et autres
Publié: (2026)
Robust image classification with multi-modal large language models
par: Villani, Francesco, et autres
Publié: (2024)
par: Villani, Francesco, et autres
Publié: (2024)
Do large language vision models understand 3D shapes?
par: Eppel, Sagi
Publié: (2024)
par: Eppel, Sagi
Publié: (2024)
DISTIL: Data-Free Inversion of Suspicious Trojan Inputs via Latent Diffusion
par: Mirzaei, Hossein, et autres
Publié: (2025)
par: Mirzaei, Hossein, et autres
Publié: (2025)
LOOC: Localizing Organs using Occupancy Networks and Body Surface Depth Images
par: Henrich, Pit, et autres
Publié: (2024)
par: Henrich, Pit, et autres
Publié: (2024)
A multi-modal vision-language model for generalizable annotation-free pathology localization
par: Yang, Hao, et autres
Publié: (2024)
par: Yang, Hao, et autres
Publié: (2024)
Multi-Flow: Multi-View-Enriched Normalizing Flows for Industrial Anomaly Detection
par: Kruse, Mathis, et autres
Publié: (2025)
par: Kruse, Mathis, et autres
Publié: (2025)
EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models
par: Bonnetto, Andy, et autres
Publié: (2025)
par: Bonnetto, Andy, et autres
Publié: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
par: Bensabath, Léore, et autres
Publié: (2024)
par: Bensabath, Léore, et autres
Publié: (2024)
BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection
par: Schween, Melissa, et autres
Publié: (2026)
par: Schween, Melissa, et autres
Publié: (2026)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
par: Wu, Yiqi, et autres
Publié: (2024)
par: Wu, Yiqi, et autres
Publié: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
par: Liu, Shuliang, et autres
Publié: (2025)
par: Liu, Shuliang, et autres
Publié: (2025)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
par: Nassar, Ahmed, et autres
Publié: (2025)
par: Nassar, Ahmed, et autres
Publié: (2025)
PosterLLaVa: Constructing a Unified Multi-modal Layout Generator with LLM
par: Yang, Tao, et autres
Publié: (2024)
par: Yang, Tao, et autres
Publié: (2024)
DeViDe: Faceted medical knowledge for improved medical vision-language pre-training
par: Luo, Haozhe, et autres
Publié: (2024)
par: Luo, Haozhe, et autres
Publié: (2024)
Multi-modal Speech Emotion Recognition via Feature Distribution Adaptation Network
par: Li, Shaokai, et autres
Publié: (2024)
par: Li, Shaokai, et autres
Publié: (2024)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
par: Xu, Lijian, et autres
Publié: (2024)
par: Xu, Lijian, et autres
Publié: (2024)
Sample-Specific Output Constraints for Neural Networks
par: Brosowsky, Mathis, et autres
Publié: (2020)
par: Brosowsky, Mathis, et autres
Publié: (2020)
CarLLaVA: Vision language models for camera-only closed-loop driving
par: Renz, Katrin, et autres
Publié: (2024)
par: Renz, Katrin, et autres
Publié: (2024)
Redundancy-Aware Pretraining of Vision-Language Foundation Models in Remote Sensing
par: Adler, Mathis Jürgen, et autres
Publié: (2025)
par: Adler, Mathis Jürgen, et autres
Publié: (2025)
Comprehensive language-image pre-training for 3D medical image understanding
par: Wald, Tassilo, et autres
Publié: (2025)
par: Wald, Tassilo, et autres
Publié: (2025)
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
par: Athanasiou, Nikos, et autres
Publié: (2023)
par: Athanasiou, Nikos, et autres
Publié: (2023)
Grid-augmented vision: A simple yet effective approach for enhanced spatial understanding in multi-modal agents
par: Chae, Joongwon, et autres
Publié: (2024)
par: Chae, Joongwon, et autres
Publié: (2024)
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models
par: Liang, Yinan, et autres
Publié: (2025)
par: Liang, Yinan, et autres
Publié: (2025)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
par: Shi, Danli, et autres
Publié: (2024)
par: Shi, Danli, et autres
Publié: (2024)
LUDO: Low-Latency Understanding of Deformable Objects using Point Cloud Occupancy Functions
par: Henrich, Pit, et autres
Publié: (2024)
par: Henrich, Pit, et autres
Publié: (2024)
bi-modal textual prompt learning for vision-language models in remote sensing
par: Kashyap, Pankhi, et autres
Publié: (2026)
par: Kashyap, Pankhi, et autres
Publié: (2026)
SplatPose & Detect: Pose-Agnostic 3D Anomaly Detection
par: Kruse, Mathis, et autres
Publié: (2024)
par: Kruse, Mathis, et autres
Publié: (2024)
Synthesizing and Identifying Noise Levels in Autonomous Vehicle Camera Radar Datasets
par: Morales, Mathis, et autres
Publié: (2025)
par: Morales, Mathis, et autres
Publié: (2025)
SYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4
par: Mayer, Kevin, et autres
Publié: (2025)
par: Mayer, Kevin, et autres
Publié: (2025)
Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models
par: Malik, Hashmat Shadab, et autres
Publié: (2025)
par: Malik, Hashmat Shadab, et autres
Publié: (2025)
UAV traffic scene understanding: A regulation embedded multi-modal network and a unified benchmark
par: Zhang, Yu, et autres
Publié: (2026)
par: Zhang, Yu, et autres
Publié: (2026)
Building and better understanding vision-language models: insights and future directions
par: Laurençon, Hugo, et autres
Publié: (2024)
par: Laurençon, Hugo, et autres
Publié: (2024)
Documents similaires
-
SuperAnimal pretrained pose estimation models for behavioral analysis
par: Ye, Shaokai, et autres
Publié: (2022) -
HOISDF: Constraining 3D Hand-Object Pose Estimation with Global Signed Distance Fields
par: Qi, Haozhe, et autres
Publié: (2024) -
Adversarially Robust Out-of-Distribution Detection Using Lyapunov-Stabilized Embeddings
par: Mirzaei, Hossein, et autres
Publié: (2024) -
PRIMA: Boosting Animal Mesh Recovery with Biological Priors and Test-Time Adaptation
par: Yu, Xiaohang, et autres
Publié: (2026) -
FMPose3D: monocular 3D pose estimation via flow matching
par: Wang, Ti, et autres
Publié: (2026)