On the Efficacy of Text-Based Input Modalities for Action Anticipation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beedu, Apoorva, Haresamudram, Harish, Samel, Karan, Essa, Irfan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Efficient Foundational Multi-modal Models for Video Summarization
von: Samel, Karan, et al.
Veröffentlicht: (2024)
von: Samel, Karan, et al.
Veröffentlicht: (2024)
HierSum: A Global and Local Attention Mechanism for Video Summarization
von: Beedu, Apoorva, et al.
Veröffentlicht: (2025)
von: Beedu, Apoorva, et al.
Veröffentlicht: (2025)
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
von: Haresamudram, Harish, et al.
Veröffentlicht: (2024)
von: Haresamudram, Harish, et al.
Veröffentlicht: (2024)
Mamba Fusion: Learning Actions Through Questioning
von: Dong, Zhikang, et al.
Veröffentlicht: (2024)
von: Dong, Zhikang, et al.
Veröffentlicht: (2024)
GMAIL: Generative Modality Alignment for generated Image Learning
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
Multi-Modal Zero-Shot Prediction of Color Trajectories in Food Drying
von: Li, Shichen, et al.
Veröffentlicht: (2025)
von: Li, Shichen, et al.
Veröffentlicht: (2025)
A Survey of IMU Based Cross-Modal Transfer Learning in Human Activity Recognition
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
von: Kamboj, Abhi, et al.
Veröffentlicht: (2024)
Separated Inter/Intra-Modal Fusion Prompts for Compositional Zero-Shot Learning
von: Jung, Sua
Veröffentlicht: (2025)
von: Jung, Sua
Veröffentlicht: (2025)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
von: Mondal, Anindya, et al.
Veröffentlicht: (2023)
U-Net in Medical Image Segmentation: A Review of Its Applications Across Modalities
von: Neha, Fnu, et al.
Veröffentlicht: (2024)
von: Neha, Fnu, et al.
Veröffentlicht: (2024)
CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
von: Yu, Xinlei, et al.
Veröffentlicht: (2025)
SAW: Toward a Surgical Action World Model via Controllable and Scalable Video Generation
von: Rapuri, Sampath, et al.
Veröffentlicht: (2026)
von: Rapuri, Sampath, et al.
Veröffentlicht: (2026)
StreamDiT: Real-Time Streaming Text-to-Video Generation
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
von: Kodaira, Akio, et al.
Veröffentlicht: (2025)
ClusMFL: A Cluster-Enhanced Framework for Modality-Incomplete Multimodal Federated Learning in Brain Imaging Analysis
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
Disease Classification and Impact of Pretrained Deep Convolution Neural Networks on Diverse Medical Imaging Datasets across Imaging Modalities
von: Borah, Jutika, et al.
Veröffentlicht: (2024)
von: Borah, Jutika, et al.
Veröffentlicht: (2024)
Realism in Action: Anomaly-Aware Diagnosis of Brain Tumors from Medical Images Using YOLOv8 and DeiT
von: Hashemi, Seyed Mohammad Hossein, et al.
Veröffentlicht: (2024)
von: Hashemi, Seyed Mohammad Hossein, et al.
Veröffentlicht: (2024)
A Novel Framework For Text Detection From Natural Scene Images With Complex Background
von: Kaladagi, Basavaraj, et al.
Veröffentlicht: (2024)
von: Kaladagi, Basavaraj, et al.
Veröffentlicht: (2024)
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
von: Wang, Hongjie, et al.
Veröffentlicht: (2024)
Enhancing GANs with Contrastive Learning-Based Multistage Progressive Finetuning SNN and RL-Based External Optimization
von: Mustafa, Osama
Veröffentlicht: (2024)
von: Mustafa, Osama
Veröffentlicht: (2024)
Deep Learning-Based Classification of Hyperkinetic Movement Disorders in Children
von: Ramamurthy, Nandika, et al.
Veröffentlicht: (2024)
von: Ramamurthy, Nandika, et al.
Veröffentlicht: (2024)
Deep Learning-Based Automatic Diagnosis System for Developmental Dysplasia of the Hip
von: Li, Yang, et al.
Veröffentlicht: (2022)
von: Li, Yang, et al.
Veröffentlicht: (2022)
From CNNs to Shift-Invariant Twin Models Based on Complex Wavelets
von: Leterme, Hubert, et al.
Veröffentlicht: (2022)
von: Leterme, Hubert, et al.
Veröffentlicht: (2022)
Efficient Cutting Tool Wear Segmentation Based on Segment Anything Model
von: Li, Zongshuo, et al.
Veröffentlicht: (2024)
von: Li, Zongshuo, et al.
Veröffentlicht: (2024)
Direct Video-Based Spatiotemporal Deep Learning for Cattle Lameness Detection
von: Sohan, Md Fahimuzzman, et al.
Veröffentlicht: (2025)
von: Sohan, Md Fahimuzzman, et al.
Veröffentlicht: (2025)
S$^3$F-Net: A Multi-Modal Approach to Medical Image Classification via Spatial-Spectral Summarizer Fusion Network
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
von: Siddiqui, Md. Saiful Bari, et al.
Veröffentlicht: (2025)
Vision Transformer-Based Time-Series Image Reconstruction for Cloud-Filling Applications
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
Novel AI-Based Quantification of Breast Arterial Calcification to Predict Cardiovascular Risk
von: Dapamede, Theodorus, et al.
Veröffentlicht: (2025)
von: Dapamede, Theodorus, et al.
Veröffentlicht: (2025)
Unit-Based Histopathology Tissue Segmentation via Multi-Level Feature Representation
von: Shakarami, Ashkan, et al.
Veröffentlicht: (2025)
von: Shakarami, Ashkan, et al.
Veröffentlicht: (2025)
DepViT-CAD: Deployable Vision Transformer-Based Cancer Diagnosis in Histopathology
von: Shakarami, Ashkan, et al.
Veröffentlicht: (2025)
von: Shakarami, Ashkan, et al.
Veröffentlicht: (2025)
Random Walks with Tweedie: A Unified View of Score-Based Diffusion Models
von: Park, Chicago Y., et al.
Veröffentlicht: (2024)
von: Park, Chicago Y., et al.
Veröffentlicht: (2024)
Generative Model-Based Fusion for Improved Few-Shot Semantic Segmentation of Infrared Images
von: Yun, Junno, et al.
Veröffentlicht: (2024)
von: Yun, Junno, et al.
Veröffentlicht: (2024)
Automated Web-Based Malaria Detection System with Machine Learning and Deep Learning Techniques
von: Taye, Abraham G, et al.
Veröffentlicht: (2024)
von: Taye, Abraham G, et al.
Veröffentlicht: (2024)
A Novel Momentum-Based Deep Learning Techniques for Medical Image Classification and Segmentation
von: Biswas, Koushik, et al.
Veröffentlicht: (2024)
von: Biswas, Koushik, et al.
Veröffentlicht: (2024)
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
von: Huang, Haiwen, et al.
Veröffentlicht: (2025)
von: Huang, Haiwen, et al.
Veröffentlicht: (2025)
Efficient Edge-Compatible CNN for Speckle-Based Material Recognition in Laser Cutting Systems
von: Salem, Mohamed Abdallah, et al.
Veröffentlicht: (2025)
von: Salem, Mohamed Abdallah, et al.
Veröffentlicht: (2025)
A Deep Convolutional Neural Network-Based Novel Class Balancing for Imbalance Data Segmentation
von: Kalsoom, Atifa, et al.
Veröffentlicht: (2025)
von: Kalsoom, Atifa, et al.
Veröffentlicht: (2025)
Drone-Based Multispectral Imaging and Deep Learning for Timely Detection of Branched Broomrape in Tomato Farms
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Narimani, Mohammadreza, et al.
Veröffentlicht: (2025)
An Enhanced Classification Method Based on Adaptive Multi-Scale Fusion for Long-tailed Multispectral Point Clouds
von: Liu, TianZhu, et al.
Veröffentlicht: (2024)
von: Liu, TianZhu, et al.
Veröffentlicht: (2024)
Deep Learning-Based Breast Cancer Detection in Mammography: A Multi-Center Validation Study in Thai Population
von: Chamveha, Isarun, et al.
Veröffentlicht: (2025)
von: Chamveha, Isarun, et al.
Veröffentlicht: (2025)
Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation
von: Wang, Wenqing, et al.
Veröffentlicht: (2024)
von: Wang, Wenqing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploring Efficient Foundational Multi-modal Models for Video Summarization
von: Samel, Karan, et al.
Veröffentlicht: (2024) -
HierSum: A Global and Local Attention Mechanism for Video Summarization
von: Beedu, Apoorva, et al.
Veröffentlicht: (2025) -
Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition -- And Ways to Overcome Them
von: Haresamudram, Harish, et al.
Veröffentlicht: (2024) -
Mamba Fusion: Learning Actions Through Questioning
von: Dong, Zhikang, et al.
Veröffentlicht: (2024) -
GMAIL: Generative Modality Alignment for generated Image Learning
von: Mo, Shentong, et al.
Veröffentlicht: (2026)