Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Jingyi, Yu, Zitong, Ni, Xiuming, He, Jia, Li, Hui |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
G$^2$V$^2$former: Graph Guided Video Vision Transformer for Face Anti-Spoofing
par: Yang, Jingyi, et autres
Publié: (2024)
par: Yang, Jingyi, et autres
Publié: (2024)
Generalized Face Anti-spoofing via Finer Domain Partition and Disentangling Liveness-irrelevant Factors
par: Yang, Jingyi, et autres
Publié: (2024)
par: Yang, Jingyi, et autres
Publié: (2024)
Language Model Guided Interpretable Video Action Reasoning
par: Wang, Ning, et autres
Publié: (2024)
par: Wang, Ning, et autres
Publié: (2024)
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
par: Wang, Haoyu, et autres
Publié: (2026)
par: Wang, Haoyu, et autres
Publié: (2026)
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
par: Jin, Qiaoqiao, et autres
Publié: (2024)
par: Jin, Qiaoqiao, et autres
Publié: (2024)
Masked Autoencoders are Parameter-Efficient Federated Continual Learners
par: He, Yuchen, et autres
Publié: (2024)
par: He, Yuchen, et autres
Publié: (2024)
Compositional Kronecker Context Optimization for Vision-Language Models
par: Ding, Kun, et autres
Publié: (2024)
par: Ding, Kun, et autres
Publié: (2024)
ActionVOS: Actions as Prompts for Video Object Segmentation
par: Ouyang, Liangyang, et autres
Publié: (2024)
par: Ouyang, Liangyang, et autres
Publié: (2024)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
par: Lin, Kun-Yu, et autres
Publié: (2024)
par: Lin, Kun-Yu, et autres
Publié: (2024)
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
par: Zhu, Yijie, et autres
Publié: (2026)
par: Zhu, Yijie, et autres
Publié: (2026)
Video Diffusion Transformers are In-Context Learners
par: Fei, Zhengcong, et autres
Publié: (2024)
par: Fei, Zhengcong, et autres
Publié: (2024)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
par: Liu, Xin, et autres
Publié: (2024)
par: Liu, Xin, et autres
Publié: (2024)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
par: Ni, Jingcheng, et autres
Publié: (2025)
par: Ni, Jingcheng, et autres
Publié: (2025)
DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing
par: Yang, Jingyi, et autres
Publié: (2025)
par: Yang, Jingyi, et autres
Publié: (2025)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
par: Zhang, Yaning, et autres
Publié: (2026)
par: Zhang, Yaning, et autres
Publié: (2026)
Contrastive Masked Autoencoders are Stronger Vision Learners
par: Huang, Zhicheng, et autres
Publié: (2022)
par: Huang, Zhicheng, et autres
Publié: (2022)
Masked Diffusion as Self-supervised Representation Learner
par: Pan, Zixuan, et autres
Publié: (2023)
par: Pan, Zixuan, et autres
Publié: (2023)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
par: Peng, Jingwei, et autres
Publié: (2025)
par: Peng, Jingwei, et autres
Publié: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
par: Zhang, Mingfang, et autres
Publié: (2024)
par: Zhang, Mingfang, et autres
Publié: (2024)
Precise Action-to-Video Generation Through Visual Action Prompts
par: Wang, Yuang, et autres
Publié: (2025)
par: Wang, Yuang, et autres
Publié: (2025)
EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
par: Chen, Zhe, et autres
Publié: (2024)
par: Chen, Zhe, et autres
Publié: (2024)
Diffusion Models as Masked Audio-Video Learners
par: Nunez, Elvis, et autres
Publié: (2023)
par: Nunez, Elvis, et autres
Publié: (2023)
AULLM++: Structural Reasoning with Large Language Models for Micro-Expression Recognition
par: Liu, Zhishu, et autres
Publié: (2026)
par: Liu, Zhishu, et autres
Publié: (2026)
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
par: Yan, Weicai, et autres
Publié: (2025)
par: Yan, Weicai, et autres
Publié: (2025)
MARMOT: Masked Autoencoder for Modeling Transient Imaging
par: Shen, Siyuan, et autres
Publié: (2025)
par: Shen, Siyuan, et autres
Publié: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
par: Fei, Jiajun, et autres
Publié: (2024)
par: Fei, Jiajun, et autres
Publié: (2024)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
par: Kang, Taewon, et autres
Publié: (2025)
par: Kang, Taewon, et autres
Publié: (2025)
Masked Diffusion Vision-Language Models for Temporal Action Localization
par: Wang, Fengshun, et autres
Publié: (2026)
par: Wang, Fengshun, et autres
Publié: (2026)
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
par: Liu, Chao, et autres
Publié: (2024)
par: Liu, Chao, et autres
Publié: (2024)
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
par: Li, Yu-Jhe, et autres
Publié: (2024)
par: Li, Yu-Jhe, et autres
Publié: (2024)
Large Language Models are Good Prompt Learners for Low-Shot Image Classification
par: Zheng, Zhaoheng, et autres
Publié: (2023)
par: Zheng, Zhaoheng, et autres
Publié: (2023)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
par: Wang, Zanyi, et autres
Publié: (2025)
par: Wang, Zanyi, et autres
Publié: (2025)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
par: Lai, Yuxiang, et autres
Publié: (2025)
par: Lai, Yuxiang, et autres
Publié: (2025)
Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
par: Huang, Zhiqi, et autres
Publié: (2024)
par: Huang, Zhiqi, et autres
Publié: (2024)
Deep Kronecker Network
par: Feng, Long, et autres
Publié: (2022)
par: Feng, Long, et autres
Publié: (2022)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
par: Wang, Yubin, et autres
Publié: (2024)
par: Wang, Yubin, et autres
Publié: (2024)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
par: Li, Kaining, et autres
Publié: (2025)
par: Li, Kaining, et autres
Publié: (2025)
Masked Angle-Aware Autoencoder for Remote Sensing Images
par: Li, Zhihao, et autres
Publié: (2024)
par: Li, Zhihao, et autres
Publié: (2024)
Recurrent Video Masked Autoencoders
par: Zoran, Daniel, et autres
Publié: (2025)
par: Zoran, Daniel, et autres
Publié: (2025)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
par: Wang, Xijun, et autres
Publié: (2023)
par: Wang, Xijun, et autres
Publié: (2023)
Documents similaires
-
G$^2$V$^2$former: Graph Guided Video Vision Transformer for Face Anti-Spoofing
par: Yang, Jingyi, et autres
Publié: (2024) -
Generalized Face Anti-spoofing via Finer Domain Partition and Disentangling Liveness-irrelevant Factors
par: Yang, Jingyi, et autres
Publié: (2024) -
Language Model Guided Interpretable Video Action Reasoning
par: Wang, Ning, et autres
Publié: (2024) -
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
par: Wang, Haoyu, et autres
Publié: (2026) -
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
par: Jin, Qiaoqiao, et autres
Publié: (2024)