Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jingyi, Yu, Zitong, Ni, Xiuming, He, Jia, Li, Hui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
G$^2$V$^2$former: Graph Guided Video Vision Transformer for Face Anti-Spoofing
by: Yang, Jingyi, et al.
Published: (2024)
by: Yang, Jingyi, et al.
Published: (2024)
Generalized Face Anti-spoofing via Finer Domain Partition and Disentangling Liveness-irrelevant Factors
by: Yang, Jingyi, et al.
Published: (2024)
by: Yang, Jingyi, et al.
Published: (2024)
Language Model Guided Interpretable Video Action Reasoning
by: Wang, Ning, et al.
Published: (2024)
by: Wang, Ning, et al.
Published: (2024)
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
by: Wang, Haoyu, et al.
Published: (2026)
by: Wang, Haoyu, et al.
Published: (2026)
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
by: Jin, Qiaoqiao, et al.
Published: (2024)
by: Jin, Qiaoqiao, et al.
Published: (2024)
Masked Autoencoders are Parameter-Efficient Federated Continual Learners
by: He, Yuchen, et al.
Published: (2024)
by: He, Yuchen, et al.
Published: (2024)
Compositional Kronecker Context Optimization for Vision-Language Models
by: Ding, Kun, et al.
Published: (2024)
by: Ding, Kun, et al.
Published: (2024)
ActionVOS: Actions as Prompts for Video Object Segmentation
by: Ouyang, Liangyang, et al.
Published: (2024)
by: Ouyang, Liangyang, et al.
Published: (2024)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
by: Zhu, Yijie, et al.
Published: (2026)
by: Zhu, Yijie, et al.
Published: (2026)
Video Diffusion Transformers are In-Context Learners
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
From Recognition to Prediction: Leveraging Sequence Reasoning for Action Anticipation
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
by: Ni, Jingcheng, et al.
Published: (2025)
by: Ni, Jingcheng, et al.
Published: (2025)
DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
GazeCLIP: Gaze-Guided CLIP with Adaptive-Enhanced Fine-Grained Language Prompt for Deepfake Attribution and Detection
by: Zhang, Yaning, et al.
Published: (2026)
by: Zhang, Yaning, et al.
Published: (2026)
Contrastive Masked Autoencoders are Stronger Vision Learners
by: Huang, Zhicheng, et al.
Published: (2022)
by: Huang, Zhicheng, et al.
Published: (2022)
Masked Diffusion as Self-supervised Representation Learner
by: Pan, Zixuan, et al.
Published: (2023)
by: Pan, Zixuan, et al.
Published: (2023)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
by: Zhang, Mingfang, et al.
Published: (2024)
by: Zhang, Mingfang, et al.
Published: (2024)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
EPE-P: Evidence-based Parameter-efficient Prompting for Multimodal Learning with Missing Modalities
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
Diffusion Models as Masked Audio-Video Learners
by: Nunez, Elvis, et al.
Published: (2023)
by: Nunez, Elvis, et al.
Published: (2023)
AULLM++: Structural Reasoning with Large Language Models for Micro-Expression Recognition
by: Liu, Zhishu, et al.
Published: (2026)
by: Liu, Zhishu, et al.
Published: (2026)
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
by: Yan, Weicai, et al.
Published: (2025)
by: Yan, Weicai, et al.
Published: (2025)
MARMOT: Masked Autoencoder for Modeling Transient Imaging
by: Shen, Siyuan, et al.
Published: (2025)
by: Shen, Siyuan, et al.
Published: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
by: Fei, Jiajun, et al.
Published: (2024)
by: Fei, Jiajun, et al.
Published: (2024)
Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Adversarial Robustness in RGB-Skeleton Action Recognition: Leveraging Attention Modality Reweighter
by: Liu, Chao, et al.
Published: (2024)
by: Liu, Chao, et al.
Published: (2024)
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
by: Li, Yu-Jhe, et al.
Published: (2024)
by: Li, Yu-Jhe, et al.
Published: (2024)
Large Language Models are Good Prompt Learners for Low-Shot Image Classification
by: Zheng, Zhaoheng, et al.
Published: (2023)
by: Zheng, Zhaoheng, et al.
Published: (2023)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
Deep Kronecker Network
by: Feng, Long, et al.
Published: (2022)
by: Feng, Long, et al.
Published: (2022)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos
by: Li, Kaining, et al.
Published: (2025)
by: Li, Kaining, et al.
Published: (2025)
Masked Angle-Aware Autoencoder for Remote Sensing Images
by: Li, Zhihao, et al.
Published: (2024)
by: Li, Zhihao, et al.
Published: (2024)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
SCP: Soft Conditional Prompt Learning for Aerial Video Action Recognition
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
Similar Items
-
G$^2$V$^2$former: Graph Guided Video Vision Transformer for Face Anti-Spoofing
by: Yang, Jingyi, et al.
Published: (2024) -
Generalized Face Anti-spoofing via Finer Domain Partition and Disentangling Liveness-irrelevant Factors
by: Yang, Jingyi, et al.
Published: (2024) -
Language Model Guided Interpretable Video Action Reasoning
by: Wang, Ning, et al.
Published: (2024) -
HalluCXR: Benchmarking and Mitigating Hallucinations in Medical Vision-Language Models for Chest Radiograph Interpretation
by: Wang, Haoyu, et al.
Published: (2026) -
AU-vMAE: Knowledge-Guide Action Units Detection via Video Masked Autoencoder
by: Jin, Qiaoqiao, et al.
Published: (2024)