Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Mengshi, Ji, Hongwei, Yun, Wulian, Zhang, Xianlin, Ma, Huadong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment
by: Yun, Wulian, et al.
Published: (2024)
by: Yun, Wulian, et al.
Published: (2024)
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation
by: Yun, Wulian, et al.
Published: (2025)
by: Yun, Wulian, et al.
Published: (2025)
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Active Multimodal Distillation for Few-shot Action Recognition
by: Feng, Weijia, et al.
Published: (2025)
by: Feng, Weijia, et al.
Published: (2025)
Towards Robust Unsupervised Attention Prediction in Autonomous Driving
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Question-Aware Evidence Ledgers for Video Relational Reasoning
by: Ou, Yilin, et al.
Published: (2026)
by: Ou, Yilin, et al.
Published: (2026)
Multi-Stage Contrastive Regression for Action Quality Assessment
by: An, Qi, et al.
Published: (2024)
by: An, Qi, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models
by: Deng, Wei, et al.
Published: (2026)
by: Deng, Wei, et al.
Published: (2026)
Robust Disentangled Counterfactual Learning for Physical Audiovisual Commonsense Reasoning
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
by: Ni, Xinzhe, et al.
Published: (2022)
by: Ni, Xinzhe, et al.
Published: (2022)
Reliable Few-shot Learning under Dual Noises
by: Zhang, Ji, et al.
Published: (2025)
by: Zhang, Ji, et al.
Published: (2025)
Few-shot Writer Adaptation via Multimodal In-Context Learning
by: Simon, Tom, et al.
Published: (2026)
by: Simon, Tom, et al.
Published: (2026)
Enhancing Temporal Action Localization: Advanced S6 Modeling with Recurrent Mechanism
by: Lee, Sangyoun, et al.
Published: (2024)
by: Lee, Sangyoun, et al.
Published: (2024)
T2SG: Traffic Topology Scene Graph for Topology Reasoning in Autonomous Driving
by: Lv, Changsheng, et al.
Published: (2024)
by: Lv, Changsheng, et al.
Published: (2024)
Learning Group Interactions and Semantic Intentions for Multi-Object Trajectory Prediction
by: Qi, Mengshi, et al.
Published: (2024)
by: Qi, Mengshi, et al.
Published: (2024)
Decomposed Vector-Quantized Variational Autoencoder for Human Grasp Generation
by: Zhao, Zhe, et al.
Published: (2024)
by: Zhao, Zhe, et al.
Published: (2024)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023)
by: Huang, Yifeng, et al.
Published: (2023)
Robo-SGG: Exploiting Layout-Oriented Normalization and Restitution Can Improve Robust Scene Graph Generation
by: Lv, Changsheng, et al.
Published: (2025)
by: Lv, Changsheng, et al.
Published: (2025)
SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action Recognition
by: Huang, Wenbo, et al.
Published: (2024)
by: Huang, Wenbo, et al.
Published: (2024)
VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought
by: Jiang, Chaoya, et al.
Published: (2025)
by: Jiang, Chaoya, et al.
Published: (2025)
Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
by: Gong, Haozhen, et al.
Published: (2025)
by: Gong, Haozhen, et al.
Published: (2025)
Siamese Transformer Networks for Few-shot Image Classification
by: Jiang, Weihao, et al.
Published: (2024)
by: Jiang, Weihao, et al.
Published: (2024)
STAR: Semantic-Temporal Adaptive Representation Learning for Few-Shot Action Recognition
by: Liu, Hongli, et al.
Published: (2026)
by: Liu, Hongli, et al.
Published: (2026)
Few-shot Semantic Encoding and Decoding for Video Surveillance
by: Cheng, Baoping, et al.
Published: (2025)
by: Cheng, Baoping, et al.
Published: (2025)
Emma-X: An Embodied Multimodal Action Model with Grounded Chain of Thought and Look-ahead Spatial Reasoning
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Multimodal Chain-of-Thought Reasoning in Language Models
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
Small Object Few-shot Segmentation for Vision-based Industrial Inspection
by: Zhang, Zilong, et al.
Published: (2024)
by: Zhang, Zilong, et al.
Published: (2024)
MVREC: A General Few-shot Defect Classification Model Using Multi-View Region-Context
by: Lyu, Shuai, et al.
Published: (2024)
by: Lyu, Shuai, et al.
Published: (2024)
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
by: Cheng, Jen-Hao, et al.
Published: (2025)
by: Cheng, Jen-Hao, et al.
Published: (2025)
Few-shot Implicit Function Generation via Equivariance
by: Huang, Suizhi, et al.
Published: (2025)
by: Huang, Suizhi, et al.
Published: (2025)
Task Consistent Prototype Learning for Incremental Few-shot Semantic Segmentation
by: Xu, Wenbo, et al.
Published: (2024)
by: Xu, Wenbo, et al.
Published: (2024)
Few-shot Open Relation Extraction with Gaussian Prototype and Adaptive Margin
by: Guo, Tianlin, et al.
Published: (2024)
by: Guo, Tianlin, et al.
Published: (2024)
Mind the Discriminability Trap in Source-Free Cross-domain Few-shot Learning
by: Zhang, Zhenyu, et al.
Published: (2026)
by: Zhang, Zhenyu, et al.
Published: (2026)
VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer
by: Lin, Rui, et al.
Published: (2026)
by: Lin, Rui, et al.
Published: (2026)
REMEMBER: Retrieval-based Explainable Multimodal Evidence-guided Modeling for Brain Evaluation and Reasoning in Zero- and Few-shot Neurodegenerative Diagnosis
by: Can, Duy-Cat, et al.
Published: (2025)
by: Can, Duy-Cat, et al.
Published: (2025)
Exploring Scalability of Self-Training for Open-Vocabulary Temporal Action Localization
by: Hyun, Jeongseok, et al.
Published: (2024)
by: Hyun, Jeongseok, et al.
Published: (2024)
Similar Items
-
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
by: Qi, Mengshi, et al.
Published: (2025) -
Semi-Supervised Teacher-Reference-Student Architecture for Action Quality Assessment
by: Yun, Wulian, et al.
Published: (2024) -
Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
by: Qi, Mengshi, et al.
Published: (2025) -
A New Teacher-Reviewer-Student Framework for Semi-supervised 2D Human Pose Estimation
by: Yun, Wulian, et al.
Published: (2025) -
Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression
by: Qi, Mengshi, et al.
Published: (2025)