LLMs are Good Action Recognizers
Fuente:
arXiv
Saved in:
| Main Authors: | Qu, Haoxuan, Cai, Yujun, Liu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Off-the-shelf ChatGPT is a Good Few-shot Human Motion Predictor
by: Qu, Haoxuan, et al.
Published: (2024)
by: Qu, Haoxuan, et al.
Published: (2024)
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose Estimation
by: Xu, Li, et al.
Published: (2023)
by: Xu, Li, et al.
Published: (2023)
DisC-GS: Discontinuity-aware Gaussian Splatting
by: Qu, Haoxuan, et al.
Published: (2024)
by: Qu, Haoxuan, et al.
Published: (2024)
A Mixed-Primitive-based Gaussian Splatting Method for Surface Reconstruction
by: Qu, Haoxuan, et al.
Published: (2025)
by: Qu, Haoxuan, et al.
Published: (2025)
Learning to Generate Cross-Task Unexploitable Examples
by: Qu, Haoxuan, et al.
Published: (2025)
by: Qu, Haoxuan, et al.
Published: (2025)
GPT-Connect: Interaction between Text-Driven Human Motion Generator and 3D Scenes in a Training-free Manner
by: Qu, Haoxuan, et al.
Published: (2024)
by: Qu, Haoxuan, et al.
Published: (2024)
Enhancing Human-Centered Dynamic Scene Understanding via Multiple LLMs Collaborated Reasoning
by: Zhang, Hang, et al.
Published: (2024)
by: Zhang, Hang, et al.
Published: (2024)
An Image-like Diffusion Method for Human-Object Interaction Detection
by: Hui, Xiaofei, et al.
Published: (2025)
by: Hui, Xiaofei, et al.
Published: (2025)
REACT: Recognize Every Action Everywhere All At Once
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
LLMs are Good Sign Language Translators
by: Gong, Jia, et al.
Published: (2024)
by: Gong, Jia, et al.
Published: (2024)
Recognizing Actions from Robotic View for Natural Human-Robot Interaction
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
CMMLoc: Advancing Text-to-PointCloud Localization with Cauchy-Mixture-Model Based Framework
by: Xu, Yanlong, et al.
Published: (2025)
by: Xu, Yanlong, et al.
Published: (2025)
Recent Advances of Continual Learning in Computer Vision: An Overview
by: Qu, Haoxuan, et al.
Published: (2021)
by: Qu, Haoxuan, et al.
Published: (2021)
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
by: Yuan, Zhenlong, et al.
Published: (2025)
by: Yuan, Zhenlong, et al.
Published: (2025)
Translating Signals to Languages for sEMG-Based Activity Recognition
by: Wang, Ming, et al.
Published: (2026)
by: Wang, Ming, et al.
Published: (2026)
When Visual Privacy Protection Meets Multimodal Large Language Models
by: Hui, Xiaofei, et al.
Published: (2026)
by: Hui, Xiaofei, et al.
Published: (2026)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
by: Wu, Fangyu, et al.
Published: (2025)
by: Wu, Fangyu, et al.
Published: (2025)
Order Matters: On Parameter-Efficient Image-to-Video Probing for Recognizing Nearly Symmetric Actions
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2025)
Making Large Vision Language Models to be Good Few-shot Learners
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
ToolFG: Towards Well-Grounded Fine-Grained Image Classification
by: Xue, Yu, et al.
Published: (2026)
by: Xue, Yu, et al.
Published: (2026)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
CRAFT-LoRA: Content-Style Personalization via Rank-Constrained Adaptation and Training-Free Fusion
by: Li, Yu, et al.
Published: (2026)
by: Li, Yu, et al.
Published: (2026)
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
by: Hao, Shuyang, et al.
Published: (2024)
by: Hao, Shuyang, et al.
Published: (2024)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
by: Zhong, Humen, et al.
Published: (2024)
by: Zhong, Humen, et al.
Published: (2024)
Recognizing Co-Speech Gestures in-the-Wild
by: Hegde, Sindhu B, et al.
Published: (2026)
by: Hegde, Sindhu B, et al.
Published: (2026)
AnimalClue: Recognizing Animals by their Traces
by: Shinoda, Risa, et al.
Published: (2025)
by: Shinoda, Risa, et al.
Published: (2025)
How Do You Perceive My Face? Recognizing Facial Expressions in Multi-Modal Context by Modeling Mental Representations
by: Blume, Florian, et al.
Published: (2024)
by: Blume, Florian, et al.
Published: (2024)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
by: Sun, Bowen, et al.
Published: (2025)
by: Sun, Bowen, et al.
Published: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
MicroscopyMatching: Towards a Ready-to-use Framework for Microscopy Image Analysis in Diverse Conditions
by: Hui, Xiaofei, et al.
Published: (2026)
by: Hui, Xiaofei, et al.
Published: (2026)
Project RISE: Recognizing Industrial Smoke Emissions
by: Hsu, Yen-Chia, et al.
Published: (2020)
by: Hsu, Yen-Chia, et al.
Published: (2020)
SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
by: Xu, Qianxun, et al.
Published: (2026)
by: Xu, Qianxun, et al.
Published: (2026)
Beyond the Label Itself: Latent Labels Enhance Semi-supervised Point Cloud Panoptic Segmentation
by: Chen, Yujun, et al.
Published: (2023)
by: Chen, Yujun, et al.
Published: (2023)
Recognize Any Surgical Object: Unleashing the Power of Weakly-Supervised Data
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
TSTMotion: Training-free Scene-aware Text-to-motion Generation
by: Guo, Ziyan, et al.
Published: (2025)
by: Guo, Ziyan, et al.
Published: (2025)
Recognize Any Regions
by: Yang, Haosen, et al.
Published: (2023)
by: Yang, Haosen, et al.
Published: (2023)
BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-Identification
by: Xu, Haoxuan, et al.
Published: (2026)
by: Xu, Haoxuan, et al.
Published: (2026)
Multi-modal In-Context Learning Makes an Ego-evolving Scene Text Recognizer
by: Zhao, Zhen, et al.
Published: (2023)
by: Zhao, Zhen, et al.
Published: (2023)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
by: Fu, Honghao, et al.
Published: (2026)
by: Fu, Honghao, et al.
Published: (2026)
Similar Items
-
Off-the-shelf ChatGPT is a Good Few-shot Human Motion Predictor
by: Qu, Haoxuan, et al.
Published: (2024) -
HyLiFormer: Hyperbolic Linear Attention for Skeleton-based Human Action Recognition
by: Li, Yue, et al.
Published: (2025) -
6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose Estimation
by: Xu, Li, et al.
Published: (2023) -
DisC-GS: Discontinuity-aware Gaussian Splatting
by: Qu, Haoxuan, et al.
Published: (2024) -
A Mixed-Primitive-based Gaussian Splatting Method for Surface Reconstruction
by: Qu, Haoxuan, et al.
Published: (2025)