Narrative Action Evaluation with Prompt-Guided Multimodal Interaction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shiyi, Bai, Sule, Chen, Guangyi, Chen, Lei, Lu, Jiwen, Wang, Junle, Tang, Yansong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
von: Zhang, Shiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Shiyi, et al.
Veröffentlicht: (2024)
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
von: Bai, Sule, et al.
Veröffentlicht: (2024)
von: Bai, Sule, et al.
Veröffentlicht: (2024)
ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
von: Lu, Guanxing, et al.
Veröffentlicht: (2024)
von: Lu, Guanxing, et al.
Veröffentlicht: (2024)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
von: Cao, Jianjian, et al.
Veröffentlicht: (2024)
von: Cao, Jianjian, et al.
Veröffentlicht: (2024)
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)
von: Bai, Sule, et al.
Veröffentlicht: (2025)
Language-free Compositional Action Generation via Decoupling Refinement
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
von: Liu, Xiao, et al.
Veröffentlicht: (2023)
Open-Vocabulary Segmentation with Semantic-Assisted Calibration
von: Liu, Yong, et al.
Veröffentlicht: (2023)
von: Liu, Yong, et al.
Veröffentlicht: (2023)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
von: Zhang, Chubin, et al.
Veröffentlicht: (2026)
GaussianToken: An Effective Image Tokenizer with 2D Gaussian Splatting
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
von: Dong, Jiajun, et al.
Veröffentlicht: (2025)
GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian Generation
von: Zhang, Chubin, et al.
Veröffentlicht: (2024)
von: Zhang, Chubin, et al.
Veröffentlicht: (2024)
Learning Dual-Level Deformable Implicit Representation for Real-World Scale Arbitrary Super-Resolution
von: Li, Zhiheng, et al.
Veröffentlicht: (2024)
von: Li, Zhiheng, et al.
Veröffentlicht: (2024)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
von: Zhang, Shiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Shiyi, et al.
Veröffentlicht: (2025)
KV-Edit: Training-Free Image Editing for Precise Background Preservation
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
von: Liu, Yong, et al.
Veröffentlicht: (2025)
von: Liu, Yong, et al.
Veröffentlicht: (2025)
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
von: Zhang, Haoji, et al.
Veröffentlicht: (2025)
Q-VLM: Post-training Quantization for Large Vision-Language Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
Towards Accurate Post-training Quantization for Diffusion Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2023)
von: Wang, Changyuan, et al.
Veröffentlicht: (2023)
Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
von: Wang, Sujia, et al.
Veröffentlicht: (2025)
von: Wang, Sujia, et al.
Veröffentlicht: (2025)
InstaRevive: One-Step Image Enhancement via Dynamic Score Matching
von: Zhu, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhu, Yixuan, et al.
Veröffentlicht: (2025)
Once for Both: Single Stage of Importance and Sparsity Search for Vision Transformer Compression
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
von: Ye, Hancheng, et al.
Veröffentlicht: (2024)
Memorize-and-Generate: Towards Long-Term Consistency in Real-Time Video Generation
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
von: Zhu, Tianrui, et al.
Veröffentlicht: (2025)
FlowIE: Efficient Image Enhancement via Rectified Flow
von: Zhu, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yixuan, et al.
Veröffentlicht: (2024)
DPMesh: Exploiting Diffusion Prior for Occluded Human Mesh Recovery
von: Zhu, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yixuan, et al.
Veröffentlicht: (2024)
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
von: Shen, Yuhao, et al.
Veröffentlicht: (2025)
ScoreHOI: Physically Plausible Reconstruction of Human-Object Interaction via Score-Guided Diffusion
von: Li, Ao, et al.
Veröffentlicht: (2025)
von: Li, Ao, et al.
Veröffentlicht: (2025)
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
Segment and Caption Anything
von: Huang, Xiaoke, et al.
Veröffentlicht: (2023)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2023)
Efficient Inference of Vision Instruction-Following Models with Elastic Cache
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
von: Liu, Zuyan, et al.
Veröffentlicht: (2024)
Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline
von: Zhao, Linqing, et al.
Veröffentlicht: (2025)
von: Zhao, Linqing, et al.
Veröffentlicht: (2025)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
von: Ma, Shilin, et al.
Veröffentlicht: (2026)
OccNeRF: Advancing 3D Occupancy Prediction in LiDAR-Free Environments
von: Zhang, Chubin, et al.
Veröffentlicht: (2023)
von: Zhang, Chubin, et al.
Veröffentlicht: (2023)
Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action Recognition
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
von: Zhang, Runzhong, et al.
Veröffentlicht: (2026)
von: Zhang, Runzhong, et al.
Veröffentlicht: (2026)
Progressive Uncertainty-Guided Evidential U-KAN for Trustworthy Medical Image Segmentation
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024)
von: Li, Honglin, et al.
Veröffentlicht: (2024)
Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
Prompt-guided Disentangled Representation for Action Recognition
von: Wu, Tianci, et al.
Veröffentlicht: (2025)
von: Wu, Tianci, et al.
Veröffentlicht: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
EgoPrompt: Prompt Learning for Egocentric Action Recognition
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
von: Zhu, Yixuan, et al.
Veröffentlicht: (2026)
von: Zhu, Yixuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LOGO: A Long-Form Video Dataset for Group Action Quality Assessment
von: Zhang, Shiyi, et al.
Veröffentlicht: (2024) -
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
von: Bai, Sule, et al.
Veröffentlicht: (2024) -
ManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic Manipulation
von: Lu, Guanxing, et al.
Veröffentlicht: (2024) -
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
von: Cao, Jianjian, et al.
Veröffentlicht: (2024) -
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
von: Bai, Sule, et al.
Veröffentlicht: (2025)