An Effective End-to-End Solution for Multimodal Action Recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Songping, Hu, Xiantao, Lyu, Yueming, Shan, Caifeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Adversarial Transferability between Kolmogorov-arnold Networks
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
by: Wang, Songping, et al.
Published: (2026)
by: Wang, Songping, et al.
Published: (2026)
Anti-Aesthetics: Protecting Facial Privacy against Customized Text-to-Image Synthesis
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
RunawayEvil: Jailbreaking the Image-to-Video Generative Models
by: Wang, Songping, et al.
Published: (2025)
by: Wang, Songping, et al.
Published: (2025)
Towards End-to-End Explainable Facial Action Unit Recognition via Vision-Language Joint Learning
by: Ge, Xuri, et al.
Published: (2024)
by: Ge, Xuri, et al.
Published: (2024)
Multimodal Action Diffusion for Robust End-to-End Autonomous Driving
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
by: Rodríguez-Vidal, Jorge Daniel, et al.
Published: (2026)
End-to-End Chess Recognition
by: Masouris, Athanasios, et al.
Published: (2023)
by: Masouris, Athanasios, et al.
Published: (2023)
An End-to-End Two-Stream Network Based on RGB Flow and Representation Flow for Human Action Recognition
by: Lai, Song-Jiang, et al.
Published: (2024)
by: Lai, Song-Jiang, et al.
Published: (2024)
CFVNet: An End-to-End Cancelable Finger Vein Network for Recognition
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
End-to-End Action Segmentation Transformer
by: Wang, Tieqiao, et al.
Published: (2025)
by: Wang, Tieqiao, et al.
Published: (2025)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
by: Zhang, Meng, et al.
Published: (2026)
by: Zhang, Meng, et al.
Published: (2026)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Towards Weakly Supervised End-to-end Learning for Long-video Action Recognition
by: Zhou, Jiaming, et al.
Published: (2023)
by: Zhou, Jiaming, et al.
Published: (2023)
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
by: Zhang, Jinrong, et al.
Published: (2023)
by: Zhang, Jinrong, et al.
Published: (2023)
MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
by: Strohmeyer, Tim, et al.
Published: (2026)
by: Strohmeyer, Tim, et al.
Published: (2026)
DLAFormer: An End-to-End Transformer For Document Layout Analysis
by: Wang, Jiawei, et al.
Published: (2024)
by: Wang, Jiawei, et al.
Published: (2024)
Efficient and Explainable End-to-End Autonomous Driving via Masked Vision-Language-Action Diffusion
by: Zhang, Jiaru, et al.
Published: (2026)
by: Zhang, Jiaru, et al.
Published: (2026)
ScrewSplat: An End-to-End Method for Articulated Object Recognition
by: Kim, Seungyeon, et al.
Published: (2025)
by: Kim, Seungyeon, et al.
Published: (2025)
REMM:Rotation-Equivariant Framework for End-to-End Multimodal Image Matching
by: Nie, Han, et al.
Published: (2024)
by: Nie, Han, et al.
Published: (2024)
End-to-End Full-Page Optical Music Recognition for Pianoform Sheet Music
by: Ríos-Vila, Antonio, et al.
Published: (2024)
by: Ríos-Vila, Antonio, et al.
Published: (2024)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
by: Liu, Shuming, et al.
Published: (2023)
by: Liu, Shuming, et al.
Published: (2023)
GoalFlow: Goal-Driven Flow Matching for Multimodal Trajectories Generation in End-to-End Autonomous Driving
by: Xing, Zebin, et al.
Published: (2025)
by: Xing, Zebin, et al.
Published: (2025)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
by: Wang, Linhan, et al.
Published: (2026)
by: Wang, Linhan, et al.
Published: (2026)
Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving
by: Wang, Linbo, et al.
Published: (2026)
by: Wang, Linbo, et al.
Published: (2026)
CryptoFace: End-to-End Encrypted Face Recognition
by: Ao, Wei, et al.
Published: (2025)
by: Ao, Wei, et al.
Published: (2025)
ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
by: Fu, Haoyu, et al.
Published: (2025)
by: Fu, Haoyu, et al.
Published: (2025)
LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization
by: Gupta, Akshita, et al.
Published: (2024)
by: Gupta, Akshita, et al.
Published: (2024)
Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs
by: Cheng, Dabing, et al.
Published: (2025)
by: Cheng, Dabing, et al.
Published: (2025)
VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
by: Zhu, Jiaying, et al.
Published: (2025)
by: Zhu, Jiaying, et al.
Published: (2025)
End-to-End Vision Tokenizer Tuning
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
E2E-MFD: Towards End-to-End Synchronous Multimodal Fusion Detection
by: Zhang, Jiaqing, et al.
Published: (2024)
by: Zhang, Jiaqing, et al.
Published: (2024)
Active Learning from Scene Embeddings for End-to-End Autonomous Driving
by: Jiang, Wenhao, et al.
Published: (2025)
by: Jiang, Wenhao, et al.
Published: (2025)
ChartE$^{3}$: A Comprehensive Benchmark for End-to-End Chart Editing
by: Li, Shuo, et al.
Published: (2026)
by: Li, Shuo, et al.
Published: (2026)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
by: You, Ling, et al.
Published: (2025)
by: You, Ling, et al.
Published: (2025)
VLA-R: Vision-Language Action Retrieval toward Open-World End-to-End Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
Skeleton-OOD: An End-to-End Skeleton-Based Model for Robust Out-of-Distribution Human Action Detection
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
LP-LLM: End-to-End Real-World Degraded License Plate Text Recognition via Large Multimodal Models
by: Gong, Haoyan, et al.
Published: (2026)
by: Gong, Haoyan, et al.
Published: (2026)
Similar Items
-
Exploring Adversarial Transferability between Kolmogorov-arnold Networks
by: Wang, Songping, et al.
Published: (2025) -
Fast Adversarial Training with Weak-to-Strong Spatial-Temporal Consistency in the Frequency Domain on Videos
by: Wang, Songping, et al.
Published: (2025) -
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
by: Wang, Songping, et al.
Published: (2026) -
Anti-Aesthetics: Protecting Facial Privacy against Customized Text-to-Image Synthesis
by: Wang, Songping, et al.
Published: (2025) -
RunawayEvil: Jailbreaking the Image-to-Video Generative Models
by: Wang, Songping, et al.
Published: (2025)