Is 3D Convolution with 5D Tensors Really Necessary for Video Analysis?
Fuente:
arXiv
Saved in:
| Main Authors: | Hajimolahoseini, Habib, Ahmed, Walid, Wen, Austin, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
by: Ataiefard, Foozhan, et al.
Published: (2024)
by: Ataiefard, Foozhan, et al.
Published: (2024)
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023)
by: Javadi, Farnoosh, et al.
Published: (2023)
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
by: Hajimolahoseini, Habib, et al.
Published: (2023)
by: Hajimolahoseini, Habib, et al.
Published: (2023)
SUNY: A Visual Interpretation Framework for Convolutional Neural Networks from a Necessary and Sufficient Perspective
by: Xuan, Xiwei, et al.
Published: (2023)
by: Xuan, Xiwei, et al.
Published: (2023)
Comparative Analysis of 3D Convolutional and 2.5D Slice-Conditioned U-Net Architectures for MRI Super-Resolution via Elucidated Diffusion Models
by: Chiche, Hendrik, et al.
Published: (2026)
by: Chiche, Hendrik, et al.
Published: (2026)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
by: Slim, Habib, et al.
Published: (2023)
by: Slim, Habib, et al.
Published: (2023)
Inductive Convolution Nuclear Norm Minimization for Tensor Completion with Arbitrary Sampling
by: Li, Wei, et al.
Published: (2026)
by: Li, Wei, et al.
Published: (2026)
VideoCAD: A Dataset and Model for Learning Long-Horizon 3D CAD UI Interactions from Video
by: Man, Brandon, et al.
Published: (2025)
by: Man, Brandon, et al.
Published: (2025)
Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
by: Lai, Zeqiang, et al.
Published: (2025)
by: Lai, Zeqiang, et al.
Published: (2025)
NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
Online Hand Gesture Recognition Using 3D Convolutional Neural Networks
by: Qin, Yinghao, et al.
Published: (2026)
by: Qin, Yinghao, et al.
Published: (2026)
WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
by: Kong, Hanyang, et al.
Published: (2025)
by: Kong, Hanyang, et al.
Published: (2025)
CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training
by: Karim, Rezaul, et al.
Published: (2026)
by: Karim, Rezaul, et al.
Published: (2026)
ViGiL3D: A Linguistically Diverse Dataset for 3D Visual Grounding
by: Wang, Austin T., et al.
Published: (2025)
by: Wang, Austin T., et al.
Published: (2025)
D3: Training-Free AI-Generated Video Detection Using Second-Order Features
by: Zheng, Chende, et al.
Published: (2025)
by: Zheng, Chende, et al.
Published: (2025)
Mixed Precision PointPillars for Efficient 3D Object Detection with TensorRT
by: Fuengfusin, Ninnart, et al.
Published: (2026)
by: Fuengfusin, Ninnart, et al.
Published: (2026)
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
by: Linghu, Xiongkun, et al.
Published: (2026)
by: Linghu, Xiongkun, et al.
Published: (2026)
How Far are AI-generated Videos from Simulating the 3D Visual World: A Learned 3D Evaluation Approach
by: Chang, Chirui, et al.
Published: (2024)
by: Chang, Chirui, et al.
Published: (2024)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
MTReD: 3D Reconstruction Dataset for Fly-over Videos of Maritime Domain
by: Yong, Rui Yi, et al.
Published: (2025)
by: Yong, Rui Yi, et al.
Published: (2025)
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
Physics3D: Learning Physical Properties of 3D Gaussians via Video Diffusion
by: Liu, Fangfu, et al.
Published: (2024)
by: Liu, Fangfu, et al.
Published: (2024)
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos
by: Liu, Yanan, et al.
Published: (2026)
by: Liu, Yanan, et al.
Published: (2026)
A Framework Combining 3D CNN and Transformer for Video-Based Behavior Recognition
by: Zhang, Xiuliang, et al.
Published: (2025)
by: Zhang, Xiuliang, et al.
Published: (2025)
JOG3R: Towards 3D-Consistent Video Generators
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
by: Huang, Chun-Hao Paul, et al.
Published: (2025)
Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video
by: Alghamdi, Eyad, et al.
Published: (2026)
by: Alghamdi, Eyad, et al.
Published: (2026)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
by: Gu, Zekai, et al.
Published: (2025)
by: Gu, Zekai, et al.
Published: (2025)
Semantic Segmentation of Video Sequences with Convolutional LSTMs
by: Pfeuffer, Andreas, et al.
Published: (2019)
by: Pfeuffer, Andreas, et al.
Published: (2019)
Improving Resnet-9 Generalization Trained on Small Datasets
by: Awad, Omar Mohamed, et al.
Published: (2023)
by: Awad, Omar Mohamed, et al.
Published: (2023)
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
by: Wu, Haoyu, et al.
Published: (2025)
by: Wu, Haoyu, et al.
Published: (2025)
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting
by: Song, Yixiao, et al.
Published: (2026)
by: Song, Yixiao, et al.
Published: (2026)
Latent Intuitive Physics: Learning to Transfer Hidden Physics from A 3D Video
by: Zhu, Xiangming, et al.
Published: (2024)
by: Zhu, Xiangming, et al.
Published: (2024)
HY3D-Bench: Generation of 3D Assets
by: Hunyuan3D, Team, et al.
Published: (2026)
by: Hunyuan3D, Team, et al.
Published: (2026)
Pomo3D: 3D-Aware Portrait Accessorizing and More
by: Liu, Tzu-Chieh, et al.
Published: (2024)
by: Liu, Tzu-Chieh, et al.
Published: (2024)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
by: Zhang, Jihai, et al.
Published: (2025)
by: Zhang, Jihai, et al.
Published: (2025)
Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection
by: Ning, Zhiwei, et al.
Published: (2026)
by: Ning, Zhiwei, et al.
Published: (2026)
Towards Efficient 3D Object Detection in Bird's-Eye-View Space for Autonomous Driving: A Convolutional-Only Approach
by: Li, Yuxin, et al.
Published: (2023)
by: Li, Yuxin, et al.
Published: (2023)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
Similar Items
-
SkipViT: Speeding Up Vision Transformers with a Token-Level Skip Connection
by: Ataiefard, Foozhan, et al.
Published: (2024) -
GQKVA: Efficient Pre-training of Transformers by Grouping Queries, Keys, and Values
by: Javadi, Farnoosh, et al.
Published: (2023) -
Training Acceleration of Low-Rank Decomposed Networks using Sequential Freezing and Rank Quantization
by: Hajimolahoseini, Habib, et al.
Published: (2023) -
SUNY: A Visual Interpretation Framework for Convolutional Neural Networks from a Necessary and Sufficient Perspective
by: Xuan, Xiwei, et al.
Published: (2023) -
Comparative Analysis of 3D Convolutional and 2.5D Slice-Conditioned U-Net Architectures for MRI Super-Resolution via Elucidated Diffusion Models
by: Chiche, Hendrik, et al.
Published: (2026)