Gespeichert in:
| Hauptverfasser: | Fadaei, Amir Hosein, Dehaqani, Mohammad-Reza A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.07277 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond still images: Temporal features and input variance resilience
von: Fadaei, Amir Hosein, et al.
Veröffentlicht: (2023)
von: Fadaei, Amir Hosein, et al.
Veröffentlicht: (2023)
SpikeReg: Energy-Efficient 3D Deformable Medical Image Registration with Spiking Neural Networks
von: Barzili, Ali Mikaeili, et al.
Veröffentlicht: (2026)
von: Barzili, Ali Mikaeili, et al.
Veröffentlicht: (2026)
Wise-SrNet: A Novel Architecture for Enhancing Image Classification by Learning Spatial Resolution of Feature Maps
von: Rahimzadeh, Mohammad, et al.
Veröffentlicht: (2021)
von: Rahimzadeh, Mohammad, et al.
Veröffentlicht: (2021)
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis
von: Li, Gary Y., et al.
Veröffentlicht: (2025)
von: Li, Gary Y., et al.
Veröffentlicht: (2025)
Improving 3D Few-Shot Segmentation with Inference-Time Pseudo-Labeling
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
von: Mozafari, Mohammad, et al.
Veröffentlicht: (2024)
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
von: Luo, Bingjun, et al.
Veröffentlicht: (2026)
Physics Context Builders: A Modular Framework for Physical Reasoning in Vision-Language Models
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2024)
von: Balazadeh, Vahid, et al.
Veröffentlicht: (2024)
Using Deep Convolutional Neural Networks to Detect Rendered Glitches in Video Games
von: Ling, Carlos Garcia, et al.
Veröffentlicht: (2024)
von: Ling, Carlos Garcia, et al.
Veröffentlicht: (2024)
Towards Neuro-Symbolic Video Understanding
von: Choi, Minkyu, et al.
Veröffentlicht: (2024)
von: Choi, Minkyu, et al.
Veröffentlicht: (2024)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
High Resolution Flood Extent Detection Using Deep Learning with Random Forest Derived Training Labels
von: Nuriddinov, Azizbek, et al.
Veröffentlicht: (2026)
von: Nuriddinov, Azizbek, et al.
Veröffentlicht: (2026)
Memory-Efficient Continual Learning Object Segmentation for Long Video
von: Nazemi, Amir, et al.
Veröffentlicht: (2023)
von: Nazemi, Amir, et al.
Veröffentlicht: (2023)
A Survey: Spatiotemporal Consistency in Video Generation
von: Yin, Zhiyu, et al.
Veröffentlicht: (2025)
von: Yin, Zhiyu, et al.
Veröffentlicht: (2025)
VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
Extracting Overlapping Microservices from Monolithic Code via Deep Semantic Embeddings and Graph Neural Network-Based Soft Clustering
von: Ziabakhsh, Morteza, et al.
Veröffentlicht: (2025)
von: Ziabakhsh, Morteza, et al.
Veröffentlicht: (2025)
Enhancing Few-Shot Image Classification through Learnable Multi-Scale Embedding and Attention Mechanisms
von: Askari, Fatemeh, et al.
Veröffentlicht: (2024)
von: Askari, Fatemeh, et al.
Veröffentlicht: (2024)
MSDNet: Multi-Scale Decoder for Few-Shot Semantic Segmentation via Transformer-Guided Prototyping
von: Fateh, Amirreza, et al.
Veröffentlicht: (2024)
von: Fateh, Amirreza, et al.
Veröffentlicht: (2024)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
Brand Visibility in Packaging: A Deep Learning Approach for Logo Detection, Saliency-Map Prediction, and Logo Placement Analysis
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
von: Liang, Yujia, et al.
Veröffentlicht: (2025)
Understanding Multimodal Deep Neural Networks: A Concept Selection View
von: Shang, Chenming, et al.
Veröffentlicht: (2024)
von: Shang, Chenming, et al.
Veröffentlicht: (2024)
Understanding Distributed Representations of Concepts in Deep Neural Networks without Supervision
von: Chang, Wonjoon, et al.
Veröffentlicht: (2023)
von: Chang, Wonjoon, et al.
Veröffentlicht: (2023)
Language-guided Recursive Spatiotemporal Graph Modeling for Video Summarization
von: Park, Jungin, et al.
Veröffentlicht: (2025)
von: Park, Jungin, et al.
Veröffentlicht: (2025)
Enhancing Long Video Understanding via Hierarchical Event-Based Memory
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
von: Cheng, Dingxin, et al.
Veröffentlicht: (2024)
Graph-Attention Network with Adversarial Domain Alignment for Robust Cross-Domain Facial Expression Recognition
von: Ghaedi, Razieh, et al.
Veröffentlicht: (2025)
von: Ghaedi, Razieh, et al.
Veröffentlicht: (2025)
StabStitch++: Unsupervised Online Video Stitching with Spatiotemporal Bidirectional Warps
von: Nie, Lang, et al.
Veröffentlicht: (2025)
von: Nie, Lang, et al.
Veröffentlicht: (2025)
Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
von: Du, Dazhao, et al.
Veröffentlicht: (2026)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
Video Panels for Long Video Understanding
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Taesiri, Mohammad Reza, et al.
Veröffentlicht: (2025)
Self-Supervised Learning for Endoscopic Video Analysis
von: Hirsch, Roy, et al.
Veröffentlicht: (2023)
von: Hirsch, Roy, et al.
Veröffentlicht: (2023)
Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
von: Hasani, Hosein, et al.
Veröffentlicht: (2025)
WhisperNetV2: SlowFast Siamese Network For Lip-Based Biometrics
von: Zakeri, Abdollah, et al.
Veröffentlicht: (2024)
von: Zakeri, Abdollah, et al.
Veröffentlicht: (2024)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
Deep Neural Networks Fused with Textures for Image Classification
von: Bera, Asish, et al.
Veröffentlicht: (2023)
von: Bera, Asish, et al.
Veröffentlicht: (2023)
CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating
von: Wang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Jiyuan, et al.
Veröffentlicht: (2026)
Personalized Video Summarization by Multimodal Video Understanding
von: Chen, Brian, et al.
Veröffentlicht: (2024)
von: Chen, Brian, et al.
Veröffentlicht: (2024)
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Deep Learning-Driven Multimodal Detection and Movement Analysis of Objects in Culinary
von: Ishat, Tahoshin Alam, et al.
Veröffentlicht: (2025)
von: Ishat, Tahoshin Alam, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Beyond still images: Temporal features and input variance resilience
von: Fadaei, Amir Hosein, et al.
Veröffentlicht: (2023) -
SpikeReg: Energy-Efficient 3D Deformable Medical Image Registration with Spiking Neural Networks
von: Barzili, Ali Mikaeili, et al.
Veröffentlicht: (2026) -
Wise-SrNet: A Novel Architecture for Enhancing Image Classification by Learning Spatial Resolution of Feature Maps
von: Rahimzadeh, Mohammad, et al.
Veröffentlicht: (2021) -
Understanding Counting Mechanisms in Large Language and Vision-Language Models
von: Hasani, Hosein, et al.
Veröffentlicht: (2025) -
Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis
von: Li, Gary Y., et al.
Veröffentlicht: (2025)