A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
Fuente:
arXiv
Salvato in:
| Autori principali: | Montello, Fabio, Güldenring, Ronja, Scardapane, Simone, Nalpantidis, Lazaros |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ClustViT: Clustering-based Token Merging for Semantic Segmentation
di: Montello, Fabio, et al.
Pubblicazione: (2025)
di: Montello, Fabio, et al.
Pubblicazione: (2025)
GFLAN: Generative Functional Layouts
di: Abouagour, Mohamed, et al.
Pubblicazione: (2025)
di: Abouagour, Mohamed, et al.
Pubblicazione: (2025)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
di: Chen, Kewei, et al.
Pubblicazione: (2026)
di: Chen, Kewei, et al.
Pubblicazione: (2026)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
di: Papyan, Narek, et al.
Pubblicazione: (2024)
di: Papyan, Narek, et al.
Pubblicazione: (2024)
Predicting Pedestrian Crossing Behavior in Germany and Japan: Insights into Model Transferability
di: Zhang, Chi, et al.
Pubblicazione: (2024)
di: Zhang, Chi, et al.
Pubblicazione: (2024)
Predicting and Analyzing Pedestrian Crossing Behavior at Unsignalized Crossings
di: Zhang, Chi, et al.
Pubblicazione: (2024)
di: Zhang, Chi, et al.
Pubblicazione: (2024)
GLL: A Differentiable Graph Learning Layer for Neural Networks
di: Brown, Jason, et al.
Pubblicazione: (2024)
di: Brown, Jason, et al.
Pubblicazione: (2024)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
di: Zhang, Junbin, et al.
Pubblicazione: (2022)
CulinaryCut-VLAP: A Vision-Language-Action-Physics Framework for Food Cutting via a Force-Aware Material Point Method
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
di: Koh, Hyunseo, et al.
Pubblicazione: (2026)
FT-NCFM: An Influence-Aware Data Distillation Framework for Efficient VLA Models
di: Chen, Kewei, et al.
Pubblicazione: (2025)
di: Chen, Kewei, et al.
Pubblicazione: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
di: Patel, Hitesh Laxmichand, et al.
Pubblicazione: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
di: Adžemović, Momir
Pubblicazione: (2025)
di: Adžemović, Momir
Pubblicazione: (2025)
MaP-AVR: A Meta-Action Planner for Agents Leveraging Vision Language Models and Retrieval-Augmented Generation
di: Guo, Zhenglong, et al.
Pubblicazione: (2025)
di: Guo, Zhenglong, et al.
Pubblicazione: (2025)
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Systematic Review and Practical Design Guidelines
di: Wimalasiri, Chathura
Pubblicazione: (2026)
di: Wimalasiri, Chathura
Pubblicazione: (2026)
Towards Ubiquitous Mapping and Localization for Dynamic Indoor Environments
di: Djerroud, Halim, et al.
Pubblicazione: (2026)
di: Djerroud, Halim, et al.
Pubblicazione: (2026)
Measuring What Matters: Scenario-Driven Evaluation for Trajectory Predictors in Autonomous Driving
di: Da, Longchao, et al.
Pubblicazione: (2025)
di: Da, Longchao, et al.
Pubblicazione: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
di: Ma, Yueen, et al.
Pubblicazione: (2024)
di: Ma, Yueen, et al.
Pubblicazione: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
di: Zeng, Zhitao, et al.
Pubblicazione: (2026)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
di: Lysyi, Andrii, et al.
Pubblicazione: (2025)
di: Lysyi, Andrii, et al.
Pubblicazione: (2025)
Key-Scan-Based Mobile Robot Navigation: Integrated Mapping, Planning, and Control using Graphs of Scan Regions
di: Latha, Dharshan Bashkaran, et al.
Pubblicazione: (2024)
di: Latha, Dharshan Bashkaran, et al.
Pubblicazione: (2024)
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
di: Li, Jinhao, et al.
Pubblicazione: (2024)
di: Li, Jinhao, et al.
Pubblicazione: (2024)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
RefineFormer3D: Efficient 3D Medical Image Segmentation via Adaptive Multi-Scale Transformer with Cross Attention Fusion
di: Tyagi, Kavyansh, et al.
Pubblicazione: (2026)
di: Tyagi, Kavyansh, et al.
Pubblicazione: (2026)
DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
di: Çalışkan, Halil Hüseyin, et al.
Pubblicazione: (2025)
di: Çalışkan, Halil Hüseyin, et al.
Pubblicazione: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
di: Zinnen, Mathias, et al.
Pubblicazione: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
di: Syed, Shahram Najam, et al.
Pubblicazione: (2025)
Context-Dependent Affordance Computation in Vision-Language Models
di: Farzulla, Murad
Pubblicazione: (2026)
di: Farzulla, Murad
Pubblicazione: (2026)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
di: Su, Yuetong, et al.
Pubblicazione: (2025)
di: Su, Yuetong, et al.
Pubblicazione: (2025)
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
di: Zeng, Zhitao, et al.
Pubblicazione: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
di: Viveiros, André G., et al.
Pubblicazione: (2025)
di: Viveiros, André G., et al.
Pubblicazione: (2025)
OpenFusion++: An Open-vocabulary Real-time Scene Understanding System
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
di: Jin, Xiaofeng, et al.
Pubblicazione: (2025)
Sequence Matters: Harnessing Video Models in 3D Super-Resolution
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
di: Ko, Hyun-kyu, et al.
Pubblicazione: (2024)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
di: Chen, Jingkun, et al.
Pubblicazione: (2025)
Better Schedules for Low Precision Training of Deep Neural Networks
di: Wolfe, Cameron R., et al.
Pubblicazione: (2024)
di: Wolfe, Cameron R., et al.
Pubblicazione: (2024)
Cost-Effective Attention Mechanisms for Low Resource Settings: Necessity & Sufficiency of Linear Transformations
di: Hosseini, Peyman, et al.
Pubblicazione: (2024)
di: Hosseini, Peyman, et al.
Pubblicazione: (2024)
Tricks and Plug-ins for Gradient Boosting in Image Classification
di: Fang, Biyi, et al.
Pubblicazione: (2025)
di: Fang, Biyi, et al.
Pubblicazione: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
di: Svystun, Serhii, et al.
Pubblicazione: (2025)
di: Svystun, Serhii, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ClustViT: Clustering-based Token Merging for Semantic Segmentation
di: Montello, Fabio, et al.
Pubblicazione: (2025) -
GFLAN: Generative Functional Layouts
di: Abouagour, Mohamed, et al.
Pubblicazione: (2025) -
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
di: Chen, Kewei, et al.
Pubblicazione: (2026) -
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
di: Papyan, Narek, et al.
Pubblicazione: (2024) -
Predicting Pedestrian Crossing Behavior in Germany and Japan: Insights into Model Transferability
di: Zhang, Chi, et al.
Pubblicazione: (2024)