YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Impraimakis, Marios, Vazquez, Daniel, Zhou, Feiyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
A deep learning approach to track eye movements based on events
von: Seth, Chirag, et al.
Veröffentlicht: (2025)
von: Seth, Chirag, et al.
Veröffentlicht: (2025)
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
von: Korolkov, Vasilii
Veröffentlicht: (2025)
von: Korolkov, Vasilii
Veröffentlicht: (2025)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
von: Kambhatla, Akhila, et al.
Veröffentlicht: (2025)
von: Kambhatla, Akhila, et al.
Veröffentlicht: (2025)
Archival Faces: Detection of Faces in Digitized Historical Documents
von: Vaško, Marek, et al.
Veröffentlicht: (2025)
von: Vaško, Marek, et al.
Veröffentlicht: (2025)
ShapBPT: Image Feature Attributions Using Data-Aware Binary Partition Trees
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
von: Rashid, Muhammad, et al.
Veröffentlicht: (2026)
Dual-sensing driving detection model
von: K, Leon C. C., et al.
Veröffentlicht: (2025)
von: K, Leon C. C., et al.
Veröffentlicht: (2025)
The Geometry of Cortical Computation: Manifold Disentanglement and Predictive Dynamics in VCNet
von: Hill, Brennen A., et al.
Veröffentlicht: (2025)
von: Hill, Brennen A., et al.
Veröffentlicht: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
CADE 2.5 - ZeResFDG: Frequency-Decoupled, Rescaled and Zero-Projected Guidance for SD/SDXL Latent Diffusion Models
von: Rychkovskiy, Denis
Veröffentlicht: (2025)
von: Rychkovskiy, Denis
Veröffentlicht: (2025)
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
von: Louison, Nikita, et al.
Veröffentlicht: (2024)
von: Louison, Nikita, et al.
Veröffentlicht: (2024)
A Landmark-Aware Visual Navigation Dataset
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
von: Johnson, Faith, et al.
Veröffentlicht: (2024)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
Objaverse++: Curated 3D Object Dataset with Quality Annotations
von: Lin, Chendi, et al.
Veröffentlicht: (2025)
von: Lin, Chendi, et al.
Veröffentlicht: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
Method of UAV Inspection of Photovoltaic Modules Using Thermal and RGB Data Fusion
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
von: Lysyi, Andrii, et al.
Veröffentlicht: (2025)
Sat-JEPA-Diff: Bridging Self-Supervised Learning and Generative Diffusion for Remote Sensing
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
von: Komurcu, Kursat, et al.
Veröffentlicht: (2026)
Wafer Map Defect Classification Using Autoencoder-Based Data Augmentation and Convolutional Neural Network
von: Bao, Yin-Yin, et al.
Veröffentlicht: (2024)
von: Bao, Yin-Yin, et al.
Veröffentlicht: (2024)
TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
von: Lentsch, Ted, et al.
Veröffentlicht: (2026)
von: Lentsch, Ted, et al.
Veröffentlicht: (2026)
UNION: Unsupervised 3D Object Detection using Object Appearance-based Pseudo-Classes
von: Lentsch, Ted, et al.
Veröffentlicht: (2024)
von: Lentsch, Ted, et al.
Veröffentlicht: (2024)
Image-based Facial Rig Inversion
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
von: Yang, Tianxiang, et al.
Veröffentlicht: (2025)
Attention Gathers, MLPs Compose: A Causal Analysis of an Action-Outcome Circuit in VideoViT
von: Chereddy, Sai V R
Veröffentlicht: (2026)
von: Chereddy, Sai V R
Veröffentlicht: (2026)
AQFusionNet: Multimodal Deep Learning for Air Quality Index Prediction with Imagery and Sensor Data
von: Kushal, Koushik Ahmed, et al.
Veröffentlicht: (2025)
von: Kushal, Koushik Ahmed, et al.
Veröffentlicht: (2025)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
von: Baydemir, Poyraz
Veröffentlicht: (2025)
von: Baydemir, Poyraz
Veröffentlicht: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
QSilk: Micrograin Stabilization and Adaptive Quantile Clipping for Detail-Friendly Latent Diffusion
von: Rychkovskiy, Denis
Veröffentlicht: (2025)
von: Rychkovskiy, Denis
Veröffentlicht: (2025)
Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
von: Roffo, Giorgio, et al.
Veröffentlicht: (2026)
von: Roffo, Giorgio, et al.
Veröffentlicht: (2026)
Polarization-Based Eye Tracking with Personalized Siamese Architectures
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
von: Kalkanli, Beyza, et al.
Veröffentlicht: (2026)
Banana Ripeness Level Classification using a Simple CNN Model Trained with Real and Synthetic Datasets
von: Chuquimarca, Luis, et al.
Veröffentlicht: (2025)
von: Chuquimarca, Luis, et al.
Veröffentlicht: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
von: Yang, Baoyao, et al.
Veröffentlicht: (2025)
μ-Net: A Deep Learning-Based Architecture for μ-CT Segmentation
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
von: Bruno, Pierangela, et al.
Veröffentlicht: (2024)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Revisiting Energy-Based Model for Out-of-Distribution Detection
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
von: Wu, Yifan, et al.
Veröffentlicht: (2024)
MVTamperBench: Evaluating Robustness of Vision-Language Models
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
Fairness Without Labels: Pseudo-Balancing for Bias Mitigation in Face Gender Classification
von: Dong, Haohua, et al.
Veröffentlicht: (2025)
von: Dong, Haohua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026) -
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025) -
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025) -
A deep learning approach to track eye movements based on events
von: Seth, Chirag, et al.
Veröffentlicht: (2025) -
Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
von: Kang, Xueyang, et al.
Veröffentlicht: (2026)