Towards Open World Detection: A Survey
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bulzan, Andrei-Stefan, Cernazanu-Glavan, Cosmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Detection and Measurement of Hailstones with Multimodal Large Language Models
von: Alker, Moritz, et al.
Veröffentlicht: (2025)
von: Alker, Moritz, et al.
Veröffentlicht: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025)
MB-DSMIL-CL-PL: Scalable Weakly Supervised Ovarian Cancer Subtype Classification and Localisation Using Contrastive and Prototype Learning with Frozen Patch Features
von: Jenkins, Marcus, et al.
Veröffentlicht: (2026)
von: Jenkins, Marcus, et al.
Veröffentlicht: (2026)
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
von: Jung, Seoik, et al.
Veröffentlicht: (2025)
A Multi-Camera Vision-Based Approach for Fine-Grained Assembly Quality Control
von: Nazeri, Ali, et al.
Veröffentlicht: (2025)
von: Nazeri, Ali, et al.
Veröffentlicht: (2025)
WildfireVLM: AI-powered Analysis for Early Wildfire Detection and Risk Assessment Using Satellite Imagery
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
von: Ayanzadeh, Aydin, et al.
Veröffentlicht: (2026)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
The Impact of Image Resolution on Face Detection: A Comparative Analysis of MTCNN, YOLOv XI and YOLOv XII models
von: Ömercikoğlu, Ahmet Can, et al.
Veröffentlicht: (2025)
von: Ömercikoğlu, Ahmet Can, et al.
Veröffentlicht: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
von: Lyu, Wenbo, et al.
Veröffentlicht: (2025)
GraphTEN: Graph Enhanced Texture Encoding Network
von: Peng, Bo, et al.
Veröffentlicht: (2025)
von: Peng, Bo, et al.
Veröffentlicht: (2025)
AniMatrix: An Anime Video Generation Model that Thinks in Art, Not Physics
von: Tencent HY Team
Veröffentlicht: (2026)
von: Tencent HY Team
Veröffentlicht: (2026)
Addressing Issues with Working Memory in Video Object Segmentation
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
von: Bromley, Clayton, et al.
Veröffentlicht: (2024)
See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
von: Dong, Zixuan, et al.
Veröffentlicht: (2025)
von: Dong, Zixuan, et al.
Veröffentlicht: (2025)
An M-Health Algorithmic Approach to Identify and Assess Physiotherapy Exercises in Real Time
von: Kandylakis, Stylianos, et al.
Veröffentlicht: (2025)
von: Kandylakis, Stylianos, et al.
Veröffentlicht: (2025)
DeepFusionNet: Autoencoder-Based Low-Light Image Enhancement and Super-Resolution
von: Çalışkan, Halil Hüseyin, et al.
Veröffentlicht: (2025)
von: Çalışkan, Halil Hüseyin, et al.
Veröffentlicht: (2025)
Unlocking UML Class Diagram Understanding in Vision Language Models
von: Naboichenko, Artem, et al.
Veröffentlicht: (2026)
von: Naboichenko, Artem, et al.
Veröffentlicht: (2026)
HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
von: Gautam, Sushant, et al.
Veröffentlicht: (2025)
Training for X-Ray Vision: Amodal Segmentation, Amodal Content Completion, and View-Invariant Object Representation from Multi-Camera Video
von: Moore, Alexander, et al.
Veröffentlicht: (2025)
von: Moore, Alexander, et al.
Veröffentlicht: (2025)
Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions
von: Hermann, Nicolai, et al.
Veröffentlicht: (2024)
von: Hermann, Nicolai, et al.
Veröffentlicht: (2024)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
von: Sajjadi, Arash, et al.
Veröffentlicht: (2025)
VideoHEDGE: Entropy-Based Hallucination Detection for Video-VLMs via Semantic Clustering and Spatiotemporal Perturbations
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
von: Gautam, Sushant, et al.
Veröffentlicht: (2026)
ARTPS: Depth-Enhanced Hybrid Anomaly Detection and Learnable Curiosity Score for Autonomous Rover Target Prioritization
von: Baydemir, Poyraz
Veröffentlicht: (2025)
von: Baydemir, Poyraz
Veröffentlicht: (2025)
How to train your VAE
von: Rivera, Mariano
Veröffentlicht: (2023)
von: Rivera, Mariano
Veröffentlicht: (2023)
TACIT Benchmark: A Programmatic Visual Reasoning Benchmark for Generative and Discriminative Models
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
JotlasNet: Joint Tensor Low-Rank and Attention-based Sparse Unrolling Network for Accelerating Dynamic MRI
von: Zhang, Yinghao, et al.
Veröffentlicht: (2025)
von: Zhang, Yinghao, et al.
Veröffentlicht: (2025)
N-DriverMotion: Driver motion learning and prediction using an event-based camera and directly trained spiking neural networks on Loihi 2
von: Chung, Hyo Jong, et al.
Veröffentlicht: (2024)
von: Chung, Hyo Jong, et al.
Veröffentlicht: (2024)
A Survey on Dynamic Neural Networks: from Computer Vision to Multi-modal Sensor Fusion
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
YOLO Ensemble for UAV-based Multispectral Defect Detection in Wind Turbine Components
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
von: Svystun, Serhii, et al.
Veröffentlicht: (2025)
Fixed-Threshold Evaluation of a Hybrid CNN-ViT for AI-Generated Image Detection Across Photos and Art
von: Khan, Md Ashik, et al.
Veröffentlicht: (2025)
von: Khan, Md Ashik, et al.
Veröffentlicht: (2025)
Representation Paradigms in AI-based 3D Radiological Image Reconstruction: A Systematic Review
von: Yang, Yuezhe, et al.
Veröffentlicht: (2025)
von: Yang, Yuezhe, et al.
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
GFLAN: Generative Functional Layouts
von: Abouagour, Mohamed, et al.
Veröffentlicht: (2025)
von: Abouagour, Mohamed, et al.
Veröffentlicht: (2025)
Fruit Deformity Classification through Single-Input and Multi-Input Architectures based on CNN Models using Real and Synthetic Images
von: Beltran, Tommy D., et al.
Veröffentlicht: (2024)
von: Beltran, Tommy D., et al.
Veröffentlicht: (2024)
Improving Object Detection for Time-Lapse Imagery Using Temporal Features in Wildlife Monitoring
von: Jenkins, Marcus, et al.
Veröffentlicht: (2024)
von: Jenkins, Marcus, et al.
Veröffentlicht: (2024)
SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
von: Hu, Yangliu, et al.
Veröffentlicht: (2025)
von: Hu, Yangliu, et al.
Veröffentlicht: (2025)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
ISO/IEC-Compliant Match-on-Card Face Verification with Short Binary Templates
von: Ganmati, Abdelilah, et al.
Veröffentlicht: (2025)
von: Ganmati, Abdelilah, et al.
Veröffentlicht: (2025)
TRACES: Temporal Recall with Contextual Embeddings for Real-Time Video Anomaly Detection
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
von: Siddiqui, Yousuf Ahmed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Detection and Measurement of Hailstones with Multimodal Large Language Models
von: Alker, Moritz, et al.
Veröffentlicht: (2025) -
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
von: Gautam, Sushant, et al.
Veröffentlicht: (2025) -
Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset
von: Zinnen, Mathias, et al.
Veröffentlicht: (2025) -
MB-DSMIL-CL-PL: Scalable Weakly Supervised Ovarian Cancer Subtype Classification and Localisation Using Contrastive and Prototype Learning with Frozen Patch Features
von: Jenkins, Marcus, et al.
Veröffentlicht: (2026) -
Short-Window Sliding Learning for Real-Time Violence Detection via LLM-based Auto-Labeling
von: Jung, Seoik, et al.
Veröffentlicht: (2025)