SAViL-Det: Semantic-Aware Vision-Language Model for Multi-Script Text Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zighem, Mohammed-En-Nadhir, Hadid, Abdenour |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recent Advances in Medical Imaging Segmentation: A Survey
von: Bougourzi, Fares, et al.
Veröffentlicht: (2025)
von: Bougourzi, Fares, et al.
Veröffentlicht: (2025)
Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation
von: Bougourzi, Fares, et al.
Veröffentlicht: (2026)
von: Bougourzi, Fares, et al.
Veröffentlicht: (2026)
C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
Harnessing the Power of Large Vision Language Models for Synthetic Image Detection
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
PE-CLIP: A Parameter-Efficient Fine-Tuning of Vision Language Models for Dynamic Facial Expression Recognition
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2025)
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2025)
VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025)
Bi-LORA: A Vision-Language Approach for Synthetic Image Detection
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)
SegDT: A Diffusion Transformer-Based Segmentation Model for Medical Imaging
von: Bekhouche, Salah Eddine, et al.
Veröffentlicht: (2025)
von: Bekhouche, Salah Eddine, et al.
Veröffentlicht: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2024)
von: Kulkarni, Yogesh, et al.
Veröffentlicht: (2024)
SPARK-IL: Spectral Retrieval-Augmented RAG for Knowledge-driven Deepfake Detection via Incremental Learning
von: Eutamene, Hessen Bougueffa, et al.
Veröffentlicht: (2026)
von: Eutamene, Hessen Bougueffa, et al.
Veröffentlicht: (2026)
Shuffle Vision Transformer: Lightweight, Fast and Efficient Recognition of Driver Facial Expression
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2024)
von: Saadi, Ibtissam, et al.
Veröffentlicht: (2024)
Conflict-Aware Multimodal Fusion for Ambivalence and Hesitancy Recognition
von: Bekhouche, Salah Eddine, et al.
Veröffentlicht: (2026)
von: Bekhouche, Salah Eddine, et al.
Veröffentlicht: (2026)
DeeCLIP: A Robust and Generalizable Transformer-Based Framework for Detecting AI-Generated Images
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
Semantic-Aware Ship Detection with Vision-Language Integration
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
von: Li, Jiahao, et al.
Veröffentlicht: (2025)
SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
von: Xu, Qing, et al.
Veröffentlicht: (2025)
von: Xu, Qing, et al.
Veröffentlicht: (2025)
Face to Cartoon Incremental Super-Resolution using Knowledge Distillation
von: Devkatte, Trinetra, et al.
Veröffentlicht: (2024)
von: Devkatte, Trinetra, et al.
Veröffentlicht: (2024)
RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
von: Keita, Mamadou, et al.
Veröffentlicht: (2025)
RGBX-DiffusionDet: A Framework for Multi-Modal RGB-X Object Detection Using DiffusionDet
von: Orfaig, Eliraz, et al.
Veröffentlicht: (2025)
von: Orfaig, Eliraz, et al.
Veröffentlicht: (2025)
RF-HiT: Rectified Flow Hierarchical Transformer for General Medical Image Segmentation
von: Djouama, Ahmed Marouane, et al.
Veröffentlicht: (2026)
von: Djouama, Ahmed Marouane, et al.
Veröffentlicht: (2026)
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
von: Qi, Jianing, et al.
Veröffentlicht: (2025)
NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
von: Huang, Chenxi, et al.
Veröffentlicht: (2024)
FMG-Det: Foundation Model Guided Robust Object Detection
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
von: Seutin, Corentin, et al.
Veröffentlicht: (2026)
von: Seutin, Corentin, et al.
Veröffentlicht: (2026)
VP-Hype: A Hybrid Mamba-Transformer Framework with Visual-Textual Prompting for Hyperspectral Image Classification
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2026)
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2026)
VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
von: Qi, Yu, et al.
Veröffentlicht: (2025)
von: Qi, Yu, et al.
Veröffentlicht: (2025)
LMM-Det: Make Large Multimodal Models Excel in Object Detection
von: Li, Jincheng, et al.
Veröffentlicht: (2025)
von: Li, Jincheng, et al.
Veröffentlicht: (2025)
Can Visual Mamba Improve AI-Generated Image Detection? An In-Depth Investigation
von: Keita, Mamadou, et al.
Veröffentlicht: (2026)
von: Keita, Mamadou, et al.
Veröffentlicht: (2026)
RemDet: Rethinking Efficient Model Design for UAV Object Detection
von: Li, Chen, et al.
Veröffentlicht: (2024)
von: Li, Chen, et al.
Veröffentlicht: (2024)
DetRefiner: Model-Agnostic Detection Refinement with Feature Fusion Transformer
von: Okazaki, Soichiro, et al.
Veröffentlicht: (2026)
von: Okazaki, Soichiro, et al.
Veröffentlicht: (2026)
Detecting Text Manipulation in Images using Vision Language Models
von: Vidit, Vidit, et al.
Veröffentlicht: (2025)
von: Vidit, Vidit, et al.
Veröffentlicht: (2025)
DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
von: Gare, Gautam Rajendrakumar, et al.
Veröffentlicht: (2026)
UniDet3D: Multi-dataset Indoor 3D Object Detection
von: Kolodiazhnyi, Maksim, et al.
Veröffentlicht: (2024)
von: Kolodiazhnyi, Maksim, et al.
Veröffentlicht: (2024)
BUSTR: Breast Ultrasound Text Reporting with a Descriptor-Aware Vision-Language Model
von: Mohammed, Rawa, et al.
Veröffentlicht: (2025)
von: Mohammed, Rawa, et al.
Veröffentlicht: (2025)
Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
von: Yang, Zhongyu, et al.
Veröffentlicht: (2025)
von: Yang, Zhongyu, et al.
Veröffentlicht: (2025)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
von: Li, Yuxuan, et al.
Veröffentlicht: (2024)
Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
von: Bai, Weimin, et al.
Veröffentlicht: (2025)
Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features
von: Sengupta, Saurav, et al.
Veröffentlicht: (2025)
von: Sengupta, Saurav, et al.
Veröffentlicht: (2025)
Training-Free Semantic Multi-Object Tracking with Vision-Language Models
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
von: Bonat, Laurence, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Recent Advances in Medical Imaging Segmentation: A Survey
von: Bougourzi, Fares, et al.
Veröffentlicht: (2025) -
Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation
von: Bougourzi, Fares, et al.
Veröffentlicht: (2026) -
C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
von: Sellam, Abdellah Zakaria, et al.
Veröffentlicht: (2025) -
FIDAVL: Fake Image Detection and Attribution using Vision-Language Model
von: Keita, Mamadou, et al.
Veröffentlicht: (2024) -
Harnessing the Power of Large Vision Language Models for Synthetic Image Detection
von: Keita, Mamadou, et al.
Veröffentlicht: (2024)