Anomaly Object Segmentation with Vision-Language Models for Steel Scrap Recycling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tanaka, Daichi, Karasawa, Takumi, Takenouchi, Shu, Kawakami, Rei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multi-Point Positional Insertion Tuning for Small Object Detection
von: Goto, Kanoko, et al.
Veröffentlicht: (2024)
von: Goto, Kanoko, et al.
Veröffentlicht: (2024)
SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation
von: Neubauer, Melanie, et al.
Veröffentlicht: (2026)
von: Neubauer, Melanie, et al.
Veröffentlicht: (2026)
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
von: An, Zhaoyi, et al.
Veröffentlicht: (2025)
Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts
von: Lee, Eunho, et al.
Veröffentlicht: (2026)
von: Lee, Eunho, et al.
Veröffentlicht: (2026)
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
von: Yajima, Masaru, et al.
Veröffentlicht: (2025)
von: Yajima, Masaru, et al.
Veröffentlicht: (2025)
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
von: Sekikawa, Yusuke, et al.
Veröffentlicht: (2024)
von: Sekikawa, Yusuke, et al.
Veröffentlicht: (2024)
Teacher-Guided Routing for Sparse Vision Mixture-of-Experts
von: Kada, Masahiro, et al.
Veröffentlicht: (2026)
von: Kada, Masahiro, et al.
Veröffentlicht: (2026)
Vision-Language Feature Alignment for Road Anomaly Segmentation
von: He, Zhuolin, et al.
Veröffentlicht: (2026)
von: He, Zhuolin, et al.
Veröffentlicht: (2026)
A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
von: Effendy, Edward, et al.
Veröffentlicht: (2025)
von: Effendy, Edward, et al.
Veröffentlicht: (2025)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Revisiting Few-Shot Object Detection with Vision-Language Models
von: Madan, Anish, et al.
Veröffentlicht: (2023)
von: Madan, Anish, et al.
Veröffentlicht: (2023)
Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models
von: Zheng, Yue, et al.
Veröffentlicht: (2025)
von: Zheng, Yue, et al.
Veröffentlicht: (2025)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
von: Zhou, Zikun, et al.
Veröffentlicht: (2024)
von: Zhou, Zikun, et al.
Veröffentlicht: (2024)
CLIP-Clique: Graph-based Correspondence Matching Augmented by Vision Language Models for Object-based Global Localization
von: Matsuzaki, Shigemichi, et al.
Veröffentlicht: (2024)
von: Matsuzaki, Shigemichi, et al.
Veröffentlicht: (2024)
Geometry Meets Light: Leveraging Geometric Priors for Universal Photometric Stereo under Limited Multi-Illumination Cues
von: Tam, King-Man, et al.
Veröffentlicht: (2025)
von: Tam, King-Man, et al.
Veröffentlicht: (2025)
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
von: Yoshihashi, Ryota, et al.
Veröffentlicht: (2026)
von: Yoshihashi, Ryota, et al.
Veröffentlicht: (2026)
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
von: Feng, Yongchao, et al.
Veröffentlicht: (2025)
von: Feng, Yongchao, et al.
Veröffentlicht: (2025)
Mechanisms of Object Localization in Vision-Language Models
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
Language-Guided Open-World Anomaly Segmentation
von: Reichard, Klara, et al.
Veröffentlicht: (2025)
von: Reichard, Klara, et al.
Veröffentlicht: (2025)
Optimizing Vision-Language Interactions Through Decoder-Only Models
von: Tanaka, Kaito, et al.
Veröffentlicht: (2024)
von: Tanaka, Kaito, et al.
Veröffentlicht: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
von: Song, Jiale, et al.
Veröffentlicht: (2026)
von: Song, Jiale, et al.
Veröffentlicht: (2026)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
von: Traub, Manuel, et al.
Veröffentlicht: (2025)
von: Traub, Manuel, et al.
Veröffentlicht: (2025)
Towards Training-free Anomaly Detection with Vision and Language Foundation Models
von: Zhang, Jinjin, et al.
Veröffentlicht: (2025)
von: Zhang, Jinjin, et al.
Veröffentlicht: (2025)
Anomaly Detection by Adapting a pre-trained Vision Language Model
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2024)
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
von: Zou, Shu, et al.
Veröffentlicht: (2025)
von: Zou, Shu, et al.
Veröffentlicht: (2025)
VLMDiff: Leveraging Vision-Language Models for Multi-Class Anomaly Detection with Diffusion
von: Hicsonmez, Samet, et al.
Veröffentlicht: (2025)
von: Hicsonmez, Samet, et al.
Veröffentlicht: (2025)
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
von: Xu, Meilong, et al.
Veröffentlicht: (2026)
von: Xu, Meilong, et al.
Veröffentlicht: (2026)
One Language-Free Foundation Model Is Enough for Universal Vision Anomaly Detection
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2026)
von: Gao, Bin-Bin, et al.
Veröffentlicht: (2026)
Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
von: Qian, Kun, et al.
Veröffentlicht: (2024)
von: Qian, Kun, et al.
Veröffentlicht: (2024)
Dual-Pathway Circuits of Object Hallucination in Vision-Language Models
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
von: Liu, Jiaxin, et al.
Veröffentlicht: (2026)
Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models
von: Jia, Kaidi, et al.
Veröffentlicht: (2026)
von: Jia, Kaidi, et al.
Veröffentlicht: (2026)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
von: Huang, Bowen, et al.
Veröffentlicht: (2024)
Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches
von: Yu, Qing, et al.
Veröffentlicht: (2024)
von: Yu, Qing, et al.
Veröffentlicht: (2024)
VISA: Reasoning Video Object Segmentation via Large Language Models
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
Text Promptable Surgical Instrument Segmentation with Vision-Language Models
von: Zhou, Zijian, et al.
Veröffentlicht: (2023)
von: Zhou, Zijian, et al.
Veröffentlicht: (2023)
Instruction-Following Evaluation of Large Vision-Language Models
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
XDR-LVLM: An Explainable Vision-Language Large Model for Diabetic Retinopathy Diagnosis
von: Ito, Masato, et al.
Veröffentlicht: (2025)
von: Ito, Masato, et al.
Veröffentlicht: (2025)
CLIPose: Category-Level Object Pose Estimation with Pre-trained Vision-Language Knowledge
von: Lin, Xiao, et al.
Veröffentlicht: (2024)
von: Lin, Xiao, et al.
Veröffentlicht: (2024)
VL4AD: Vision-Language Models Improve Pixel-wise Anomaly Detection
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
von: Zhong, Liangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Multi-Point Positional Insertion Tuning for Small Object Detection
von: Goto, Kanoko, et al.
Veröffentlicht: (2024) -
SteelDS: A High-Resolution Video Dataset of E40 Steel Scrap for Object Detection and Instance Segmentation
von: Neubauer, Melanie, et al.
Veröffentlicht: (2026) -
Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
von: An, Zhaoyi, et al.
Veröffentlicht: (2025) -
Leveraging Multimodal Large Language Models for All-in-One Image Restoration via a Mixture of Frequency Experts
von: Lee, Eunho, et al.
Veröffentlicht: (2026) -
Zero-Shot Peg Insertion: Identifying Mating Holes and Estimating SE(2) Poses with Vision-Language Models
von: Yajima, Masaru, et al.
Veröffentlicht: (2025)