Gespeichert in:
| Hauptverfasser: | Lin, Dongheng, Qu, Mengxue, Han, Kunyang, Jiao, Jianbo, Jin, Xiaojie, Wei, Yunchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2511.00962 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
von: Lin, Dongheng, et al.
Veröffentlicht: (2025)
von: Lin, Dongheng, et al.
Veröffentlicht: (2025)
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
von: Han, Kunyang, et al.
Veröffentlicht: (2024)
von: Han, Kunyang, et al.
Veröffentlicht: (2024)
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
von: Zhao, Shifang, et al.
Veröffentlicht: (2025)
von: Zhao, Shifang, et al.
Veröffentlicht: (2025)
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
von: Chen, Wenxin, et al.
Veröffentlicht: (2025)
von: Chen, Wenxin, et al.
Veröffentlicht: (2025)
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
von: Ren, Zhongwei, et al.
Veröffentlicht: (2025)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2025)
Lane2Seq: Towards Unified Lane Detection via Sequence Generation
von: Zhou, Kunyang
Veröffentlicht: (2024)
von: Zhou, Kunyang
Veröffentlicht: (2024)
Bridge the Points: Graph-based Few-shot Segment Anything Semantically
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
von: Zhang, Anqi, et al.
Veröffentlicht: (2024)
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
AnyAnomaly: Zero-Shot Customizable Video Anomaly Detection with LVLM
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2025)
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2025)
Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection
von: Qu, Zhen, et al.
Veröffentlicht: (2025)
von: Qu, Zhen, et al.
Veröffentlicht: (2025)
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
von: Ren, Zhongwei, et al.
Veröffentlicht: (2026)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2026)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
AnomalyPainter: Vision-Language-Diffusion Synergy for Zero-Shot Realistic and Diverse Industrial Anomaly Synthesis
von: Lai, Zhangyu, et al.
Veröffentlicht: (2025)
von: Lai, Zhangyu, et al.
Veröffentlicht: (2025)
PackCache: A Training-Free Acceleration Method for Unified Autoregressive Video Generation via Compact KV-Cache
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
von: Li, Kunyang, et al.
Veröffentlicht: (2026)
CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
von: Fang, Kai, et al.
Veröffentlicht: (2025)
von: Fang, Kai, et al.
Veröffentlicht: (2025)
Rethinking MLLM Itself as a Segmenter with a Single Segmentation Token
von: Zhang, Anqi, et al.
Veröffentlicht: (2026)
von: Zhang, Anqi, et al.
Veröffentlicht: (2026)
CoPS: Conditional Prompt Synthesis for Zero-Shot Anomaly Detection
von: Chen, Qiyu, et al.
Veröffentlicht: (2025)
von: Chen, Qiyu, et al.
Veröffentlicht: (2025)
Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning
von: Liu, Man, et al.
Veröffentlicht: (2024)
von: Liu, Man, et al.
Veröffentlicht: (2024)
Exploring Image Representation with Decoupled Classical Visual Descriptors
von: Qu, Chenyuan, et al.
Veröffentlicht: (2025)
von: Qu, Chenyuan, et al.
Veröffentlicht: (2025)
HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
von: Meng, Yihao, et al.
Veröffentlicht: (2025)
Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
von: Deng, Hanqiu, et al.
Veröffentlicht: (2023)
von: Deng, Hanqiu, et al.
Veröffentlicht: (2023)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
von: Dai, Zunkai, et al.
Veröffentlicht: (2026)
Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
von: Li, Kaiqiang, et al.
Veröffentlicht: (2026)
von: Li, Kaiqiang, et al.
Veröffentlicht: (2026)
MAGIC: Map-Guided Few-Shot Audio-Visual Acoustics Modeling
von: Huang, Diwei, et al.
Veröffentlicht: (2024)
von: Huang, Diwei, et al.
Veröffentlicht: (2024)
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
von: Tan, Chuangchuang, et al.
Veröffentlicht: (2025)
VISD: Enhancing Video Reasoning via Structured Self-Distillation
von: Lin, Hao, et al.
Veröffentlicht: (2026)
von: Lin, Hao, et al.
Veröffentlicht: (2026)
Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
von: Yin, Xiaojie, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Adapter Shield: A Unified Framework with Built-in Authentication for Preventing Unauthorized Zero-Shot Image-to-Image Generation
von: Jia, Jun, et al.
Veröffentlicht: (2025)
von: Jia, Jun, et al.
Veröffentlicht: (2025)
Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language Models
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
von: Xu, Jiacong, et al.
Veröffentlicht: (2025)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
von: Jin, Woojeong, et al.
Veröffentlicht: (2026)
von: Jin, Woojeong, et al.
Veröffentlicht: (2026)
The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
Are Video Models Emerging as Zero-Shot Learners and Reasoners in Medical Imaging?
von: Lai, Yuxiang, et al.
Veröffentlicht: (2025)
von: Lai, Yuxiang, et al.
Veröffentlicht: (2025)
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
von: Kao, Shiu-hong, et al.
Veröffentlicht: (2025)
DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models
von: Yeh, Chang-Han, et al.
Veröffentlicht: (2024)
von: Yeh, Chang-Han, et al.
Veröffentlicht: (2024)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
FlowSeg: Dynamic Semantic Guidance for LLM-Conditioned Segmentation
von: Zhang, Zekang, et al.
Veröffentlicht: (2026)
von: Zhang, Zekang, et al.
Veröffentlicht: (2026)
On the Problem of Consistent Anomalies in Zero-Shot Industrial Anomaly Detection
von: Le-Gia, Tai, et al.
Veröffentlicht: (2025)
von: Le-Gia, Tai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Time Tells Us? An Explorative Study of Time Awareness Learned from Static Images
von: Lin, Dongheng, et al.
Veröffentlicht: (2025) -
ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
von: Han, Kunyang, et al.
Veröffentlicht: (2024) -
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024) -
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
von: Zhao, Shifang, et al.
Veröffentlicht: (2025) -
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
von: Kang, Weitai, et al.
Veröffentlicht: (2024)