Discrete Latent Perspective Learning for Segmentation and Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Deyi, Zhao, Feng, Zhu, Lanyun, Jin, Wenwei, Lu, Hongtao, Ye, Jieping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PPTFormer: Pseudo Multi-Perspective Transformer for UAV Segmentation
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
Structural and Statistical Texture Knowledge Distillation and Learning for Segmentation
von: Ji, Deyi, et al.
Veröffentlicht: (2025)
von: Ji, Deyi, et al.
Veröffentlicht: (2025)
LLaFS: When Large Language Models Meet Few-Shot Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2023)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2023)
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
View-Centric Multi-Object Tracking with Homographic Matching in Moving UAV
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
IBD: Alleviating Hallucinations in Large Vision-Language Models via Image-Biased Decoding
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2024)
ChangeNet: Multi-Temporal Asymmetric Change Detection Dataset
von: Ji, Deyi, et al.
Veröffentlicht: (2023)
von: Ji, Deyi, et al.
Veröffentlicht: (2023)
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
SAM3-Adapter: Efficient Adaptation of Segment Anything 3 for Camouflage Object Segmentation, Shadow Detection, and Medical Image Segmentation
von: Chen, Tianrun, et al.
Veröffentlicht: (2025)
von: Chen, Tianrun, et al.
Veröffentlicht: (2025)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
POPEN: Preference-Based Optimization and Ensemble for LVLM-Based Reasoning Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025)
Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation
von: Ji, Deyi, et al.
Veröffentlicht: (2023)
von: Ji, Deyi, et al.
Veröffentlicht: (2023)
Breaking the Box: Enhancing Remote Sensing Image Segmentation with Freehand Sketches
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
Video-Zero: Self-Evolution Video Understanding
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
von: Zhang, Ruixu, et al.
Veröffentlicht: (2026)
Let Human Sketches Help: Empowering Challenging Image Segmentation Task with Freehand Sketches
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
von: Liu, Xuanyi, et al.
Veröffentlicht: (2026)
xLSTM-UNet can be an Effective 2D & 3D Medical Image Segmentation Backbone with Vision-LSTM (ViL) better than its Mamba Counterpart
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
von: Zang, Ying, et al.
Veröffentlicht: (2026)
von: Zang, Ying, et al.
Veröffentlicht: (2026)
Hybrid Mamba for Few-Shot Segmentation
von: Xu, Qianxiong, et al.
Veröffentlicht: (2024)
von: Xu, Qianxiong, et al.
Veröffentlicht: (2024)
Unlocking the Power of SAM 2 for Few-Shot Segmentation
von: Xu, Qianxiong, et al.
Veröffentlicht: (2025)
von: Xu, Qianxiong, et al.
Veröffentlicht: (2025)
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
von: Zang, Ying, et al.
Veröffentlicht: (2026)
von: Zang, Ying, et al.
Veröffentlicht: (2026)
From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching
von: Zang, Ying, et al.
Veröffentlicht: (2025)
von: Zang, Ying, et al.
Veröffentlicht: (2025)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
RESMatch: Referring Expression Segmentation in a Semi-Supervised Manner
von: Zang, Ying, et al.
Veröffentlicht: (2024)
von: Zang, Ying, et al.
Veröffentlicht: (2024)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
von: Tong, Jintao, et al.
Veröffentlicht: (2025)
HD-VGGT: High-Resolution Visual Geometry Transformer
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
von: Chen, Tianrun, et al.
Veröffentlicht: (2026)
WaterWave: Bridging Underwater Image Enhancement into Video Streams via Wavelet-based Temporal Consistency Field
von: Zhu, Qi, et al.
Veröffentlicht: (2025)
von: Zhu, Qi, et al.
Veröffentlicht: (2025)
Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
von: Zhu, Deyi, et al.
Veröffentlicht: (2026)
Tree-of-Table: Unleashing the Power of LLMs for Enhanced Large-Scale Table Understanding
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
von: Ji, Deyi, et al.
Veröffentlicht: (2024)
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
von: Sheng, Hualian, et al.
Veröffentlicht: (2024)
von: Sheng, Hualian, et al.
Veröffentlicht: (2024)
ESOD: Efficient Small Object Detection on High-Resolution Images
von: Liu, Kai, et al.
Veröffentlicht: (2024)
von: Liu, Kai, et al.
Veröffentlicht: (2024)
Category-Extensible Out-of-Distribution Detection via Hierarchical Context Descriptions
von: Liu, Kai, et al.
Veröffentlicht: (2024)
von: Liu, Kai, et al.
Veröffentlicht: (2024)
Rethinking Out-of-Distribution Detection on Imbalanced Data Distribution
von: Liu, Kai, et al.
Veröffentlicht: (2024)
von: Liu, Kai, et al.
Veröffentlicht: (2024)
CoL3D: Collaborative Learning of Single-view Depth and Camera Intrinsics for Metric 3D Shape Recovery
von: Zhang, Chenghao, et al.
Veröffentlicht: (2025)
von: Zhang, Chenghao, et al.
Veröffentlicht: (2025)
LADMIM: Logical Anomaly Detection with Masked Image Modeling in Discrete Latent Space
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2024)
von: Sakai, Shunsuke, et al.
Veröffentlicht: (2024)
Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation
von: Zhu, Vince, et al.
Veröffentlicht: (2024)
von: Zhu, Vince, et al.
Veröffentlicht: (2024)
KAnoCLIP: Zero-Shot Anomaly Detection through Knowledge-Driven Prompt Learning and Enhanced Cross-Modal Integration
von: Li, Chengyuan, et al.
Veröffentlicht: (2025)
von: Li, Chengyuan, et al.
Veröffentlicht: (2025)
Self-Learning Symmetric Multi-view Probabilistic Clustering
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
Video Generation with Predictive Latents
von: Zhao, Yian, et al.
Veröffentlicht: (2026)
von: Zhao, Yian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
PPTFormer: Pseudo Multi-Perspective Transformer for UAV Segmentation
von: Ji, Deyi, et al.
Veröffentlicht: (2024) -
Structural and Statistical Texture Knowledge Distillation and Learning for Segmentation
von: Ji, Deyi, et al.
Veröffentlicht: (2025) -
LLaFS: When Large Language Models Meet Few-Shot Segmentation
von: Zhu, Lanyun, et al.
Veröffentlicht: (2023) -
Not Every Patch is Needed: Towards a More Efficient and Effective Backbone for Video-based Person Re-identification
von: Zhu, Lanyun, et al.
Veröffentlicht: (2025) -
View-Centric Multi-Object Tracking with Homographic Matching in Moving UAV
von: Ji, Deyi, et al.
Veröffentlicht: (2024)