Dictionary-based Framework for Interpretable and Consistent Object Parsing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Tiezheng, Yu, Qihang, Yuille, Alan, He, Ju |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023)
von: He, Ju, et al.
Veröffentlicht: (2023)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Beyond Masks: The Case for Medical Image Parsing
von: Gupta, Siddharth, et al.
Veröffentlicht: (2026)
von: Gupta, Siddharth, et al.
Veröffentlicht: (2026)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Leveraging AI Predicted and Expert Revised Annotations in Interactive Segmentation: Continual Tuning or Full Training?
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2024)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
von: Zhang, Tiezheng, et al.
Veröffentlicht: (2025)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
von: Paul, Soumava, et al.
Veröffentlicht: (2026)
ViT-5: Vision Transformers for The Mid-2020s
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
von: Xie, Yunfei, et al.
Veröffentlicht: (2024)
NOVUM: Neural Object Volumes for Robust Object Classification
von: Jesslen, Artur, et al.
Veröffentlicht: (2023)
von: Jesslen, Artur, et al.
Veröffentlicht: (2023)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
von: Yuille, Alan, et al.
Veröffentlicht: (2026)
von: Yuille, Alan, et al.
Veröffentlicht: (2026)
Interaction-aware Representation Modeling with Co-occurrence Consistency for Egocentric Hand-Object Parsing
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
von: Su, Yuejiao, et al.
Veröffentlicht: (2026)
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
von: Wang, Xingrui, et al.
Veröffentlicht: (2025)
ParseCaps: An Interpretable Parsing Capsule Network for Medical Image Diagnosis
von: Geng, Xinyu, et al.
Veröffentlicht: (2024)
von: Geng, Xinyu, et al.
Veröffentlicht: (2024)
Efficient Large Multi-modal Models via Visual Context Compression
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
von: He, Ju, et al.
Veröffentlicht: (2025)
von: He, Ju, et al.
Veröffentlicht: (2025)
GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
von: He, Zhenghao, et al.
Veröffentlicht: (2025)
von: He, Zhenghao, et al.
Veröffentlicht: (2025)
AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
von: Zhou, Qihang, et al.
Veröffentlicht: (2023)
von: Zhou, Qihang, et al.
Veröffentlicht: (2023)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
von: Kaushik, Prakhar, et al.
Veröffentlicht: (2024)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
von: Zhang, Guofeng, et al.
Veröffentlicht: (2026)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
von: Paul, Soumava, et al.
Veröffentlicht: (2024)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
von: Wang, Feng, et al.
Veröffentlicht: (2023)
von: Wang, Feng, et al.
Veröffentlicht: (2023)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
von: Chen, Yixiong, et al.
Veröffentlicht: (2024)
von: Chen, Yixiong, et al.
Veröffentlicht: (2024)
Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
DINeMo: Learning Neural Mesh Models with no 3D Annotations
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
von: Guo, Weijie, et al.
Veröffentlicht: (2025)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
von: Ma, Wufei, et al.
Veröffentlicht: (2024)
Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xiao, et al.
Veröffentlicht: (2025)
Randomized Autoregressive Visual Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
Learning a Category-level Object Pose Estimator without Pose Annotations
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
von: Tian, Fengrui, et al.
Veröffentlicht: (2024)
Parsing Objects at a Finer Granularity: A Survey
von: Zhao, Yifan, et al.
Veröffentlicht: (2022)
von: Zhao, Yifan, et al.
Veröffentlicht: (2022)
AbdomenAtlas-8K: Annotating 8,000 CT Volumes for Multi-Organ Segmentation in Three Weeks
von: Qu, Chongyu, et al.
Veröffentlicht: (2023)
von: Qu, Chongyu, et al.
Veröffentlicht: (2023)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects
von: Cao, Qihang, et al.
Veröffentlicht: (2024)
von: Cao, Qihang, et al.
Veröffentlicht: (2024)
TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
von: Fang, Hao, et al.
Veröffentlicht: (2025)
von: Fang, Hao, et al.
Veröffentlicht: (2025)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
von: Ma, Wufei, et al.
Veröffentlicht: (2026)
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
von: Zhang, Jinlu, et al.
Veröffentlicht: (2025)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023) -
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
von: Ren, Sucheng, et al.
Veröffentlicht: (2025) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025) -
Beyond Masks: The Case for Medical Image Parsing
von: Gupta, Siddharth, et al.
Veröffentlicht: (2026) -
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)