COCONut: Modernizing COCO Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Xueqing, Yu, Qihang, Wang, Peng, Shen, Xiaohui, Chen, Liang-Chieh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
von: Deng, Xueqing, et al.
Veröffentlicht: (2025)
von: Deng, Xueqing, et al.
Veröffentlicht: (2025)
Randomized Autoregressive Visual Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023)
von: He, Ju, et al.
Veröffentlicht: (2023)
An Image is Worth 32 Tokens for Reconstruction and Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
von: Yu, Qihang, et al.
Veröffentlicht: (2024)
MaskBit: Embedding-free Image Generation via Bit Tokens
von: Weber, Mark, et al.
Veröffentlicht: (2024)
von: Weber, Mark, et al.
Veröffentlicht: (2024)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
von: Athar, Ali, et al.
Veröffentlicht: (2024)
von: Athar, Ali, et al.
Veröffentlicht: (2024)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
von: Liu, Qihao, et al.
Veröffentlicht: (2024)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
von: Ren, Sucheng, et al.
Veröffentlicht: (2026)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
von: Shin, Inkyu, et al.
Veröffentlicht: (2024)
von: Shin, Inkyu, et al.
Veröffentlicht: (2024)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
von: Kim, Dongwon, et al.
Veröffentlicht: (2025)
From COCO to COCO-FP: A Deep Dive into Background False Positives for COCO Detectors
von: Liu, Longfei, et al.
Veröffentlicht: (2024)
von: Liu, Longfei, et al.
Veröffentlicht: (2024)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
von: He, Ju, et al.
Veröffentlicht: (2025)
von: He, Ju, et al.
Veröffentlicht: (2025)
1.58-bit FLUX
von: Yang, Chenglin, et al.
Veröffentlicht: (2024)
von: Yang, Chenglin, et al.
Veröffentlicht: (2024)
Stable Diffusion for Data Augmentation in COCO and Weed Datasets
von: Deng, Boyang
Veröffentlicht: (2023)
von: Deng, Boyang
Veröffentlicht: (2023)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
von: Ren, Sucheng, et al.
Veröffentlicht: (2025)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
von: Wei, Wenbo, et al.
Veröffentlicht: (2024)
von: Wei, Wenbo, et al.
Veröffentlicht: (2024)
Autoregressive Image Generation with Masked Bit Modeling
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
von: Yu, Qihang, et al.
Veröffentlicht: (2026)
3D-COCO: extension of MS-COCO dataset for image detection and 3D reconstruction modules
von: Bideaux, Maxence, et al.
Veröffentlicht: (2024)
von: Bideaux, Maxence, et al.
Veröffentlicht: (2024)
COCO is "ALL'' You Need for Visual Instruction Fine-tuning
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
RoCOCO: Robustness Benchmark of MS-COCO to Stress-test Image-Text Matching Models
von: Park, Seulki, et al.
Veröffentlicht: (2023)
von: Park, Seulki, et al.
Veröffentlicht: (2023)
Semi-Supervised Multi-Modal Medical Image Segmentation for Complex Situations
von: Meng, Dongdong, et al.
Veröffentlicht: (2025)
von: Meng, Dongdong, et al.
Veröffentlicht: (2025)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
von: Kim, Seungwook, et al.
Veröffentlicht: (2024)
von: Kim, Seungwook, et al.
Veröffentlicht: (2024)
Benchmarking Object Detectors with COCO: A New Path Forward
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
von: Singh, Shweta, et al.
Veröffentlicht: (2024)
ObjVariantEnsemble: Advancing Point Cloud LLM Evaluation in Challenging Scenes with Subtly Distinguished Objects
von: Cao, Qihang, et al.
Veröffentlicht: (2024)
von: Cao, Qihang, et al.
Veröffentlicht: (2024)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
Multi-Sequence Parotid Gland Lesion Segmentation via Expert Text-Guided Segment Anything Model
von: Wu, Zhongyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhongyuan, et al.
Veröffentlicht: (2025)
Full-Duplex Strategy for Video Object Segmentation
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
von: Ji, Ge-Peng, et al.
Veröffentlicht: (2021)
Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
von: Zhang, Xueqing, et al.
Veröffentlicht: (2024)
Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
von: Zhang, Jiangning, et al.
Veröffentlicht: (2024)
COCO-Tree: Compositional Hierarchical Concept Trees for Enhanced Reasoning in Vision Language Models
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2025)
ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
CmFNet: Cross-modal Fusion Network for Weakly-supervised Segmentation of Medical Images
von: Meng, Dongdong, et al.
Veröffentlicht: (2025)
von: Meng, Dongdong, et al.
Veröffentlicht: (2025)
Deeply Supervised Flow-Based Generative Models
von: Shin, Inkyu, et al.
Veröffentlicht: (2025)
von: Shin, Inkyu, et al.
Veröffentlicht: (2025)
Segment Anything for Video: A Comprehensive Review of Video Object Segmentation and Tracking from Past to Future
von: Xu, Guoping, et al.
Veröffentlicht: (2025)
von: Xu, Guoping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
von: Deng, Xueqing, et al.
Veröffentlicht: (2025) -
Randomized Autoregressive Visual Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024) -
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
von: He, Ju, et al.
Veröffentlicht: (2023) -
An Image is Worth 32 Tokens for Reconstruction and Generation
von: Yu, Qihang, et al.
Veröffentlicht: (2024) -
MaskBit: Embedding-free Image Generation via Bit Tokens
von: Weber, Mark, et al.
Veröffentlicht: (2024)