Enregistré dans:
| Auteurs principaux: | Deng, Xueqing, Yu, Qihang, Athar, Ali, Yang, Chenglin, Yang, Linjie, Jin, Xiaojie, Shen, Xiaohui, Chen, Liang-Chieh |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2502.02589 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
COCONut: Modernizing COCO Segmentation
par: Deng, Xueqing, et autres
Publié: (2024)
par: Deng, Xueqing, et autres
Publié: (2024)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
par: Athar, Ali, et autres
Publié: (2024)
par: Athar, Ali, et autres
Publié: (2024)
Randomized Autoregressive Visual Generation
par: Yu, Qihang, et autres
Publié: (2024)
par: Yu, Qihang, et autres
Publié: (2024)
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
par: He, Ju, et autres
Publié: (2023)
par: He, Ju, et autres
Publié: (2023)
PanDepth: Joint Panoptic Segmentation and Depth Completion
par: Lagos, Juan, et autres
Publié: (2022)
par: Lagos, Juan, et autres
Publié: (2022)
1.58-bit FLUX
par: Yang, Chenglin, et autres
Publié: (2024)
par: Yang, Chenglin, et autres
Publié: (2024)
An Image is Worth 32 Tokens for Reconstruction and Generation
par: Yu, Qihang, et autres
Publié: (2024)
par: Yu, Qihang, et autres
Publié: (2024)
MaskBit: Embedding-free Image Generation via Bit Tokens
par: Weber, Mark, et autres
Publié: (2024)
par: Weber, Mark, et autres
Publié: (2024)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
par: Kim, Dongwon, et autres
Publié: (2025)
par: Kim, Dongwon, et autres
Publié: (2025)
FingerCap: Fine-grained Finger-level Hand Motion Captioning
par: Shen, Xin, et autres
Publié: (2025)
par: Shen, Xin, et autres
Publié: (2025)
PanORama: Multiview Consistent Panoptic Segmentation in Operating Rooms
par: Gürbüz, Tuna, et autres
Publié: (2026)
par: Gürbüz, Tuna, et autres
Publié: (2026)
PanSR: An Object-Centric Mask Transformer for Panoptic Segmentation
par: Žust, Lojze, et autres
Publié: (2024)
par: Žust, Lojze, et autres
Publié: (2024)
PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
par: Pan, Yining, et autres
Publié: (2026)
par: Pan, Yining, et autres
Publié: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
par: Oliveira, Daniel A. P., et autres
Publié: (2025)
PanSt3R: Multi-view Consistent Panoptic Segmentation
par: Zust, Lojze, et autres
Publié: (2025)
par: Zust, Lojze, et autres
Publié: (2025)
CompCap: Improving Multimodal Large Language Models with Composite Captions
par: Chen, Xiaohui, et autres
Publié: (2024)
par: Chen, Xiaohui, et autres
Publié: (2024)
SPORTS: Simultaneous Panoptic Odometry, Rendering, Tracking and Segmentation for Urban Scenes Understanding
par: Yang, Zhiliu, et autres
Publié: (2025)
par: Yang, Zhiliu, et autres
Publié: (2025)
FSAR-Cap: A Fine-Grained Two-Stage Annotated Dataset for SAR Image Captioning
par: Zhang, Jinqi, et autres
Publié: (2025)
par: Zhang, Jinqi, et autres
Publié: (2025)
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding
par: Zheng, Lihao, et autres
Publié: (2026)
par: Zheng, Lihao, et autres
Publié: (2026)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
par: Chen, Jieneng, et autres
Publié: (2024)
par: Chen, Jieneng, et autres
Publié: (2024)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
par: Li, Xiangtai, et autres
Publié: (2023)
par: Li, Xiangtai, et autres
Publié: (2023)
Panoptic Captioning: An Equivalence Bridge for Image and Text
par: Lin, Kun-Yu, et autres
Publié: (2025)
par: Lin, Kun-Yu, et autres
Publié: (2025)
Enhancing Temporal Consistency in Video Editing by Reconstructing Videos with 3D Gaussian Splatting
par: Shin, Inkyu, et autres
Publié: (2024)
par: Shin, Inkyu, et autres
Publié: (2024)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
par: Zhang, Xu, et autres
Publié: (2026)
par: Zhang, Xu, et autres
Publié: (2026)
VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
par: Wang, Ziteng, et autres
Publié: (2025)
par: Wang, Ziteng, et autres
Publié: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
par: Uijlings, Jasper, et autres
Publié: (2025)
par: Uijlings, Jasper, et autres
Publié: (2025)
ProCap: Projection-Aware Captioning for Spatial Augmented Reality
par: Cao, Zimo, et autres
Publié: (2026)
par: Cao, Zimo, et autres
Publié: (2026)
MC-PanDA: Mask Confidence for Panoptic Domain Adaptation
par: Martinović, Ivan, et autres
Publié: (2024)
par: Martinović, Ivan, et autres
Publié: (2024)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
par: Ren, Sucheng, et autres
Publié: (2025)
par: Ren, Sucheng, et autres
Publié: (2025)
Frequency-Aware Flow Matching for High-Quality Image Generation
par: Ren, Sucheng, et autres
Publié: (2026)
par: Ren, Sucheng, et autres
Publié: (2026)
Alleviating Distortion in Image Generation via Multi-Resolution Diffusion Models and Time-Dependent Layer Normalization
par: Liu, Qihao, et autres
Publié: (2024)
par: Liu, Qihao, et autres
Publié: (2024)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
par: Ren, Sucheng, et autres
Publié: (2024)
par: Ren, Sucheng, et autres
Publié: (2024)
Deeply Supervised Flow-Based Generative Models
par: Shin, Inkyu, et autres
Publié: (2025)
par: Shin, Inkyu, et autres
Publié: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
par: Yang, Chenglin, et autres
Publié: (2023)
par: Yang, Chenglin, et autres
Publié: (2023)
Video ReCap: Recursive Captioning of Hour-Long Videos
par: Islam, Md Mohaiminul, et autres
Publié: (2024)
par: Islam, Md Mohaiminul, et autres
Publié: (2024)
Rational Design Strategies in DNA‐Encoded Libraries for Drug Discovery
par: Xudong Wang, et autres
Publié: (2025)
par: Xudong Wang, et autres
Publié: (2025)
COCO-OLAC: A Benchmark for Occluded Panoptic Segmentation and Image Understanding
par: Wei, Wenbo, et autres
Publié: (2024)
par: Wei, Wenbo, et autres
Publié: (2024)
ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
par: Cao, Shuo, et autres
Publié: (2025)
par: Cao, Shuo, et autres
Publié: (2025)
CaReBench: A Fine-Grained Benchmark for Video Captioning and Retrieval
par: Xu, Yifan, et autres
Publié: (2024)
par: Xu, Yifan, et autres
Publié: (2024)
Open-World Panoptic Segmentation
par: Sodano, Matteo, et autres
Publié: (2024)
par: Sodano, Matteo, et autres
Publié: (2024)
Documents similaires
-
COCONut: Modernizing COCO Segmentation
par: Deng, Xueqing, et autres
Publié: (2024) -
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
par: Athar, Ali, et autres
Publié: (2024) -
Randomized Autoregressive Visual Generation
par: Yu, Qihang, et autres
Publié: (2024) -
A Simple Video Segmenter by Tracking Objects Along Axial Trajectories
par: He, Ju, et autres
Publié: (2023) -
PanDepth: Joint Panoptic Segmentation and Depth Completion
par: Lagos, Juan, et autres
Publié: (2022)