Mixed-Query Transformer: A Unified Image Segmentation Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Pei, Cai, Zhaowei, Yang, Hao, Swaminathan, Ashwin, Manmatha, R., Soatto, Stefano |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
by: Tao, Chaofan, et al.
Published: (2024)
by: Tao, Chaofan, et al.
Published: (2024)
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023)
by: Golatkar, Aditya, et al.
Published: (2023)
Fast Sparse View Guided NeRF Update for Object Reconfigurations
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
A Quantitative Evaluation of Score Distillation Sampling Based Text-to-3D
by: Fei, Xiaohan, et al.
Published: (2024)
by: Fei, Xiaohan, et al.
Published: (2024)
THRONE: An Object-based Hallucination Benchmark for the Free-form Generations of Large Vision-Language Models
by: Kaul, Prannay, et al.
Published: (2024)
by: Kaul, Prannay, et al.
Published: (2024)
CPR: Retrieval Augmented Generation for Copyright Protection
by: Golatkar, Aditya, et al.
Published: (2024)
by: Golatkar, Aditya, et al.
Published: (2024)
Rethinking Query-based Transformer for Continual Image Segmentation
by: Zhu, Yuchen, et al.
Published: (2025)
by: Zhu, Yuchen, et al.
Published: (2025)
Multi-Modal Hallucination Control by Visual Information Grounding
by: Favero, Alessandro, et al.
Published: (2024)
by: Favero, Alessandro, et al.
Published: (2024)
Diffusion Soup: Model Merging for Text-to-Image Diffusion Models
by: Biggs, Benjamin, et al.
Published: (2024)
by: Biggs, Benjamin, et al.
Published: (2024)
On the Viability of Monocular Depth Pre-training for Semantic Segmentation
by: Lao, Dong, et al.
Published: (2022)
by: Lao, Dong, et al.
Published: (2022)
DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
by: Li, Xiao-Hui, et al.
Published: (2025)
by: Li, Xiao-Hui, et al.
Published: (2025)
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Grounded Compositional and Diverse Text-to-3D with Pretrained Multi-View Diffusion Model
by: Li, Xiaolong, et al.
Published: (2024)
by: Li, Xiaolong, et al.
Published: (2024)
Diffeomorphic Template Registration for Atmospheric Turbulence Mitigation
by: Lao, Dong, et al.
Published: (2024)
by: Lao, Dong, et al.
Published: (2024)
Visual Reasoning through Tool-supervised Reinforcement Learning
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Musketeer: Joint Training for Multi-task Vision Language Model with Task Explanation Prompts
by: Zhang, Zhaoyang, et al.
Published: (2023)
by: Zhang, Zhaoyang, et al.
Published: (2023)
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
by: Tan, Jing, et al.
Published: (2026)
by: Tan, Jing, et al.
Published: (2026)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
by: Song, Yuxin, et al.
Published: (2025)
by: Song, Yuxin, et al.
Published: (2025)
DQFormer: Towards Unified LiDAR Panoptic Segmentation with Decoupled Queries
by: Yang, Yu, et al.
Published: (2024)
by: Yang, Yu, et al.
Published: (2024)
Adaptive Mix for Semi-Supervised Medical Image Segmentation
by: Shen, Zhiqiang, et al.
Published: (2024)
by: Shen, Zhiqiang, et al.
Published: (2024)
UTSRMorph: A Unified Transformer and Superresolution Network for Unsupervised Medical Image Registration
by: Zhang, Runshi, et al.
Published: (2024)
by: Zhang, Runshi, et al.
Published: (2024)
BEFUnet: A Hybrid CNN-Transformer Architecture for Precise Medical Image Segmentation
by: Manzari, Omid Nejati, et al.
Published: (2024)
by: Manzari, Omid Nejati, et al.
Published: (2024)
When Swin Transformer Meets KANs: An Improved Transformer Architecture for Medical Image Segmentation
by: Sapkota, Nishchal, et al.
Published: (2025)
by: Sapkota, Nishchal, et al.
Published: (2025)
Gasformer: A Transformer-based Architecture for Segmenting Methane Emissions from Livestock in Optical Gas Imaging
by: Sarker, Toqi Tahamid, et al.
Published: (2024)
by: Sarker, Toqi Tahamid, et al.
Published: (2024)
SEG-SAM: Semantic-Guided SAM for Unified Medical Image Segmentation
by: Huang, Shuangping, et al.
Published: (2024)
by: Huang, Shuangping, et al.
Published: (2024)
Sub-token ViT Embedding via Stochastic Resonance Transformers
by: Lao, Dong, et al.
Published: (2023)
by: Lao, Dong, et al.
Published: (2023)
Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
by: Swaminathan, Arjhun, et al.
Published: (2025)
by: Swaminathan, Arjhun, et al.
Published: (2025)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Textual Query-Driven Mask Transformer for Domain Generalized Segmentation
by: Pak, Byeonghyun, et al.
Published: (2024)
by: Pak, Byeonghyun, et al.
Published: (2024)
USegMix: Unsupervised Segment Mix for Efficient Data Augmentation in Pathology Images
by: Wang, Jiamu, et al.
Published: (2025)
by: Wang, Jiamu, et al.
Published: (2025)
MM-UNet: A Mixed MLP Architecture for Improved Ophthalmic Image Segmentation
by: Xiao, Zunjie, et al.
Published: (2024)
by: Xiao, Zunjie, et al.
Published: (2024)
A Novel Unified Architecture for Low-Shot Counting by Detection and Segmentation
by: Pelhan, Jer, et al.
Published: (2024)
by: Pelhan, Jer, et al.
Published: (2024)
SegStitch: Multidimensional Transformer for Robust and Efficient Medical Imaging Segmentation
by: Tan, Shengbo, et al.
Published: (2024)
by: Tan, Shengbo, et al.
Published: (2024)
PanopticQuery: Unified Query-Time Reasoning for 4D Scenes
by: Tang, Ruilin, et al.
Published: (2026)
by: Tang, Ruilin, et al.
Published: (2026)
HyCTAS: Multi-Objective Hybrid Convolution-Transformer Architecture Search for Real-Time Image Segmentation
by: Yu, Hongyuan, et al.
Published: (2024)
by: Yu, Hongyuan, et al.
Published: (2024)
QMambaBSR: Burst Image Super-Resolution with Query State Space Model
by: Di, Xin, et al.
Published: (2024)
by: Di, Xin, et al.
Published: (2024)
Unified Restoration-Perception Learning: Maritime Infrared-Visible Image Fusion and Segmentation
by: Cai, Weichao, et al.
Published: (2026)
by: Cai, Weichao, et al.
Published: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
by: Cai, Qi, et al.
Published: (2026)
by: Cai, Qi, et al.
Published: (2026)
Similar Items
-
NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality
by: Tao, Chaofan, et al.
Published: (2024) -
Efficient Scaling of Diffusion Transformers for Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024) -
On the Scalability of Diffusion-based Text-to-Image Generation
by: Li, Hao, et al.
Published: (2024) -
Training Data Protection with Compositional Diffusion Models
by: Golatkar, Aditya, et al.
Published: (2023) -
Fast Sparse View Guided NeRF Update for Object Reconfigurations
by: Lu, Ziqi, et al.
Published: (2024)