Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Cheng, Tianheng, Wang, Xinggang, Liao, Junchao, Liu, Wenyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
por: Li, Yongkang, et al.
Publicado: (2024)
por: Li, Yongkang, et al.
Publicado: (2024)
Occupancy as Set of Points
por: Shi, Yiang, et al.
Publicado: (2024)
por: Shi, Yiang, et al.
Publicado: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
por: Cheng, Tianheng, et al.
Publicado: (2024)
por: Cheng, Tianheng, et al.
Publicado: (2024)
Polar Parametrization for Vision-based Surround-View 3D Detection
por: Chen, Shaoyu, et al.
Publicado: (2022)
por: Chen, Shaoyu, et al.
Publicado: (2022)
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
por: Liao, Bencheng, et al.
Publicado: (2025)
por: Liao, Bencheng, et al.
Publicado: (2025)
AFRDA: Attentive Feature Refinement for Domain Adaptive Semantic Segmentation
por: Khan, Md. Al-Masrur, et al.
Publicado: (2025)
por: Khan, Md. Al-Masrur, et al.
Publicado: (2025)
EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
por: Zhang, Yuxuan, et al.
Publicado: (2024)
por: Zhang, Yuxuan, et al.
Publicado: (2024)
Lane Graph as Path: Continuity-preserving Path-wise Modeling for Online Lane Graph Construction
por: Liao, Bencheng, et al.
Publicado: (2023)
por: Liao, Bencheng, et al.
Publicado: (2023)
GaussTR: Foundation Model-Aligned Gaussian Transformer for Self-Supervised 3D Spatial Understanding
por: Jiang, Haoyi, et al.
Publicado: (2024)
por: Jiang, Haoyi, et al.
Publicado: (2024)
Matte Anything: Interactive Natural Image Matting with Segment Anything Models
por: Yao, Jingfeng, et al.
Publicado: (2023)
por: Yao, Jingfeng, et al.
Publicado: (2023)
Causality-inspired Discriminative Feature Learning in Triple Domains for Gait Recognition
por: Xiong, Haijun, et al.
Publicado: (2024)
por: Xiong, Haijun, et al.
Publicado: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
por: Zhu, Lianghui, et al.
Publicado: (2023)
por: Zhu, Lianghui, et al.
Publicado: (2023)
MaTVLM: Hybrid Mamba-Transformer for Efficient Vision-Language Modeling
por: Li, Yingyue, et al.
Publicado: (2025)
por: Li, Yingyue, et al.
Publicado: (2025)
DyGLNet: Hybrid Global-Local Feature Fusion with Dynamic Upsampling for Medical Image Segmentation
por: Zhao, Yican, et al.
Publicado: (2025)
por: Zhao, Yican, et al.
Publicado: (2025)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
por: Yao, Jingfeng, et al.
Publicado: (2024)
por: Yao, Jingfeng, et al.
Publicado: (2024)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
por: Hu, Bin, et al.
Publicado: (2024)
por: Hu, Bin, et al.
Publicado: (2024)
GaitGS: Temporal Feature Learning in Granularity and Span Dimension for Gait Recognition
por: Xiong, Haijun, et al.
Publicado: (2023)
por: Xiong, Haijun, et al.
Publicado: (2023)
ControlAR: Controllable Image Generation with Autoregressive Models
por: Li, Zongming, et al.
Publicado: (2024)
por: Li, Zongming, et al.
Publicado: (2024)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
por: Zou, Jialv, et al.
Publicado: (2025)
por: Zou, Jialv, et al.
Publicado: (2025)
MIM4D: Masked Modeling with Multi-View Video for Autonomous Driving Representation Learning
por: Zou, Jialv, et al.
Publicado: (2024)
por: Zou, Jialv, et al.
Publicado: (2024)
LENS: Learning to Segment Anything with Unified Reinforced Reasoning
por: Zhu, Lianghui, et al.
Publicado: (2025)
por: Zhu, Lianghui, et al.
Publicado: (2025)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
por: Hu, Rui, et al.
Publicado: (2025)
por: Hu, Rui, et al.
Publicado: (2025)
ViTGaze: Gaze Following with Interaction Features in Vision Transformers
por: Song, Yuehao, et al.
Publicado: (2024)
por: Song, Yuehao, et al.
Publicado: (2024)
Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
por: Seo, Minseok, et al.
Publicado: (2025)
por: Seo, Minseok, et al.
Publicado: (2025)
MolSight: Optical Chemical Structure Recognition with SMILES Pretraining, Multi-Granularity Learning and Reinforcement Learning
por: Zhang, Wenrui, et al.
Publicado: (2025)
por: Zhang, Wenrui, et al.
Publicado: (2025)
GaraMoSt: Parallel Multi-Granularity Motion and Structural Modeling for Efficient Multi-Frame Interpolation in DSA Images
por: Xu, Ziyang, et al.
Publicado: (2024)
por: Xu, Ziyang, et al.
Publicado: (2024)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
por: Zeng, Lunbin, et al.
Publicado: (2025)
por: Zeng, Lunbin, et al.
Publicado: (2025)
Weighted Reverse Convolution for Feature Upsampling
por: Li, Wentong, et al.
Publicado: (2026)
por: Li, Wentong, et al.
Publicado: (2026)
Segment Any 4D Gaussians
por: Ji, Shengxiang, et al.
Publicado: (2024)
por: Ji, Shengxiang, et al.
Publicado: (2024)
WeakSAM: Segment Anything Meets Weakly-supervised Instance-level Recognition
por: Zhu, Lianghui, et al.
Publicado: (2024)
por: Zhu, Lianghui, et al.
Publicado: (2024)
PixelHacker: Image Inpainting with Structural and Semantic Consistency
por: Xu, Ziyang, et al.
Publicado: (2025)
por: Xu, Ziyang, et al.
Publicado: (2025)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
por: Xiong, Haijun, et al.
Publicado: (2025)
por: Xiong, Haijun, et al.
Publicado: (2025)
Graph-Boosted Attentive Network for Semantic Body Parsing
por: Wang, Tinghuai, et al.
Publicado: (2024)
por: Wang, Tinghuai, et al.
Publicado: (2024)
DeltaMIL: Gated Memory Integration for Efficient and Discriminative Whole Slide Image Analysis
por: Zhu, Yueting, et al.
Publicado: (2025)
por: Zhu, Yueting, et al.
Publicado: (2025)
STP4D: Spatio-Temporal-Prompt Consistent Modeling for Text-to-4D Gaussian Splatting
por: Deng, Yunze, et al.
Publicado: (2025)
por: Deng, Yunze, et al.
Publicado: (2025)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
por: Zhu, Lianghui, et al.
Publicado: (2024)
por: Zhu, Lianghui, et al.
Publicado: (2024)
DiveUp: Learning Feature Upsampling from Diverse Vision Foundation Models
por: Liu, Xiaoqiong, et al.
Publicado: (2026)
por: Liu, Xiaoqiong, et al.
Publicado: (2026)
Boosting Few-Shot Learning via Attentive Feature Regularization
por: Zhu, Xingyu, et al.
Publicado: (2024)
por: Zhu, Xingyu, et al.
Publicado: (2024)
LKCell: Efficient Cell Nuclei Instance Segmentation with Large Convolution Kernels
por: Cui, Ziwei, et al.
Publicado: (2024)
por: Cui, Ziwei, et al.
Publicado: (2024)
SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised Segmentation
por: Bi, Xiuli, et al.
Publicado: (2025)
por: Bi, Xiuli, et al.
Publicado: (2025)
Ejemplares similares
-
Mask-Adapter: The Devil is in the Masks for Open-Vocabulary Segmentation
por: Li, Yongkang, et al.
Publicado: (2024) -
Occupancy as Set of Points
por: Shi, Yiang, et al.
Publicado: (2024) -
YOLO-World: Real-Time Open-Vocabulary Object Detection
por: Cheng, Tianheng, et al.
Publicado: (2024) -
Polar Parametrization for Vision-based Surround-View 3D Detection
por: Chen, Shaoyu, et al.
Publicado: (2022) -
Multimodal Mamba: Decoder-only Multimodal State Space Model via Quadratic to Linear Distillation
por: Liao, Bencheng, et al.
Publicado: (2025)