Vision Transformers: From Semantic Segmentation to Dense Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Li, Lu, Jiachen, Zheng, Sixiao, Zhao, Xinxuan, Zhu, Xiatian, Fu, Yanwei, Xiang, Tao, Feng, Jianfeng, Torr, Philip H. S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Softmax-free Linear Transformers
von: Lu, Jiachen, et al.
Veröffentlicht: (2022)
von: Lu, Jiachen, et al.
Veröffentlicht: (2022)
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
4D Gaussian Splatting: Modeling Dynamic Scenes with Native 4D Primitives
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
Influencer Backdoor Attack on Semantic Segmentation
von: Lan, Haoheng, et al.
Veröffentlicht: (2023)
von: Lan, Haoheng, et al.
Veröffentlicht: (2023)
Mind-of-Director: Multi-modal Agent-Driven Film Previsualization via Collaborative Decision-Making
von: Nan, Shufeng, et al.
Veröffentlicht: (2026)
von: Nan, Shufeng, et al.
Veröffentlicht: (2026)
Behind Every Domain There is a Shift: Adapting Distortion-aware Vision Transformers for Panoramic Semantic Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2022)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2022)
Enhancing High-Resolution 3D Generation through Pixel-wise Gradient Clipping
von: Pan, Zijie, et al.
Veröffentlicht: (2023)
von: Pan, Zijie, et al.
Veröffentlicht: (2023)
Unified Domain Adaptive Semantic Segmentation
von: Zhang, Zhe, et al.
Veröffentlicht: (2023)
von: Zhang, Zhe, et al.
Veröffentlicht: (2023)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
Preconditioned Score-based Generative Models
von: Ma, Hengyuan, et al.
Veröffentlicht: (2023)
von: Ma, Hengyuan, et al.
Veröffentlicht: (2023)
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
Online Dense Point Tracking with Streaming Memory
von: Dong, Qiaole, et al.
Veröffentlicht: (2025)
von: Dong, Qiaole, et al.
Veröffentlicht: (2025)
DeepInteraction++: Multi-Modality Interaction for Autonomous Driving
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
MacFormer: Semantic Segmentation with Fine Object Boundaries
von: Xu, Guoan, et al.
Veröffentlicht: (2024)
von: Xu, Guoan, et al.
Veröffentlicht: (2024)
RD-ViT: Recurrent-Depth Vision Transformer for Semantic Segmentation with Reduced Data Dependence Extending the Recurrent-Depth Transformer Architecture to Dense Prediction
von: He, Renjie
Veröffentlicht: (2026)
von: He, Renjie
Veröffentlicht: (2026)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding and Reasoning
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
The Pictorial Cortex: Zero-Shot Cross-Subject fMRI-to-Image Reconstruction via Compositional Latent Modeling
von: Huo, Jingyang, et al.
Veröffentlicht: (2026)
von: Huo, Jingyang, et al.
Veröffentlicht: (2026)
Making Your Dreams A Reality: Decoding the Dreams into a Coherent Video Story from fMRI Signals
von: Fu, Yanwei, et al.
Veröffentlicht: (2025)
von: Fu, Yanwei, et al.
Veröffentlicht: (2025)
KANs for Computer Vision: An Experimental Study
von: Mohan, Karthik, et al.
Veröffentlicht: (2024)
von: Mohan, Karthik, et al.
Veröffentlicht: (2024)
Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
von: Zhou, Jinxin, et al.
Veröffentlicht: (2025)
von: Zhou, Jinxin, et al.
Veröffentlicht: (2025)
Revitalizing Dense Material Segmentation: Stabilized Vision Transformers and the Generalization Paradox
von: Kazakov, Allan, et al.
Veröffentlicht: (2026)
von: Kazakov, Allan, et al.
Veröffentlicht: (2026)
Exploring Token-Level Augmentation in Vision Transformer for Semi-Supervised Semantic Segmentation
von: Zhang, Dengke, et al.
Veröffentlicht: (2025)
von: Zhang, Dengke, et al.
Veröffentlicht: (2025)
3D Audio-Visual Segmentation
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
von: Sokolov, Artem, et al.
Veröffentlicht: (2024)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
von: Xia, Chunlong, et al.
Veröffentlicht: (2024)
von: Xia, Chunlong, et al.
Veröffentlicht: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
von: Hong, Yuanduo, et al.
Veröffentlicht: (2022)
Polyphony: Diffusion-based Dual-Hand Action Segmentation with Alternating Vision Transformer and Semantic Conditioning
von: Zheng, Hao, et al.
Veröffentlicht: (2026)
von: Zheng, Hao, et al.
Veröffentlicht: (2026)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
von: Rezaei, Razieh, et al.
Veröffentlicht: (2024)
CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2025)
Task Indicating Transformer for Task-conditional Dense Predictions
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
von: Lu, Yuxiang, et al.
Veröffentlicht: (2024)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
ConceptHash: Interpretable Fine-Grained Hashing via Concept Discovery
von: Ng, Kam Woh, et al.
Veröffentlicht: (2024)
von: Ng, Kam Woh, et al.
Veröffentlicht: (2024)
PartCraft: Crafting Creative Objects by Parts
von: Ng, Kam Woh, et al.
Veröffentlicht: (2024)
von: Ng, Kam Woh, et al.
Veröffentlicht: (2024)
MinD-3D: Reconstruct High-quality 3D objects in Human Brain
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2023)
MinD-3D++: Advancing fMRI-Based 3D Reconstruction with High-Quality Textured Mesh Generation and a Comprehensive Dataset
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
von: Gao, Jianxiong, et al.
Veröffentlicht: (2024)
MERGETUNE: Continued Fine-Tuning of Vision-Language Models
von: Wang, Wenqing, et al.
Veröffentlicht: (2026)
von: Wang, Wenqing, et al.
Veröffentlicht: (2026)
MemFlow: Optical Flow Estimation and Prediction with Memory
von: Dong, Qiaole, et al.
Veröffentlicht: (2024)
von: Dong, Qiaole, et al.
Veröffentlicht: (2024)
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
von: Lu, Xinxuan, et al.
Veröffentlicht: (2026)
von: Lu, Xinxuan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Softmax-free Linear Transformers
von: Lu, Jiachen, et al.
Veröffentlicht: (2022) -
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024) -
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024) -
4D Gaussian Splatting: Modeling Dynamic Scenes with Native 4D Primitives
von: Yang, Zeyu, et al.
Veröffentlicht: (2024) -
Influencer Backdoor Attack on Semantic Segmentation
von: Lan, Haoheng, et al.
Veröffentlicht: (2023)