Fully Attentional Networks with Self-emerging Token Labeling
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Bingyin, Yu, Zhiding, Lan, Shiyi, Cheng, Yutao, Anandkumar, Anima, Lao, Yingjie, Alvarez, Jose M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What is Point Supervision Worth in Video Instance Segmentation?
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
Improving Distant 3D Object Detection Using 2D Box Supervision
by: Yang, Zetong, et al.
Published: (2024)
by: Yang, Zetong, et al.
Published: (2024)
Learning Calibrated Uncertainties for Domain Shift: A Distributionally Robust Learning Approach
by: Wang, Haoxuan, et al.
Published: (2020)
by: Wang, Haoxuan, et al.
Published: (2020)
Prismer: A Vision-Language Model with Multi-Task Experts
by: Liu, Shikun, et al.
Published: (2023)
by: Liu, Shikun, et al.
Published: (2023)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
by: Li, Zhenxin, et al.
Published: (2025)
by: Li, Zhenxin, et al.
Published: (2025)
Exploring Camera Encoder Designs for Autonomous Driving Perception
by: Lakshmanan, Barath, et al.
Published: (2024)
by: Lakshmanan, Barath, et al.
Published: (2024)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
by: Li, Zhiqi, et al.
Published: (2023)
by: Li, Zhiqi, et al.
Published: (2023)
Enhancing Autonomous Driving Safety with Collision Scenario Integration
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
T-Stitch: Accelerating Sampling in Pre-Trained Diffusion Models with Trajectory Stitching
by: Pan, Zizheng, et al.
Published: (2024)
by: Pan, Zizheng, et al.
Published: (2024)
Diffusion State-Guided Projected Gradient for Inverse Problems
by: Zirvi, Rayhan, et al.
Published: (2024)
by: Zirvi, Rayhan, et al.
Published: (2024)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2025)
by: Wang, Shihao, et al.
Published: (2025)
Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
by: Li, Kailin, et al.
Published: (2025)
by: Li, Kailin, et al.
Published: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
by: Wang, Shihao, et al.
Published: (2024)
by: Wang, Shihao, et al.
Published: (2024)
StreamChat: Chatting with Streaming Video
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
BEVNeXt: Reviving Dense BEV Frameworks for 3D Object Detection
by: Li, Zhenxin, et al.
Published: (2023)
by: Li, Zhenxin, et al.
Published: (2023)
Fast Training of Diffusion Models with Masked Transformers
by: Zheng, Hongkai, et al.
Published: (2023)
by: Zheng, Hongkai, et al.
Published: (2023)
Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
by: Tang, Yutao, et al.
Published: (2025)
by: Tang, Yutao, et al.
Published: (2025)
Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Centaur: Robust End-to-End Autonomous Driving with Test-Time Training
by: Sima, Chonghao, et al.
Published: (2025)
by: Sima, Chonghao, et al.
Published: (2025)
Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
by: Li, Zhenxin, et al.
Published: (2024)
by: Li, Zhenxin, et al.
Published: (2024)
SEGIC: Unleashing the Emergent Correspondence for In-Context Segmentation
by: Meng, Lingchen, et al.
Published: (2023)
by: Meng, Lingchen, et al.
Published: (2023)
Resolution-Independent Neural Operators for Multi-Rate Sparse-View CT
by: Datta, Aujasvit, et al.
Published: (2025)
by: Datta, Aujasvit, et al.
Published: (2025)
DMin: Scalable Training Data Influence Estimation for Diffusion Models
by: Lin, Huawei, et al.
Published: (2024)
by: Lin, Huawei, et al.
Published: (2024)
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
by: Qiao, Bingtian, et al.
Published: (2026)
by: Qiao, Bingtian, et al.
Published: (2026)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
MixFormerV2: Efficient Fully Transformer Tracking
by: Cui, Yutao, et al.
Published: (2023)
by: Cui, Yutao, et al.
Published: (2023)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
Attention Debiasing for Token Pruning in Vision Language Models
by: Zhao, Kai, et al.
Published: (2025)
by: Zhao, Kai, et al.
Published: (2025)
Progressively Normalized Self-Attention Network for Video Polyp Segmentation
by: Ji, Ge-Peng, et al.
Published: (2021)
by: Ji, Ge-Peng, et al.
Published: (2021)
UltraClean: A Simple Framework to Train Robust Neural Networks against Backdoor Attacks
by: Zhao, Bingyin, et al.
Published: (2023)
by: Zhao, Bingyin, et al.
Published: (2023)
FashionMV: Product-Level Composed Image Retrieval with Multi-View Fashion Data
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
Improving Diffusion Inverse Problem Solving with Decoupled Noise Annealing
by: Zhang, Bingliang, et al.
Published: (2024)
by: Zhang, Bingliang, et al.
Published: (2024)
Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms
by: Li, Fengyun, et al.
Published: (2025)
by: Li, Fengyun, et al.
Published: (2025)
ARDuP: Active Region Video Diffusion for Universal Policies
by: Huang, Shuaiyi, et al.
Published: (2024)
by: Huang, Shuaiyi, et al.
Published: (2024)
3DEditSafe: Defending 3D Editing Pipelines from Unsafe Generation
by: Meng, Nicole, et al.
Published: (2026)
by: Meng, Nicole, et al.
Published: (2026)
Cross-Layer Feature Self-Attention Module for Multi-Scale Object Detection
by: Xie, Dingzhou, et al.
Published: (2025)
by: Xie, Dingzhou, et al.
Published: (2025)
Reconstruction-Based Anomaly Localization via Knowledge-Informed Self-Training
by: Qian, Cheng, et al.
Published: (2024)
by: Qian, Cheng, et al.
Published: (2024)
Graph Attention Transformer Network for Multi-Label Image Classification
by: Yuan, Jin, et al.
Published: (2022)
by: Yuan, Jin, et al.
Published: (2022)
Fully Automatic Data Labeling for Ultrasound Screen Detection
by: Gomez, Alberto, et al.
Published: (2025)
by: Gomez, Alberto, et al.
Published: (2025)
Attention-Based 3D Seismic Fault Segmentation Training by a Few 2D Slice Labels
by: Dou, YiMin, et al.
Published: (2021)
by: Dou, YiMin, et al.
Published: (2021)
Similar Items
-
What is Point Supervision Worth in Video Instance Segmentation?
by: Huang, Shuaiyi, et al.
Published: (2024) -
Improving Distant 3D Object Detection Using 2D Box Supervision
by: Yang, Zetong, et al.
Published: (2024) -
Learning Calibrated Uncertainties for Domain Shift: A Distributionally Robust Learning Approach
by: Wang, Haoxuan, et al.
Published: (2020) -
Prismer: A Vision-Language Model with Multi-Task Experts
by: Liu, Shikun, et al.
Published: (2023) -
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
by: Li, Zhenxin, et al.
Published: (2025)