VST++: Efficient and Stronger Visual Saliency Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Nian, Luo, Ziyang, Zhang, Ni, Han, Junwei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
by: Li, Long, et al.
Published: (2025)
by: Li, Long, et al.
Published: (2025)
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs
by: Luo, Ziyang, et al.
Published: (2026)
by: Luo, Ziyang, et al.
Published: (2026)
AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning
by: Luo, Ziyang, et al.
Published: (2023)
by: Luo, Ziyang, et al.
Published: (2023)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
TopoVST: Toward Topology-fidelitous Vessel Skeleton Tracking
by: Liu, Yaoyu, et al.
Published: (2026)
by: Liu, Yaoyu, et al.
Published: (2026)
Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering Refinement
by: Hou, Xiuquan, et al.
Published: (2024)
by: Hou, Xiuquan, et al.
Published: (2024)
Point Transformer V3: Simpler, Faster, Stronger
by: Wu, Xiaoyang, et al.
Published: (2023)
by: Wu, Xiaoyang, et al.
Published: (2023)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions
by: Su, Junhao, et al.
Published: (2025)
by: Su, Junhao, et al.
Published: (2025)
UniVST: A Unified Framework for Training-free Localized Video Style Transfer
by: Song, Quanjian, et al.
Published: (2024)
by: Song, Quanjian, et al.
Published: (2024)
VST-Pose: A Velocity-Integrated Spatiotem-poral Attention Network for Human WiFi Pose Estimation
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
LitePT: Lighter Yet Stronger Point Transformer
by: Yue, Yuanwen, et al.
Published: (2025)
by: Yue, Yuanwen, et al.
Published: (2025)
Faster or Stronger: Towards Flexible Visual Place Recognition via Weighted Aggregation and Token Pruning
by: Zeng, Zichao, et al.
Published: (2026)
by: Zeng, Zichao, et al.
Published: (2026)
Robust Saliency-Aware Distillation for Few-shot Fine-grained Visual Recognition
by: Liu, Haiqi, et al.
Published: (2023)
by: Liu, Haiqi, et al.
Published: (2023)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
by: Parikh, Dhruv, et al.
Published: (2026)
by: Parikh, Dhruv, et al.
Published: (2026)
PointOBB-v2: Towards Simpler, Faster, and Stronger Single Point Supervised Oriented Object Detection
by: Ren, Botao, et al.
Published: (2024)
by: Ren, Botao, et al.
Published: (2024)
Finding Visual Saliency in Continuous Spike Stream
by: Zhu, Lin, et al.
Published: (2024)
by: Zhu, Lin, et al.
Published: (2024)
Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
by: Zhang, Yifei, et al.
Published: (2023)
by: Zhang, Yifei, et al.
Published: (2023)
Saliency Suppressed, Semantics Surfaced: Visual Transformations in Neural Networks and the Brain
by: Opiełka, Gustaw, et al.
Published: (2024)
by: Opiełka, Gustaw, et al.
Published: (2024)
Contextual Encoder-Decoder Network for Visual Saliency Prediction
by: Kroner, Alexander, et al.
Published: (2019)
by: Kroner, Alexander, et al.
Published: (2019)
Stronger, Steadier & Superior: Geometric Consistency in Depth VFM Forges Domain Generalized Semantic Segmentation
by: Chen, Siyu, et al.
Published: (2025)
by: Chen, Siyu, et al.
Published: (2025)
Stronger Normalization-Free Transformers
by: Chen, Mingzhi, et al.
Published: (2025)
by: Chen, Mingzhi, et al.
Published: (2025)
Unsupervised Pre-training with Language-Vision Prompts for Low-Data Instance Segmentation
by: Zhang, Dingwen, et al.
Published: (2024)
by: Zhang, Dingwen, et al.
Published: (2024)
SGDM: Static-Guided Dynamic Module Make Stronger Visual Models
by: Xing, Wenjie, et al.
Published: (2024)
by: Xing, Wenjie, et al.
Published: (2024)
Exploring Stronger Transformer Representation Learning for Occluded Person Re-Identification
by: Ji, Zhangjian, et al.
Published: (2024)
by: Ji, Zhangjian, et al.
Published: (2024)
Saliency Guided Longitudinal Medical Visual Question Answering
by: Wu, Jialin, et al.
Published: (2025)
by: Wu, Jialin, et al.
Published: (2025)
Learning with Unmasked Tokens Drives Stronger Vision Learners
by: Kim, Taekyung, et al.
Published: (2023)
by: Kim, Taekyung, et al.
Published: (2023)
Increasing Interpretability of Neural Networks By Approximating Human Visual Saliency
by: Boyd, Aidan, et al.
Published: (2024)
by: Boyd, Aidan, et al.
Published: (2024)
Relevance-guided Audio Visual Fusion for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2024)
by: Yu, Li, et al.
Published: (2024)
SleepVST: Sleep Staging from Near-Infrared Video Signals using Pre-Trained Transformers
by: Carter, Jonathan F., et al.
Published: (2024)
by: Carter, Jonathan F., et al.
Published: (2024)
Pixel Distillation: A New Knowledge Distillation Scheme for Low-Resolution Image Recognition
by: Guo, Guangyu, et al.
Published: (2021)
by: Guo, Guangyu, et al.
Published: (2021)
DyDiT++: Diffusion Transformers with Timestep and Spatial Dynamics for Efficient Visual Generation
by: Zhao, Wangbo, et al.
Published: (2025)
by: Zhao, Wangbo, et al.
Published: (2025)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
DTFSal: Audio-Visual Dynamic Token Fusion for Video Saliency Prediction
by: Hooshanfar, Kiana, et al.
Published: (2025)
by: Hooshanfar, Kiana, et al.
Published: (2025)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
Efficient Multi-Crop Saliency Partitioning for Automatic Image Cropping
by: Hamara, Andrew, et al.
Published: (2025)
by: Hamara, Andrew, et al.
Published: (2025)
Saliency Driven Imagery Preprocessing for Efficient Compression -- Industrial Paper
by: Downes, Justin, et al.
Published: (2026)
by: Downes, Justin, et al.
Published: (2026)
Stronger Semantic Encoders Can Harm Relighting Performance: Probing Visual Priors via Augmented Latent Intrinsics
by: Xing, Xiaoyan, et al.
Published: (2026)
by: Xing, Xiaoyan, et al.
Published: (2026)
Depth-induced Saliency Comparison Network for Diagnosis of Alzheimer's Disease via Jointly Analysis of Visual Stimuli and Eye Movements
by: Liu, Yu, et al.
Published: (2024)
by: Liu, Yu, et al.
Published: (2024)
Similar Items
-
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
by: Li, Long, et al.
Published: (2025) -
Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs
by: Luo, Ziyang, et al.
Published: (2026) -
AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
by: Luo, Ziyang, et al.
Published: (2025) -
VSCode: General Visual Salient and Camouflaged Object Detection with 2D Prompt Learning
by: Luo, Ziyang, et al.
Published: (2023) -
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)