SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Duy-Kien, Oswald, Martin R., Snoek, Cees G. M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FVO: Fast Visual Odometry with Transformers
di: Yugay, Vlardimir, et al.
Pubblicazione: (2025)
di: Yugay, Vlardimir, et al.
Pubblicazione: (2025)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
di: Bhowmik, Aritra, et al.
Pubblicazione: (2023)
di: Bhowmik, Aritra, et al.
Pubblicazione: (2023)
R-MAE: Regions Meet Masked Autoencoders
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2023)
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2023)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
di: Dorkenwald, Michael, et al.
Pubblicazione: (2024)
di: Dorkenwald, Michael, et al.
Pubblicazione: (2024)
Segment Any 3D-Part in a Scene from a Sentence
di: Wu, Hongyu, et al.
Pubblicazione: (2025)
di: Wu, Hongyu, et al.
Pubblicazione: (2025)
Dual Guidance Semi-Supervised Action Detection
di: Singh, Ankit, et al.
Pubblicazione: (2025)
di: Singh, Ankit, et al.
Pubblicazione: (2025)
TWIST & SCOUT: Grounding Multimodal LLM-Experts by Forget-Free Tuning
di: Bhowmik, Aritra, et al.
Pubblicazione: (2024)
di: Bhowmik, Aritra, et al.
Pubblicazione: (2024)
Training-Free Semantic Segmentation via LLM-Supervision
di: Sun, Wenfang, et al.
Pubblicazione: (2024)
di: Sun, Wenfang, et al.
Pubblicazione: (2024)
Low-Resource Vision Challenges for Foundation Models
di: Zhang, Yunhua, et al.
Pubblicazione: (2024)
di: Zhang, Yunhua, et al.
Pubblicazione: (2024)
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
di: Sun, Wenfang, et al.
Pubblicazione: (2024)
di: Sun, Wenfang, et al.
Pubblicazione: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
di: Liu, Huabin, et al.
Pubblicazione: (2025)
di: Liu, Huabin, et al.
Pubblicazione: (2025)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
di: Du, Yingjun, et al.
Pubblicazione: (2023)
di: Du, Yingjun, et al.
Pubblicazione: (2023)
IPO: Interpretable Prompt Optimization for Vision-Language Models
di: Du, Yingjun, et al.
Pubblicazione: (2024)
di: Du, Yingjun, et al.
Pubblicazione: (2024)
Crane: Context-Guided Prompt Learning and Attention Refinement for Zero-Shot Anomaly Detection
di: Salehi, Alireza, et al.
Pubblicazione: (2025)
di: Salehi, Alireza, et al.
Pubblicazione: (2025)
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
GeneralAD: Anomaly Detection Across Domains by Attending to Distorted Features
di: Sträter, Luc P. J., et al.
Pubblicazione: (2024)
di: Sträter, Luc P. J., et al.
Pubblicazione: (2024)
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
di: Bhowmik, Aritra, et al.
Pubblicazione: (2025)
di: Bhowmik, Aritra, et al.
Pubblicazione: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
di: Sun, Wenfang, et al.
Pubblicazione: (2026)
di: Sun, Wenfang, et al.
Pubblicazione: (2026)
Beyond Coarse-Grained Matching in Video-Text Retrieval
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
SegMaFormer: A Hybrid State-Space and Transformer Model for Efficient Segmentation
di: Nguyen, Duy D., et al.
Pubblicazione: (2026)
di: Nguyen, Duy D., et al.
Pubblicazione: (2026)
SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements
di: Xie, Haiyang, et al.
Pubblicazione: (2025)
di: Xie, Haiyang, et al.
Pubblicazione: (2025)
Elastic ViTs from Pretrained Models without Retraining
di: Simoncini, Walter, et al.
Pubblicazione: (2025)
di: Simoncini, Walter, et al.
Pubblicazione: (2025)
Lost in Time: A New Temporal Benchmark for VideoLLMs
di: Cores, Daniel, et al.
Pubblicazione: (2024)
di: Cores, Daniel, et al.
Pubblicazione: (2024)
Any-Shift Prompting for Generalization over Distributions
di: Xiao, Zehao, et al.
Pubblicazione: (2024)
di: Xiao, Zehao, et al.
Pubblicazione: (2024)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
di: Le, Minh Khoa, et al.
Pubblicazione: (2026)
SimLTD: Simple Supervised and Semi-Supervised Long-Tailed Object Detection
di: Tran, Phi Vu
Pubblicazione: (2024)
di: Tran, Phi Vu
Pubblicazione: (2024)
SimA: Simple Softmax-free Attention for Vision Transformers
di: Koohpayegani, Soroush Abbasi, et al.
Pubblicazione: (2022)
di: Koohpayegani, Soroush Abbasi, et al.
Pubblicazione: (2022)
SimToken: A Simple Baseline for Referring Audio-Visual Segmentation
di: Jin, Dian, et al.
Pubblicazione: (2025)
di: Jin, Dian, et al.
Pubblicazione: (2025)
MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2025)
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2025)
NeoBabel: A Multilingual Open Tower for Visual Generation
di: Derakhshani, Mohammad Mahdi, et al.
Pubblicazione: (2025)
di: Derakhshani, Mohammad Mahdi, et al.
Pubblicazione: (2025)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
di: Traub, Manuel, et al.
Pubblicazione: (2025)
di: Traub, Manuel, et al.
Pubblicazione: (2025)
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
di: Nguyen, Toan, et al.
Pubblicazione: (2025)
di: Nguyen, Toan, et al.
Pubblicazione: (2025)
Joint Instance Segmentation and Geometric Attribute Regression for Roof Structures in Aerial Imagery
di: Versteeg, Luuk, et al.
Pubblicazione: (2026)
di: Versteeg, Luuk, et al.
Pubblicazione: (2026)
SIGMA: Sinkhorn-Guided Masked Video Modeling
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
Auto-Vocabulary Semantic Segmentation
di: Ülger, Osman, et al.
Pubblicazione: (2023)
di: Ülger, Osman, et al.
Pubblicazione: (2023)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
di: Vu, Kiet Dang, et al.
Pubblicazione: (2026)
di: Vu, Kiet Dang, et al.
Pubblicazione: (2026)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
di: Wang, Ziqi, et al.
Pubblicazione: (2025)
di: Wang, Ziqi, et al.
Pubblicazione: (2025)
Cross-Cluster Shifting for Efficient and Effective 3D Object Detection in Autonomous Driving
di: Chen, Zhili, et al.
Pubblicazione: (2024)
di: Chen, Zhili, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FVO: Fast Visual Odometry with Transformers
di: Yugay, Vlardimir, et al.
Pubblicazione: (2025) -
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2024) -
Union-over-Intersections: Object Detection beyond Winner-Takes-All
di: Bhowmik, Aritra, et al.
Pubblicazione: (2023) -
R-MAE: Regions Meet Masked Autoencoders
di: Nguyen, Duy-Kien, et al.
Pubblicazione: (2023) -
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)