Progressive Token Length Scaling in Transformer Encoders for Efficient Universal Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Aich, Abhishek, Suh, Yumin, Schulter, Samuel, Chandraker, Manmohan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
di: Yao, Manyi, et al.
Pubblicazione: (2024)
di: Yao, Manyi, et al.
Pubblicazione: (2024)
Generating Enhanced Negatives for Training Language-Based Object Detectors
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
Taming Self-Training for Open-Vocabulary Object Detection
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
di: Zhao, Shiyu, et al.
Pubblicazione: (2023)
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
di: Zhangli, Qilong, et al.
Pubblicazione: (2024)
di: Zhangli, Qilong, et al.
Pubblicazione: (2024)
What to Test Next: Interpretable Coverage Gap Discovery in Driving VLMs
di: Aich, Abhishek, et al.
Pubblicazione: (2026)
di: Aich, Abhishek, et al.
Pubblicazione: (2026)
Tuned Contrastive Learning
di: Animesh, Chaitanya, et al.
Pubblicazione: (2023)
di: Animesh, Chaitanya, et al.
Pubblicazione: (2023)
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
di: Yao, Manyi, et al.
Pubblicazione: (2025)
di: Yao, Manyi, et al.
Pubblicazione: (2025)
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes
di: Soroco, Mauricio, et al.
Pubblicazione: (2026)
di: Soroco, Mauricio, et al.
Pubblicazione: (2026)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
AIDE: An Automatic Data Engine for Object Detection in Autonomous Driving
di: Liang, Mingfu, et al.
Pubblicazione: (2024)
di: Liang, Mingfu, et al.
Pubblicazione: (2024)
Locally Orderless Images for Optimization in Differentiable Rendering
di: Mehta, Ishit, et al.
Pubblicazione: (2025)
di: Mehta, Ishit, et al.
Pubblicazione: (2025)
UDA-Bench: Revisiting Common Assumptions in Unsupervised Domain Adaptation Using a Standardized Framework
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion
di: Li, Haodong, et al.
Pubblicazione: (2026)
di: Li, Haodong, et al.
Pubblicazione: (2026)
Instantaneous Perception of Moving Objects in 3D
di: Liu, Di, et al.
Pubblicazione: (2024)
di: Liu, Di, et al.
Pubblicazione: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
di: Kalluri, Tarun, et al.
Pubblicazione: (2024)
Mapillary Vistas Validation for Fine-Grained Traffic Signs: A Benchmark Revealing Vision-Language Model Limitations
di: Garg, Sparsh, et al.
Pubblicazione: (2025)
di: Garg, Sparsh, et al.
Pubblicazione: (2025)
LidaRF: Delving into Lidar for Neural Radiance Field on Street Scenes
di: Sun, Shanlin, et al.
Pubblicazione: (2024)
di: Sun, Shanlin, et al.
Pubblicazione: (2024)
Drive-1-to-3: Enriching Diffusion Priors for Novel View Synthesis of Real Vehicles
di: Lin, Chuang, et al.
Pubblicazione: (2024)
di: Lin, Chuang, et al.
Pubblicazione: (2024)
PhyCo: Learning Controllable Physical Priors for Generative Motion
di: Narayanan, Sriram, et al.
Pubblicazione: (2026)
di: Narayanan, Sriram, et al.
Pubblicazione: (2026)
LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents
di: He, Yun, et al.
Pubblicazione: (2025)
di: He, Yun, et al.
Pubblicazione: (2025)
HorizonForge: Driving Scene Editing with Any Trajectories and Any Vehicles
di: Wang, Yifan, et al.
Pubblicazione: (2026)
di: Wang, Yifan, et al.
Pubblicazione: (2026)
Open-world Instance Segmentation: Top-down Learning with Bottom-up Supervision
di: Kalluri, Tarun, et al.
Pubblicazione: (2023)
di: Kalluri, Tarun, et al.
Pubblicazione: (2023)
LLM-Assist: Enhancing Closed-Loop Planning with Language-Based Reasoning
di: Sharan, S P, et al.
Pubblicazione: (2023)
di: Sharan, S P, et al.
Pubblicazione: (2023)
Tokenizing Semantic Segmentation with Run Length Encoding
di: Singh, Abhineet, et al.
Pubblicazione: (2026)
di: Singh, Abhineet, et al.
Pubblicazione: (2026)
RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
di: Ghosh, Anurag, et al.
Pubblicazione: (2026)
di: Ghosh, Anurag, et al.
Pubblicazione: (2026)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
Efficient Universal Perception Encoder
di: Zhu, Chenchen, et al.
Pubblicazione: (2026)
di: Zhu, Chenchen, et al.
Pubblicazione: (2026)
GBlobs: Explicit Local Structure via Gaussian Blobs for Improved Cross-Domain LiDAR-based 3D Object Detection
di: Malić, Dušan, et al.
Pubblicazione: (2025)
di: Malić, Dušan, et al.
Pubblicazione: (2025)
LiSu: A Dataset and Method for LiDAR Surface Normal Estimation
di: Malić, Dušan, et al.
Pubblicazione: (2025)
di: Malić, Dušan, et al.
Pubblicazione: (2025)
PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
di: Kim, Donghyun, et al.
Pubblicazione: (2024)
T-MPEDNet: Unveiling the Synergy of Transformer-aware Multiscale Progressive Encoder-Decoder Network with Feature Recalibration for Tumor and Liver Segmentation
di: Raghaw, Chandravardhan Singh, et al.
Pubblicazione: (2025)
di: Raghaw, Chandravardhan Singh, et al.
Pubblicazione: (2025)
NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
di: Jain, Seemandhar, et al.
Pubblicazione: (2026)
di: Jain, Seemandhar, et al.
Pubblicazione: (2026)
Oil Spill Segmentation using Deep Encoder-Decoder models
di: Satyanarayana, Abhishek Ramanathapura, et al.
Pubblicazione: (2023)
di: Satyanarayana, Abhishek Ramanathapura, et al.
Pubblicazione: (2023)
AutoScape: Geometry-Consistent Long-Horizon Scene Generation
di: Chen, Jiacheng, et al.
Pubblicazione: (2025)
di: Chen, Jiacheng, et al.
Pubblicazione: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
di: Xia, Zunhui, et al.
Pubblicazione: (2025)
di: Xia, Zunhui, et al.
Pubblicazione: (2025)
METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
di: Liu, Yuchen, et al.
Pubblicazione: (2025)
di: Liu, Yuchen, et al.
Pubblicazione: (2025)
EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive Generation
di: Xiong, Tianwei, et al.
Pubblicazione: (2026)
di: Xiong, Tianwei, et al.
Pubblicazione: (2026)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Image-Specific Adaptation of Transformer Encoders for Compute-Efficient Segmentation
di: Yao, Manyi, et al.
Pubblicazione: (2024) -
Generating Enhanced Negatives for Training Language-Based Object Detectors
di: Zhao, Shiyu, et al.
Pubblicazione: (2023) -
Taming Self-Training for Open-Vocabulary Object Detection
di: Zhao, Shiyu, et al.
Pubblicazione: (2023) -
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement
di: Khan, Zaid, et al.
Pubblicazione: (2024) -
Resolving Inconsistent Semantics in Multi-Dataset Image Segmentation
di: Zhangli, Qilong, et al.
Pubblicazione: (2024)