PixFoundation: Are We Heading in the Right Direction with Pixel-level Vision Foundation Models?
Fuente:
arXiv
Saved in:
| Main Author: | Siam, Mennatullah |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025)
by: Siam, Mennatullah
Published: (2025)
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025)
by: Cheshmi, Leila, et al.
Published: (2025)
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
by: Teuber, Carolin, et al.
Published: (2026)
by: Teuber, Carolin, et al.
Published: (2026)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023)
by: Goyal, Raghav, et al.
Published: (2023)
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)
by: Karim, Rezaul, et al.
Published: (2023)
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
by: Li, Zhenyu, et al.
Published: (2025)
by: Li, Zhenyu, et al.
Published: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
by: Jin, Youngwan, et al.
Published: (2024)
by: Jin, Youngwan, et al.
Published: (2024)
Boosting Gaze Object Prediction via Pixel-level Supervision from Vision Foundation Model
by: Jin, Yang, et al.
Published: (2024)
by: Jin, Yang, et al.
Published: (2024)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
PixNerd: Pixel Neural Field Diffusion
by: Wang, Shuai, et al.
Published: (2025)
by: Wang, Shuai, et al.
Published: (2025)
Towards Foundation Models for 3D Vision: How Close Are We?
by: Zuo, Yiming, et al.
Published: (2024)
by: Zuo, Yiming, et al.
Published: (2024)
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2025)
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
by: Hossain, Mir Rayat Imtiaz, et al.
Published: (2024)
PixOOD: Pixel-Level Out-of-Distribution Detection
by: Vojíř, Tomáš, et al.
Published: (2024)
by: Vojíř, Tomáš, et al.
Published: (2024)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
by: Ou, Ruizhe, et al.
Published: (2025)
by: Ou, Ruizhe, et al.
Published: (2025)
Pixelis: Reasoning in Pixels, from Seeing to Acting
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Can We Simplify Slide-level Fine-tuning of Pathology Foundation Models?
by: Li, Jiawen, et al.
Published: (2025)
by: Li, Jiawen, et al.
Published: (2025)
Generalized Few-Shot Semantic Segmentation in Remote Sensing: Challenge and Benchmark
by: Broni-Bediako, Clifford, et al.
Published: (2024)
by: Broni-Bediako, Clifford, et al.
Published: (2024)
PixIE: Prompted Pixel-Space Low-Light Image Enhancement
by: Lin, Ruirui, et al.
Published: (2026)
by: Lin, Ruirui, et al.
Published: (2026)
Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
by: Fuster-Barceló, Caterina, et al.
Published: (2026)
Are We on the Right Way for Evaluating Large Vision-Language Models?
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
A Vision Centric Remote Sensing Benchmark
by: Adejumo, Abduljaleel, et al.
Published: (2025)
by: Adejumo, Abduljaleel, et al.
Published: (2025)
Sapiens: Foundation for Human Vision Models
by: Khirodkar, Rawal, et al.
Published: (2024)
by: Khirodkar, Rawal, et al.
Published: (2024)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
Cracks in the Foundation: A Civil Infrastructure Dataset to Challenge Vision Foundation Models
by: Farronato, Nicola, et al.
Published: (2026)
by: Farronato, Nicola, et al.
Published: (2026)
Beyond Pixel-Wise Supervision for Medical Image Segmentation: From Traditional Models to Foundation Models
by: Shi, Yuyan, et al.
Published: (2024)
by: Shi, Yuyan, et al.
Published: (2024)
Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
by: Shen, Ruiqi, et al.
Published: (2025)
by: Shen, Ruiqi, et al.
Published: (2025)
Explainability for Vision Foundation Models: A Survey
by: Kazmierczak, Rémi, et al.
Published: (2025)
by: Kazmierczak, Rémi, et al.
Published: (2025)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
by: Guo, Jianyuan, et al.
Published: (2024)
by: Guo, Jianyuan, et al.
Published: (2024)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
by: Zheng, Yaoyan, et al.
Published: (2025)
by: Zheng, Yaoyan, et al.
Published: (2025)
Even the "Devil" has Rights!
by: Siam, Mennatullah
Published: (2024)
by: Siam, Mennatullah
Published: (2024)
Insect-Foundation: A Foundation Model and Large Multimodal Dataset for Vision-Language Insect Understanding
by: Truong, Thanh-Dat, et al.
Published: (2025)
by: Truong, Thanh-Dat, et al.
Published: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
by: Liang, Wenqi, et al.
Published: (2025)
by: Liang, Wenqi, et al.
Published: (2025)
Quantifying and Learning Static vs. Dynamic Information in Deep Spatiotemporal Networks
by: Kowal, Matthew, et al.
Published: (2022)
by: Kowal, Matthew, et al.
Published: (2022)
CanViT: Toward Active-Vision Foundation Models
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
by: Berreby, Yohaï-Eliel, et al.
Published: (2026)
Vision Foundation Models as Generalist Tokenizers for Image Generation
by: Zheng, Anlin, et al.
Published: (2026)
by: Zheng, Anlin, et al.
Published: (2026)
Bootstrapping SparseFormers from Vision Foundation Models
by: Gao, Ziteng, et al.
Published: (2023)
by: Gao, Ziteng, et al.
Published: (2023)
Annotation Free Semantic Segmentation with Vision Foundation Models
by: Seifi, Soroush, et al.
Published: (2024)
by: Seifi, Soroush, et al.
Published: (2024)
Similar Items
-
PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
by: Siam, Mennatullah
Published: (2025) -
Multiscale Video Transformers for Class Agnostic Segmentation in Autonomous Driving
by: Cheshmi, Leila, et al.
Published: (2025) -
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
by: Teuber, Carolin, et al.
Published: (2026) -
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
by: Goyal, Raghav, et al.
Published: (2023) -
MED-VT++: Unifying Multimodal Learning with a Multiscale Encoder-Decoder Video Transformer
by: Karim, Rezaul, et al.
Published: (2023)