Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Tao, Li, Xiangtai, Huang, Zilong, Li, Yanwei, Lei, Weixian, Deng, Xueqing, Chen, Shihao, Ji, Shunping, Feng, Jiashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
von: Lei, Weixian, et al.
Veröffentlicht: (2025)
von: Lei, Weixian, et al.
Veröffentlicht: (2025)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
von: Zhang, Tao, et al.
Veröffentlicht: (2024)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
von: Athar, Ali, et al.
Veröffentlicht: (2024)
von: Athar, Ali, et al.
Veröffentlicht: (2024)
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)
Dense360: Dense Understanding from Omnidirectional Panoramas
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
von: Li, Xiangtai, et al.
Veröffentlicht: (2025)
von: Li, Xiangtai, et al.
Veröffentlicht: (2025)
PixelThink: Towards Efficient Chain-of-Pixel Reasoning
von: Wang, Song, et al.
Veröffentlicht: (2025)
von: Wang, Song, et al.
Veröffentlicht: (2025)
Goal2Pixel: Grounding Goals to Pixels for Vision-Language Navigation
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
von: Bao, Muyi, et al.
Veröffentlicht: (2026)
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
von: Zhou, Yikang, et al.
Veröffentlicht: (2024)
Transforming Weather Data from Pixel to Latent Space
von: Zhao, Sijie, et al.
Veröffentlicht: (2025)
von: Zhao, Sijie, et al.
Veröffentlicht: (2025)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level
von: Deng, Andong, et al.
Veröffentlicht: (2024)
von: Deng, Andong, et al.
Veröffentlicht: (2024)
4th PVUW MeViS 3rd Place Report: Sa2VA
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
von: Shabbir, Akashah, et al.
Veröffentlicht: (2025)
PixelDiT: Pixel Diffusion Transformers for Image Generation
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
FARMER: Flow AutoRegressive Transformer over Pixels
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
von: Zong, Yongshuo, et al.
Veröffentlicht: (2025)
von: Zong, Yongshuo, et al.
Veröffentlicht: (2025)
The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
von: Niu, Quanzhu, et al.
Veröffentlicht: (2025)
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
von: Liu, Chang, et al.
Veröffentlicht: (2026)
von: Liu, Chang, et al.
Veröffentlicht: (2026)
T-Pixel2Mesh: Combining Global and Local Transformer for 3D Mesh Generation from a Single Image
von: Zhang, Shijie, et al.
Veröffentlicht: (2024)
von: Zhang, Shijie, et al.
Veröffentlicht: (2024)
Dual-Scale Transformer for Large-Scale Single-Pixel Imaging
von: Qu, Gang, et al.
Veröffentlicht: (2024)
von: Qu, Gang, et al.
Veröffentlicht: (2024)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2025)
Beyond Pixels: Text Enhances Generalization in Real-World Image Restoration
von: Sun, Haoze, et al.
Veröffentlicht: (2024)
von: Sun, Haoze, et al.
Veröffentlicht: (2024)
VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
von: Liang, Wenqi, et al.
Veröffentlicht: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
von: Chen, Shoufa, et al.
Veröffentlicht: (2025)
von: Chen, Shoufa, et al.
Veröffentlicht: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
von: Zhao, Weiheng, et al.
Veröffentlicht: (2025)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Pixel Is Not a Barrier: An Effective Evasion Attack for Pixel-Domain Diffusion Models
von: Shih, Chun-Yen, et al.
Veröffentlicht: (2024)
von: Shih, Chun-Yen, et al.
Veröffentlicht: (2024)
One Pixel is All I Need
von: Siqin, Deng, et al.
Veröffentlicht: (2024)
von: Siqin, Deng, et al.
Veröffentlicht: (2024)
HyperDiT: Hyper-Connected Transformers for High-Fidelity Pixel-Space Diffusion
von: He, Yu, et al.
Veröffentlicht: (2026)
von: He, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
von: Lei, Weixian, et al.
Veröffentlicht: (2025) -
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
von: Zhang, Tao, et al.
Veröffentlicht: (2024) -
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025) -
ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation
von: Athar, Ali, et al.
Veröffentlicht: (2024) -
PixelLM: Pixel Reasoning with Large Multimodal Model
von: Ren, Zhongwei, et al.
Veröffentlicht: (2023)