Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Peiyu, Peng, Yi, Gan, Yimeng, Hu, Liang, Xie, Tianyidan, Wang, Xiaokun, Wei, Yichen, Tang, Chuanxin, Zhu, Bo, Li, Changshi, Wei, Hongyang, Li, Eric, Song, Xuchen, Liu, Yang, Zhou, Yahui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
von: Wei, Hongyang, et al.
Veröffentlicht: (2026)
von: Wei, Hongyang, et al.
Veröffentlicht: (2026)
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
von: Wang, Xiaokun, et al.
Veröffentlicht: (2025)
von: Wang, Xiaokun, et al.
Veröffentlicht: (2025)
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
von: Wang, Peiyu, et al.
Veröffentlicht: (2025)
von: Wang, Peiyu, et al.
Veröffentlicht: (2025)
Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
von: Peng, Yi, et al.
Veröffentlicht: (2025)
von: Peng, Yi, et al.
Veröffentlicht: (2025)
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs
von: Zeng, Liang, et al.
Veröffentlicht: (2025)
von: Zeng, Liang, et al.
Veröffentlicht: (2025)
Skywork-R1V3 Technical Report
von: Shen, Wei, et al.
Veröffentlicht: (2025)
von: Shen, Wei, et al.
Veröffentlicht: (2025)
LongSkywork: A Training Recipe for Efficiently Extending Context Length in Large Language Models
von: Zhao, Liang, et al.
Veröffentlicht: (2024)
von: Zhao, Liang, et al.
Veröffentlicht: (2024)
Skywork Open Reasoner 1 Technical Report
von: He, Jujie, et al.
Veröffentlicht: (2025)
von: He, Jujie, et al.
Veröffentlicht: (2025)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
von: Zhang, Ruiheng, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiheng, et al.
Veröffentlicht: (2026)
Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
von: Wei, Tianwen, et al.
Veröffentlicht: (2024)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
von: Jian, Ai, et al.
Veröffentlicht: (2025)
von: Jian, Ai, et al.
Veröffentlicht: (2025)
UniShield: Unified Face Attack Detection via KG-Informed Multimodal Reasoning
von: Li, Hongrui, et al.
Veröffentlicht: (2026)
von: Li, Hongrui, et al.
Veröffentlicht: (2026)
UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation
von: Wang, Jiayun, et al.
Veröffentlicht: (2026)
von: Wang, Jiayun, et al.
Veröffentlicht: (2026)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2024)
UniVBench: Towards Unified Evaluation for Video Foundation Models
von: Wei, Jianhui, et al.
Veröffentlicht: (2026)
von: Wei, Jianhui, et al.
Veröffentlicht: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
von: Wang, Dianyi, et al.
Veröffentlicht: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
Unified Medical Image Tokenizer for Autoregressive Synthesis and Understanding
von: Ma, Chenglong, et al.
Veröffentlicht: (2025)
von: Ma, Chenglong, et al.
Veröffentlicht: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
von: Fan, Lijie, et al.
Veröffentlicht: (2025)
von: Fan, Lijie, et al.
Veröffentlicht: (2025)
UniT: Unified Geometry Learning with Group Autoregressive Transformer
von: Wang, Haotian, et al.
Veröffentlicht: (2026)
von: Wang, Haotian, et al.
Veröffentlicht: (2026)
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
von: Guan, Wenhao, et al.
Veröffentlicht: (2025)
UniARM: Towards a Unified Autoregressive Reward Model for Multi-Objective Test-Time Alignment
von: Xie, Hongyan, et al.
Veröffentlicht: (2026)
von: Xie, Hongyan, et al.
Veröffentlicht: (2026)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2025)
von: Liu, Chris Yuhao, et al.
Veröffentlicht: (2025)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
UniECG: Understanding and Generating ECG in One Unified Model
von: Jin, Jiarui, et al.
Veröffentlicht: (2025)
von: Jin, Jiarui, et al.
Veröffentlicht: (2025)
UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
von: Lin, Bin, et al.
Veröffentlicht: (2025)
von: Lin, Bin, et al.
Veröffentlicht: (2025)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
von: Wei, Hongyang, et al.
Veröffentlicht: (2025)
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
von: Zeng, Liang, et al.
Veröffentlicht: (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
UniMo: Unified Motion Generation and Understanding with Chain of Thought
von: Wang, Guocun, et al.
Veröffentlicht: (2026)
von: Wang, Guocun, et al.
Veröffentlicht: (2026)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
UniQueR: Unified Query-based Feedforward 3D Reconstruction
von: Peng, Chensheng, et al.
Veröffentlicht: (2026)
von: Peng, Chensheng, et al.
Veröffentlicht: (2026)
UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
von: Li, Jinke, et al.
Veröffentlicht: (2025)
von: Li, Jinke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
von: Wei, Hongyang, et al.
Veröffentlicht: (2025) -
Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling
von: Wei, Hongyang, et al.
Veröffentlicht: (2026) -
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
von: Wang, Xiaokun, et al.
Veröffentlicht: (2025) -
Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
von: Wang, Peiyu, et al.
Veröffentlicht: (2025) -
Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
von: Peng, Yi, et al.
Veröffentlicht: (2025)