A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Liang, Tan, Sinan, Cai, Zefan, Xie, Weichu, Zhao, Haozhe, Zhang, Yichi, Lin, Junyang, Bai, Jinze, Liu, Tianyu, Chang, Baobao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)
Mitigating Language-Level Performance Disparity in mPLMs via Teacher Language Selection and Cross-lingual Self-Distillation
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
di: Zhao, Haozhe, et al.
Pubblicazione: (2023)
Improving the Robustness of Distantly-Supervised Named Entity Recognition via Uncertainty-Aware Teacher Learning and Student-Student Collaborative Learning
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
di: Chen, Liang, et al.
Pubblicazione: (2024)
di: Chen, Liang, et al.
Pubblicazione: (2024)
SANTA: Separate Strategies for Inaccurate and Incomplete Annotation Noise in Distantly-Supervised Named Entity Recognition
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
di: Si, Shuzheng, et al.
Pubblicazione: (2023)
BabyVision: Visual Reasoning Beyond Language
di: Chen, Liang, et al.
Pubblicazione: (2026)
di: Chen, Liang, et al.
Pubblicazione: (2026)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
di: Huang, Jinsheng, et al.
Pubblicazione: (2024)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
LLM Critics Help Catch Bugs in Mathematics: Towards a Better Mathematical Verifier with Natural Language Feedback
di: Gao, Bofei, et al.
Pubblicazione: (2024)
di: Gao, Bofei, et al.
Pubblicazione: (2024)
VeCAF: Vision-language Collaborative Active Finetuning with Training Objective Awareness
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
di: Zhang, Rongyu, et al.
Pubblicazione: (2024)
A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
di: Zhao, Haozhe, et al.
Pubblicazione: (2024)
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
di: An, Kaikai, et al.
Pubblicazione: (2024)
di: An, Kaikai, et al.
Pubblicazione: (2024)
Autoregressive Image Generation with Vision Full-view Prompt
di: Cai, Miaomiao, et al.
Pubblicazione: (2025)
di: Cai, Miaomiao, et al.
Pubblicazione: (2025)
Improving Event Definition Following For Zero-Shot Event Detection
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
di: Tang, Xiangru, et al.
Pubblicazione: (2023)
Revisiting Non-Autoregressive Transformers for Efficient Image Synthesis
di: Ni, Zanlin, et al.
Pubblicazione: (2024)
di: Ni, Zanlin, et al.
Pubblicazione: (2024)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
di: Tang, Haotian, et al.
Pubblicazione: (2024)
di: Tang, Haotian, et al.
Pubblicazione: (2024)
Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark
di: Bai, Zeyu
Pubblicazione: (2026)
di: Bai, Zeyu
Pubblicazione: (2026)
NEP: Autoregressive Image Editing via Next Editing Token Prediction
di: Wu, Huimin, et al.
Pubblicazione: (2025)
di: Wu, Huimin, et al.
Pubblicazione: (2025)
A Cost‐Aware and Latency‐Benefit Evaluation‐Based Task Scheduling Optimization Strategy in Apache Spark
di: Qingsong Xu, et al.
Pubblicazione: (2025)
di: Qingsong Xu, et al.
Pubblicazione: (2025)
NeuronMLP: Efficient LLM Inference via Singular Value Decomposition Compression and Tiling on AWS Trainium
di: Song, Dinghong, et al.
Pubblicazione: (2025)
di: Song, Dinghong, et al.
Pubblicazione: (2025)
V2Flow: Unifying Visual Tokenization and Large Language Model Vocabularies for Autoregressive Image Generation
di: Zhang, Guiwei, et al.
Pubblicazione: (2025)
di: Zhang, Guiwei, et al.
Pubblicazione: (2025)
Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression
di: Zhen, Dingcheng, et al.
Pubblicazione: (2025)
di: Zhen, Dingcheng, et al.
Pubblicazione: (2025)
A GPU Implementation of Multi-Guiding Spark Fireworks Algorithm for Efficient Black-Box Neural Network Optimization
di: Meng, Xiangrui, et al.
Pubblicazione: (2025)
di: Meng, Xiangrui, et al.
Pubblicazione: (2025)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
Blur-aware Spatio-temporal Sparse Transformer for Video Deblurring
di: Zhang, Huicong, et al.
Pubblicazione: (2024)
di: Zhang, Huicong, et al.
Pubblicazione: (2024)
Universal Approximation of Visual Autoregressive Transformers
di: Chen, Yifang, et al.
Pubblicazione: (2025)
di: Chen, Yifang, et al.
Pubblicazione: (2025)
Revisiting Multimodal Positional Encoding in Vision-Language Models
di: Huang, Jie, et al.
Pubblicazione: (2025)
di: Huang, Jie, et al.
Pubblicazione: (2025)
From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
di: Cai, Zefan, et al.
Pubblicazione: (2025)
di: Cai, Zefan, et al.
Pubblicazione: (2025)
Head-Aware Key-Value Compression for Efficient Autoregressive Image Generation
di: Liang, Guotao, et al.
Pubblicazione: (2026)
di: Liang, Guotao, et al.
Pubblicazione: (2026)
MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining
di: Wu, Ruiqi, et al.
Pubblicazione: (2025)
di: Wu, Ruiqi, et al.
Pubblicazione: (2025)
Decoupled Diffusion Sparks Adaptive Scene Generation
di: Zhou, Yunsong, et al.
Pubblicazione: (2025)
di: Zhou, Yunsong, et al.
Pubblicazione: (2025)
Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
di: Chen, Zhuokun, et al.
Pubblicazione: (2025)
di: Chen, Zhuokun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
di: Chen, Liang, et al.
Pubblicazione: (2025) -
An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
di: Chen, Liang, et al.
Pubblicazione: (2024) -
PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain
di: Chen, Liang, et al.
Pubblicazione: (2024) -
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
di: Zhao, Haozhe, et al.
Pubblicazione: (2024) -
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
di: Zhao, Haozhe, et al.
Pubblicazione: (2025)