Semantic-Aware Prefix Learning for Token-Efficient Image Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Qingfeng, Zhang, Haoxian, He, Xu, Tang, Songlin, Fang, Zhixue, Liu, Xiaoqiang, Li, Pengfei Wan Guoqi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
por: Fang, Zhixue, et al.
Publicado: (2026)
por: Fang, Zhixue, et al.
Publicado: (2026)
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
por: He, Xu, et al.
Publicado: (2025)
por: He, Xu, et al.
Publicado: (2025)
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
por: Peng, Ziqiao, et al.
Publicado: (2025)
por: Peng, Ziqiao, et al.
Publicado: (2025)
IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation
por: Xu, Zhufeng, et al.
Publicado: (2026)
por: Xu, Zhufeng, et al.
Publicado: (2026)
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
por: Chen, Ming, et al.
Publicado: (2025)
por: Chen, Ming, et al.
Publicado: (2025)
Learning Semantic-Aware Threshold for Multi-Label Image Recognition with Partial Labels
por: Ruan, Haoxian, et al.
Publicado: (2025)
por: Ruan, Haoxian, et al.
Publicado: (2025)
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
por: Chen, Hejia, et al.
Publicado: (2025)
por: Chen, Hejia, et al.
Publicado: (2025)
GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation
por: Hu, Wentao, et al.
Publicado: (2025)
por: Hu, Wentao, et al.
Publicado: (2025)
Kling-MotionControl Technical Report
por: Kling Team, et al.
Publicado: (2026)
por: Kling Team, et al.
Publicado: (2026)
Implicit Priors Editing in Stable Diffusion via Targeted Token Adjustment
por: He, Feng, et al.
Publicado: (2024)
por: He, Feng, et al.
Publicado: (2024)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
por: Liao, Xinyao, et al.
Publicado: (2026)
por: Liao, Xinyao, et al.
Publicado: (2026)
Learning Semantic-Aware Representation in Visual-Language Models for Multi-Label Recognition with Partial Labels
por: Ruan, Haoxian, et al.
Publicado: (2024)
por: Ruan, Haoxian, et al.
Publicado: (2024)
CATANet: Efficient Content-Aware Token Aggregation for Lightweight Image Super-Resolution
por: Liu, Xin, et al.
Publicado: (2025)
por: Liu, Xin, et al.
Publicado: (2025)
Autoregressive Image Generation with Randomized Parallel Decoding
por: Li, Haopeng, et al.
Publicado: (2025)
por: Li, Haopeng, et al.
Publicado: (2025)
Beyond Inserting: Learning Identity Embedding for Semantic-Fidelity Personalized Diffusion Generation
por: Li, Yang, et al.
Publicado: (2024)
por: Li, Yang, et al.
Publicado: (2024)
PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
por: Wang, Ao, et al.
Publicado: (2024)
por: Wang, Ao, et al.
Publicado: (2024)
FaceGCD: Generalized Face Discovery via Dynamic Prefix Generation
por: Oh, Yunseok, et al.
Publicado: (2025)
por: Oh, Yunseok, et al.
Publicado: (2025)
Training-Free Efficient Video Generation via Dynamic Token Carving
por: Zhang, Yuechen, et al.
Publicado: (2025)
por: Zhang, Yuechen, et al.
Publicado: (2025)
SETA: Semantic-Aware Token Augmentation for Domain Generalization
por: Guo, Jintao, et al.
Publicado: (2024)
por: Guo, Jintao, et al.
Publicado: (2024)
SEAL: Semantic-Aware Hierarchical Learning for Generalized Category Discovery
por: He, Zhenqi, et al.
Publicado: (2025)
por: He, Zhenqi, et al.
Publicado: (2025)
KlingAvatar 2.0 Technical Report
por: Kling Team, et al.
Publicado: (2025)
por: Kling Team, et al.
Publicado: (2025)
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
por: Guo, Jianzhu, et al.
Publicado: (2024)
por: Guo, Jianzhu, et al.
Publicado: (2024)
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models
por: Liu, Yingen, et al.
Publicado: (2024)
por: Liu, Yingen, et al.
Publicado: (2024)
Semantic-Aware Representation Learning via Conditional Transport for Multi-Label Image Classification
por: Xie, Ren-Dong, et al.
Publicado: (2025)
por: Xie, Ren-Dong, et al.
Publicado: (2025)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
por: Han, Minghao, et al.
Publicado: (2025)
por: Han, Minghao, et al.
Publicado: (2025)
Scalable Autoregressive Image Generation with Mamba
por: Li, Haopeng, et al.
Publicado: (2024)
por: Li, Haopeng, et al.
Publicado: (2024)
MetaSegNet: Metadata-collaborative Vision-Language Representation Learning for Semantic Segmentation of Remote Sensing Images
por: Wang, Libo, et al.
Publicado: (2023)
por: Wang, Libo, et al.
Publicado: (2023)
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning
por: Wang, Xuan, et al.
Publicado: (2023)
por: Wang, Xuan, et al.
Publicado: (2023)
AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation
por: Yin, Xuanhua, et al.
Publicado: (2026)
por: Yin, Xuanhua, et al.
Publicado: (2026)
Learning Spectral-Decomposed Tokens for Domain Generalized Semantic Segmentation
por: Yi, Jingjun, et al.
Publicado: (2024)
por: Yi, Jingjun, et al.
Publicado: (2024)
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
por: Li, Jiaze, et al.
Publicado: (2025)
por: Li, Jiaze, et al.
Publicado: (2025)
CATP: Confidence-Aware Token Pruning for Camouflaged Object Detection
por: Gao, Yuhan, et al.
Publicado: (2026)
por: Gao, Yuhan, et al.
Publicado: (2026)
Semantic One-Dimensional Tokenizer for Image Reconstruction and Generation
por: Qu, Yunpeng, et al.
Publicado: (2026)
por: Qu, Yunpeng, et al.
Publicado: (2026)
Importance-Based Token Merging for Efficient Image and Video Generation
por: Wu, Haoyu, et al.
Publicado: (2024)
por: Wu, Haoyu, et al.
Publicado: (2024)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
por: Zou, Shihao, et al.
Publicado: (2025)
por: Zou, Shihao, et al.
Publicado: (2025)
RASR: Retrieval-Augmented Semantic Reasoning for Fake News Video Detection
por: Li, Hui, et al.
Publicado: (2026)
por: Li, Hui, et al.
Publicado: (2026)
CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
por: Li, Yanshu, et al.
Publicado: (2025)
por: Li, Yanshu, et al.
Publicado: (2025)
Vision Transformer with Super Token Sampling
por: Huang, Huaibo, et al.
Publicado: (2022)
por: Huang, Huaibo, et al.
Publicado: (2022)
Caption Generation for Dongba Paintings via Prompt Learning and Semantic Fusion
por: Qian, Shuangwu, et al.
Publicado: (2026)
por: Qian, Shuangwu, et al.
Publicado: (2026)
DuPL: Dual Student with Trustworthy Progressive Learning for Robust Weakly Supervised Semantic Segmentation
por: Wu, Yuanchen, et al.
Publicado: (2024)
por: Wu, Yuanchen, et al.
Publicado: (2024)
Ejemplares similares
-
3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
por: Fang, Zhixue, et al.
Publicado: (2026) -
From Inpainting to Editing: Unlocking Robust Mask-Free Visual Dubbing via Generative Bootstrapping
por: He, Xu, et al.
Publicado: (2025) -
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
por: Peng, Ziqiao, et al.
Publicado: (2025) -
IM-Animation: An Implicit Motion Representation for Identity-decoupled Character Animation
por: Xu, Zhufeng, et al.
Publicado: (2026) -
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
por: Chen, Ming, et al.
Publicado: (2025)