Grid: Omni Visual Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Cong, Luo, Xiangyang, Luo, Hao, Cai, Zijian, Song, Yiren, Zhao, Yunlong, Bai, Yifan, Wang, Fan, He, Yuhang, Gong, Yihong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
by: Wan, Cong, et al.
Published: (2024)
by: Wan, Cong, et al.
Published: (2024)
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
by: Zhao, Zeyang, et al.
Published: (2024)
by: Zhao, Zeyang, et al.
Published: (2024)
OmniPSD: Layered PSD Generation with Diffusion Transformer
by: Liu, Cheng, et al.
Published: (2025)
by: Liu, Cheng, et al.
Published: (2025)
IOTA: Corrective Knowledge-Guided Prompt Learning via Black-White Box Framework
by: Wang, Shaokun, et al.
Published: (2026)
by: Wang, Shaokun, et al.
Published: (2026)
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
by: Bai, Yifan, et al.
Published: (2023)
by: Bai, Yifan, et al.
Published: (2023)
DualCP: Rehearsal-Free Domain-Incremental Learning via Dual-Level Concept Prototype
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
OmniRefiner: Reinforcement-Guided Local Diffusion Refinement
by: Liu, Yaoli, et al.
Published: (2025)
by: Liu, Yaoli, et al.
Published: (2025)
ReMoT: Reinforcement Learning with Motion Contrast Triplets
by: Wan, Cong, et al.
Published: (2026)
by: Wan, Cong, et al.
Published: (2026)
CanonSwap: High-Fidelity and Consistent Video Face Swapping via Canonical Space Modulation
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
by: Song, Yiren, et al.
Published: (2026)
by: Song, Yiren, et al.
Published: (2026)
Continual Novel Class Discovery via Feature Enhancement and Adaptation
by: Yu, Yifan, et al.
Published: (2024)
by: Yu, Yifan, et al.
Published: (2024)
VINO: A Unified Visual Generator with Interleaved OmniModal Context
by: Chen, Junyi, et al.
Published: (2026)
by: Chen, Junyi, et al.
Published: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
by: Peng, Haosong, et al.
Published: (2025)
by: Peng, Haosong, et al.
Published: (2025)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
GFPL: Generative Federated Prototype Learning for Resource-Constrained and Data-Imbalanced Vision Task
by: Lu, Shiwei, et al.
Published: (2026)
by: Lu, Shiwei, et al.
Published: (2026)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
by: Zhang, Yiman, et al.
Published: (2025)
by: Zhang, Yiman, et al.
Published: (2025)
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
by: Zhang, Xiaoyan, et al.
Published: (2026)
by: Zhang, Xiaoyan, et al.
Published: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
by: Zheng, Yiren, et al.
Published: (2026)
by: Zheng, Yiren, et al.
Published: (2026)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
by: Zhou, Jiazhou, et al.
Published: (2025)
by: Zhou, Jiazhou, et al.
Published: (2025)
Beyond Prompt Learning: Continual Adapter for Efficient Rehearsal-Free Continual Learning
by: Gao, Xinyuan, et al.
Published: (2024)
by: Gao, Xinyuan, et al.
Published: (2024)
Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
by: Liang, Susan, et al.
Published: (2026)
by: Liang, Susan, et al.
Published: (2026)
Unleashing the Potential of All Test Samples: Mean-Shift Guided Test-Time Adaptation
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
Consistent Supervised-Unsupervised Alignment for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2025)
by: Han, Jizhou, et al.
Published: (2025)
MSP-ReID: Hairstyle-Robust Cloth-Changing Person Re-Identification
by: He, Xiangyang, et al.
Published: (2026)
by: He, Xiangyang, et al.
Published: (2026)
P2L-CA: An Effective Parameter Tuning Framework for Rehearsal-Free Multi-Label Class-Incremental Learning
by: Dong, Songlin, et al.
Published: (2026)
by: Dong, Songlin, et al.
Published: (2026)
OmniParser: A Unified Framework for Text Spotting, Key Information Extraction and Table Recognition
by: Wan, Jianqiang, et al.
Published: (2024)
by: Wan, Jianqiang, et al.
Published: (2024)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
by: Jia, Yiduo, et al.
Published: (2026)
by: Jia, Yiduo, et al.
Published: (2026)
Beyond CLIP Generalization: Against Forward&Backward Forgetting Adapter for Continual Learning of Vision-Language Models
by: Dong, Songlin, et al.
Published: (2025)
by: Dong, Songlin, et al.
Published: (2025)
GOAL: Geometrically Optimal Alignment for Continual Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2026)
by: Han, Jizhou, et al.
Published: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
Adding Additional Control to One-Step Diffusion with Joint Distribution Matching
by: Luo, Yihong, et al.
Published: (2025)
by: Luo, Yihong, et al.
Published: (2025)
Learning Like Humans: Analogical Concept Learning for Generalized Category Discovery
by: Han, Jizhou, et al.
Published: (2026)
by: Han, Jizhou, et al.
Published: (2026)
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
Trajectory-Diversity-Driven Robust Vision-and-Language Navigation
by: Li, Jiangyang, et al.
Published: (2026)
by: Li, Jiangyang, et al.
Published: (2026)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
Similar Items
-
Prompt-Agnostic Adversarial Perturbation for Customized Diffusion Models
by: Wan, Cong, et al.
Published: (2024) -
Projecting Points to Axes: Oriented Object Detection via Point-Axis Representation
by: Zhao, Zeyang, et al.
Published: (2024) -
OmniPSD: Layered PSD Generation with Diffusion Transformer
by: Liu, Cheng, et al.
Published: (2025) -
IOTA: Corrective Knowledge-Guided Prompt Learning via Black-White Box Framework
by: Wang, Shaokun, et al.
Published: (2026) -
ARTrackV2: Prompting Autoregressive Tracker Where to Look and How to Describe
by: Bai, Yifan, et al.
Published: (2023)