Saved in:
| Main Authors: | Li, Yangyang, Liu, Daqing, Liu, Wu, He, Allen, Liu, Xinchen, Zhang, Yongdong, Jin, Guoqing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.12242 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026)
by: He, Allen, et al.
Published: (2026)
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024)
by: Xiong, Tianwei, et al.
Published: (2024)
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
by: Mao, Fangyuan, et al.
Published: (2025)
by: Mao, Fangyuan, et al.
Published: (2025)
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)
by: Erogullari, Eren, et al.
Published: (2025)
Disentangling Regional Primitives for Image Generation
by: Chen, Zhengting, et al.
Published: (2024)
by: Chen, Zhengting, et al.
Published: (2024)
OmniGen: Unified Image Generation
by: Xiao, Shitao, et al.
Published: (2024)
by: Xiao, Shitao, et al.
Published: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
by: Wang, Lifu, et al.
Published: (2025)
by: Wang, Lifu, et al.
Published: (2025)
CoDeGAN: Contrastive Disentanglement for Generative Adversarial Network
by: Zhao, Jiangwei, et al.
Published: (2021)
by: Zhao, Jiangwei, et al.
Published: (2021)
Controllable Video Generation with Provable Disentanglement
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Learning to Infer Generative Template Programs for Visual Concepts
by: Jones, R. Kenny, et al.
Published: (2024)
by: Jones, R. Kenny, et al.
Published: (2024)
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
by: Kim, Hyeonjin, et al.
Published: (2026)
by: Kim, Hyeonjin, et al.
Published: (2026)
DEAL: Disentangle and Localize Concept-level Explanations for VLMs
by: Li, Tang, et al.
Published: (2024)
by: Li, Tang, et al.
Published: (2024)
ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
by: Choi, Jinho, et al.
Published: (2025)
by: Choi, Jinho, et al.
Published: (2025)
VideoPrism: A Foundational Visual Encoder for Video Understanding
by: Zhao, Long, et al.
Published: (2024)
by: Zhao, Long, et al.
Published: (2024)
Boosting the Transferability of Adversarial Examples via Local Mixup and Adaptive Step Size
by: Liu, Junlin, et al.
Published: (2024)
by: Liu, Junlin, et al.
Published: (2024)
SigFormer: Sparse Signal-Guided Transformer for Multi-Modal Human Action Segmentation
by: Liu, Qi, et al.
Published: (2023)
by: Liu, Qi, et al.
Published: (2023)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
by: Feng, Chun, et al.
Published: (2024)
by: Feng, Chun, et al.
Published: (2024)
Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
by: Qian, Yilue, et al.
Published: (2023)
by: Qian, Yilue, et al.
Published: (2023)
Explore the Limits of Omni-modal Pretraining at Scale
by: Zhang, Yiyuan, et al.
Published: (2024)
by: Zhang, Yiyuan, et al.
Published: (2024)
Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
by: Liu, Chenruo, et al.
Published: (2025)
by: Liu, Chenruo, et al.
Published: (2025)
Pre-Training Meta-Rule Selection Policy for Visual Generative Abductive Learning
by: Jin, Yu, et al.
Published: (2025)
by: Jin, Yu, et al.
Published: (2025)
Symbolic Disentangled Representations for Images
by: Korchemnyi, Alexandr, et al.
Published: (2024)
by: Korchemnyi, Alexandr, et al.
Published: (2024)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
by: Jiang, Dongzhi, et al.
Published: (2025)
by: Jiang, Dongzhi, et al.
Published: (2025)
Pretrained Reversible Generation as Unsupervised Visual Representation Learning
by: Xue, Rongkun, et al.
Published: (2024)
by: Xue, Rongkun, et al.
Published: (2024)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
Unsupervised Synthetic Image Attribution: Alignment and Disentanglement
by: Liu, Zongfang, et al.
Published: (2026)
by: Liu, Zongfang, et al.
Published: (2026)
DTL: Disentangled Transfer Learning for Visual Recognition
by: Fu, Minghao, et al.
Published: (2023)
by: Fu, Minghao, et al.
Published: (2023)
Denoising Multi-Beta VAE: Representation Learning for Disentanglement and Generation
by: Uppal, Anshuk, et al.
Published: (2025)
by: Uppal, Anshuk, et al.
Published: (2025)
Omni-SILA: Towards Omni-scene Driven Visual Sentiment Identifying, Locating and Attributing in Videos
by: Luo, Jiamin, et al.
Published: (2025)
by: Luo, Jiamin, et al.
Published: (2025)
Prism: Spectral-Aware Block-Sparse Attention
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Explaining Generalization Power of a DNN Using Interactive Concepts
by: Zhou, Huilin, et al.
Published: (2023)
by: Zhou, Huilin, et al.
Published: (2023)
Pix2Code: Learning to Compose Neural Visual Concepts as Programs
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
Omni-Dish: Photorealistic and Faithful Image Generation and Editing for Arbitrary Chinese Dishes
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
Learning Generalized and Flexible Trajectory Models from Omni-Semantic Supervision
by: Zhu, Yuanshao, et al.
Published: (2025)
by: Zhu, Yuanshao, et al.
Published: (2025)
MVP-CBM:Multi-layer Visual Preference-enhanced Concept Bottleneck Model for Explainable Medical Image Classification
by: Wang, Chunjiang, et al.
Published: (2025)
by: Wang, Chunjiang, et al.
Published: (2025)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
by: Zhang, Yiman, et al.
Published: (2025)
by: Zhang, Yiman, et al.
Published: (2025)
Understanding Visual Concepts Across Models
by: Trabucco, Brandon, et al.
Published: (2024)
by: Trabucco, Brandon, et al.
Published: (2024)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Automatically Generating Visual Hallucination Test Cases for Multimodal Large Language Models
by: Liu, Zhongye, et al.
Published: (2024)
by: Liu, Zhongye, et al.
Published: (2024)
Similar Items
-
A Paradigm Shift: Fully End-to-End Training for Temporal Sentence Grounding in Videos
by: He, Allen, et al.
Published: (2026) -
VideoDPO: Omni-Preference Alignment for Video Diffusion Generation
by: Liu, Runtao, et al.
Published: (2024) -
Editing Massive Concepts in Text-to-Image Diffusion Models
by: Xiong, Tianwei, et al.
Published: (2024) -
Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation
by: Mao, Fangyuan, et al.
Published: (2025) -
Post-Hoc Concept Disentanglement: From Correlated to Isolated Concept Representations
by: Erogullari, Eren, et al.
Published: (2025)