Generalized Visual Relation Detection with Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Kaifeng, Chen, Siqi, Zhang, Hanwang, Xiao, Jun, Zhuang, Yueting, Sun, Qianru |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024)
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024)
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024)
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024)
Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models
von: Zhu, Beier, et al.
Veröffentlicht: (2023)
von: Zhu, Beier, et al.
Veröffentlicht: (2023)
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
Few-shot Learner Parameterization by Diffusion Time-steps
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
Adaptive Begin-of-Video Tokens for Autoregressive Video Diffusion Models
von: Cheng, Tianle, et al.
Veröffentlicht: (2025)
von: Cheng, Tianle, et al.
Veröffentlicht: (2025)
Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-Of-Distribution Generalization
von: Qi, Jiaxin, et al.
Veröffentlicht: (2022)
von: Qi, Jiaxin, et al.
Veröffentlicht: (2022)
Exploring Diffusion Time-steps for Unsupervised Representation Learning
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)
3D Question Answering via only 2D Vision-Language Models
von: Wang, Fengyun, et al.
Veröffentlicht: (2025)
von: Wang, Fengyun, et al.
Veröffentlicht: (2025)
DiT as Real-Time Rerenderer: Streaming Video Stylization with Autoregressive Diffusion Transformer
von: Lyu, Hengye, et al.
Veröffentlicht: (2026)
von: Lyu, Hengye, et al.
Veröffentlicht: (2026)
Real-Time Motion-Controllable Autoregressive Video Diffusion
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
von: Zhao, Kesen, et al.
Veröffentlicht: (2025)
Video Anomaly Detection and Explanation via Large Language Models
von: Lv, Hui, et al.
Veröffentlicht: (2024)
von: Lv, Hui, et al.
Veröffentlicht: (2024)
Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
From Easy to Hard: Learning Curricular Shape-aware Features for Robust Panoptic Scene Graph Generation
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
Physically Plausible Human-Object Rendering from Sparse Views via 3D Gaussian Splatting
von: Wang, Weiquan, et al.
Veröffentlicht: (2025)
von: Wang, Weiquan, et al.
Veröffentlicht: (2025)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
Rendering Multi-Human and Multi-Object with 3D Gaussian Splatting
von: Wang, Weiquan, et al.
Veröffentlicht: (2026)
von: Wang, Weiquan, et al.
Veröffentlicht: (2026)
FlowDC: Flow-Based Decoupling-Decay for Complex Image Editing
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
von: Jiang, Yilei, et al.
Veröffentlicht: (2025)
Distilling Parallel Gradients for Fast ODE Solvers of Diffusion Models
von: Zhu, Beier, et al.
Veröffentlicht: (2025)
von: Zhu, Beier, et al.
Veröffentlicht: (2025)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
von: Zhang, Yuxi, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxi, et al.
Veröffentlicht: (2025)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
von: Gao, Minghe, et al.
Veröffentlicht: (2025)
NICEST: Noisy Label Correction and Training for Robust Scene Graph Generation
von: Li, Lin, et al.
Veröffentlicht: (2022)
von: Li, Lin, et al.
Veröffentlicht: (2022)
Non-confusing Generation of Customized Concepts in Diffusion Models
von: Lin, Wang, et al.
Veröffentlicht: (2024)
von: Lin, Wang, et al.
Veröffentlicht: (2024)
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
IDPro: Flexible Interactive Video Object Segmentation by ID-queried Concurrent Propagation
von: Li, Kexin, et al.
Veröffentlicht: (2024)
von: Li, Kexin, et al.
Veröffentlicht: (2024)
Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Reducing Class-Wise Performance Disparity via Margin Regularization
von: Zhu, Beier, et al.
Veröffentlicht: (2026)
von: Zhu, Beier, et al.
Veröffentlicht: (2026)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
Incorporating Visual Correspondence into Diffusion Model for Virtual Try-On
von: Wan, Siqi, et al.
Veröffentlicht: (2025)
von: Wan, Siqi, et al.
Veröffentlicht: (2025)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Towards Better Alignment: Training Diffusion Models with Reinforcement Learning Against Sparse Rewards
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
von: Hu, Zijing, et al.
Veröffentlicht: (2025)
Diffusion Time-step Curriculum for One Image to 3D Generation
von: Yi, Xuanyu, et al.
Veröffentlicht: (2024)
von: Yi, Xuanyu, et al.
Veröffentlicht: (2024)
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2025)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
Two Causal Principles for Improving Visual Dialog
von: Qi, Jiaxin, et al.
Veröffentlicht: (2019)
von: Qi, Jiaxin, et al.
Veröffentlicht: (2019)
Learning De-Biased Representations for Remote-Sensing Imagery
von: Tian, Zichen, et al.
Veröffentlicht: (2024)
von: Tian, Zichen, et al.
Veröffentlicht: (2024)
CrossView Suite: Harnessing Cross-view Spatial Intelligence of MLLMs with Dataset, Model and Benchmark
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ViD-GPT: Introducing GPT-style Autoregressive Generation in Video Diffusion Models
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024) -
Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing
von: Gao, Kaifeng, et al.
Veröffentlicht: (2024) -
Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models
von: Zhu, Beier, et al.
Veröffentlicht: (2023) -
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
von: Zhao, Kesen, et al.
Veröffentlicht: (2025) -
Few-shot Learner Parameterization by Diffusion Time-steps
von: Yue, Zhongqi, et al.
Veröffentlicht: (2024)