See Further When Clear: Curriculum Consistency Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yunpeng, Liu, Boxiao, Zhang, Yi, Hou, Xingzhong, Song, Guanglu, Liu, Yu, You, Haihang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting
by: Hou, Xingzhong, et al.
Published: (2025)
by: Hou, Xingzhong, et al.
Published: (2025)
Brain-Inspired Efficient Pruning: Exploiting Criticality in Spiking Neural Networks
by: Chen, Shuo, et al.
Published: (2023)
by: Chen, Shuo, et al.
Published: (2023)
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023)
by: Xue, Zeyue, et al.
Published: (2023)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026)
by: Liu, Yi, et al.
Published: (2026)
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
Enhancing Vision-Language Model with Unmasked Token Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
SeeClear: Semantic Distillation Enhances Pixel Condensation for Video Super-Resolution
by: Tang, Qi, et al.
Published: (2024)
by: Tang, Qi, et al.
Published: (2024)
Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
by: Ma, Bingqi, et al.
Published: (2024)
by: Ma, Bingqi, et al.
Published: (2024)
Seeing Clearly, Reasoning Confidently: Plug-and-Play Remedies for Vision Language Model Blindness
by: Hu, Xin, et al.
Published: (2026)
by: Hu, Xin, et al.
Published: (2026)
Seeing Further on the Shoulders of Giants: Knowledge Inheritance for Vision Foundation Models
by: Huang, Jiabo, et al.
Published: (2025)
by: Huang, Jiabo, et al.
Published: (2025)
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation
by: Ge, Xingtong, et al.
Published: (2026)
by: Ge, Xingtong, et al.
Published: (2026)
FouriScale: A Frequency Perspective on Training-Free High-Resolution Image Synthesis
by: Huang, Linjiang, et al.
Published: (2024)
by: Huang, Linjiang, et al.
Published: (2024)
Phased Consistency Models
by: Wang, Fu-Yun, et al.
Published: (2024)
by: Wang, Fu-Yun, et al.
Published: (2024)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation
by: Chang, Chun-Peng, et al.
Published: (2025)
by: Chang, Chun-Peng, et al.
Published: (2025)
Pixelis: Reasoning in Pixels, from Seeing to Acting
by: Zhou, Yunpeng
Published: (2026)
by: Zhou, Yunpeng
Published: (2026)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
by: Wu, Zhiheng, et al.
Published: (2026)
by: Wu, Zhiheng, et al.
Published: (2026)
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
by: Tang, Feilong, et al.
Published: (2025)
by: Tang, Feilong, et al.
Published: (2025)
Learning to See in the Extremely Dark
by: Jiang, Hai, et al.
Published: (2025)
by: Jiang, Hai, et al.
Published: (2025)
AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
by: He, Dailan, et al.
Published: (2026)
by: He, Dailan, et al.
Published: (2026)
Seeing Cells Clearly: Evaluating Machine Vision Strategies for Microglia Centroid Detection in 3D Images
by: Zhang, Youjia
Published: (2025)
by: Zhang, Youjia
Published: (2025)
ADT: Tuning Diffusion Models with Adversarial Supervision
by: Shen, Dazhong, et al.
Published: (2025)
by: Shen, Dazhong, et al.
Published: (2025)
Semantics and Content Matter: Towards Multi-Prior Hierarchical Mamba for Image Deraining
by: Yu, Zhaocheng, et al.
Published: (2025)
by: Yu, Zhaocheng, et al.
Published: (2025)
SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification
by: Wang, Xiaoying, et al.
Published: (2026)
by: Wang, Xiaoying, et al.
Published: (2026)
Acc3D: Accelerating Single Image to 3D Diffusion Models via Edge Consistency Guided Score Distillation
by: Liu, Kendong, et al.
Published: (2025)
by: Liu, Kendong, et al.
Published: (2025)
Supervise Less, See More: Training-free Nuclear Instance Segmentation with Prototype-Guided Prompting
by: Zhang, Wen, et al.
Published: (2025)
by: Zhang, Wen, et al.
Published: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
by: Zong, Zhuofan, et al.
Published: (2024)
by: Zong, Zhuofan, et al.
Published: (2024)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
by: Liu, Mengmeng, et al.
Published: (2025)
by: Liu, Mengmeng, et al.
Published: (2025)
Channel Consistency Prior and Self-Reconstruction Strategy Based Unsupervised Image Deraining
by: Dong, Guanglu, et al.
Published: (2025)
by: Dong, Guanglu, et al.
Published: (2025)
Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model
by: Zhang, Yuxuan, et al.
Published: (2024)
by: Zhang, Yuxuan, et al.
Published: (2024)
Adaptive Whole-Body PET Image Denoising Using 3D Diffusion Models with ControlNet
by: Yu, Boxiao, et al.
Published: (2024)
by: Yu, Boxiao, et al.
Published: (2024)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
Rethinking the Spatial Inconsistency in Classifier-Free Diffusion Guidance
by: Shen, Dazhong, et al.
Published: (2024)
by: Shen, Dazhong, et al.
Published: (2024)
LookOut: Real-World Humanoid Egocentric Navigation
by: Pan, Boxiao, et al.
Published: (2025)
by: Pan, Boxiao, et al.
Published: (2025)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
by: Zhang, Xiaofeng, et al.
Published: (2024)
by: Zhang, Xiaofeng, et al.
Published: (2024)
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
by: Song, Yiren, et al.
Published: (2025)
by: Song, Yiren, et al.
Published: (2025)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
by: Yang, Siqi, et al.
Published: (2025)
by: Yang, Siqi, et al.
Published: (2025)
Seeing Motion at Nighttime with an Event Camera
by: Liu, Haoyue, et al.
Published: (2024)
by: Liu, Haoyue, et al.
Published: (2024)
Similar Items
-
Towards Seamless Borders: A Method for Mitigating Inconsistencies in Image Inpainting and Outpainting
by: Hou, Xingzhong, et al.
Published: (2025) -
Brain-Inspired Efficient Pruning: Exploiting Criticality in Spiking Neural Networks
by: Chen, Shuo, et al.
Published: (2023) -
RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
by: Xue, Zeyue, et al.
Published: (2023) -
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
by: Liu, Yi, et al.
Published: (2026) -
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
by: Liu, Yexin, et al.
Published: (2024)