Coarse-Guided Visual Generation via Weighted h-Transform Sampling
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yanghao, Jiang, Ziqi, Wang, Zhen, Chen, Long |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024)
by: Jiang, Ziqi, et al.
Published: (2024)
Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation
by: Chen, Hongxu, et al.
Published: (2026)
by: Chen, Hongxu, et al.
Published: (2026)
Exploring Cross-Modal Flows for Few-Shot Learning
by: Jiang, Ziqi, et al.
Published: (2025)
by: Jiang, Ziqi, et al.
Published: (2025)
Treat Stillness with Movement: Remote Sensing Change Detection via Coarse-grained Temporal Foregrounds Mining
by: Wang, Xixi, et al.
Published: (2024)
by: Wang, Xixi, et al.
Published: (2024)
Target-aware Image Editing via Cycle-consistent Constraints
by: Wang, Yanghao, et al.
Published: (2025)
by: Wang, Yanghao, et al.
Published: (2025)
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
by: Chen, Hongxu, et al.
Published: (2026)
by: Chen, Hongxu, et al.
Published: (2026)
AG-VAS: Anchor-Guided Zero-Shot Visual Anomaly Segmentation with Large Multimodal Models
by: Qu, Zhen, et al.
Published: (2026)
by: Qu, Zhen, et al.
Published: (2026)
CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
by: Cao, Zongsheng, et al.
Published: (2025)
by: Cao, Zongsheng, et al.
Published: (2025)
UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards
by: Wang, Jun, et al.
Published: (2026)
by: Wang, Jun, et al.
Published: (2026)
Simulating the Visual World with Artificial Intelligence: A Roadmap
by: Yue, Jingtong, et al.
Published: (2025)
by: Yue, Jingtong, et al.
Published: (2025)
StrLoRA: Towards Streaming Continual Visual Instruction Tuning for MLLMs
by: Che, Chang, et al.
Published: (2026)
by: Che, Chang, et al.
Published: (2026)
Attention Guided CAM: Visual Explanations of Vision Transformer Guided by Self-Attention
by: Leem, Saebom, et al.
Published: (2024)
by: Leem, Saebom, et al.
Published: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
by: Dai, Ming, et al.
Published: (2025)
by: Dai, Ming, et al.
Published: (2025)
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling
by: Li, Wenjie, et al.
Published: (2026)
by: Li, Wenjie, et al.
Published: (2026)
TinyDrop: Tiny Model Guided Token Dropping for Vision Transformers
by: Wang, Guoxin, et al.
Published: (2025)
by: Wang, Guoxin, et al.
Published: (2025)
RandAR: Decoder-only Autoregressive Visual Generation in Random Orders
by: Pang, Ziqi, et al.
Published: (2024)
by: Pang, Ziqi, et al.
Published: (2024)
SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning
by: Wang, Ziqi, et al.
Published: (2024)
by: Wang, Ziqi, et al.
Published: (2024)
EVE: Verifiable Self-Evolution of MLLMs via Executable Visual Transformations
by: Heng, Yongrui, et al.
Published: (2026)
by: Heng, Yongrui, et al.
Published: (2026)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
DVGT: Driving Visual Geometry Transformer
by: Zuo, Sicheng, et al.
Published: (2025)
by: Zuo, Sicheng, et al.
Published: (2025)
Fine-Grained Scene Graph Generation via Sample-Level Bias Prediction
by: Li, Yansheng, et al.
Published: (2024)
by: Li, Yansheng, et al.
Published: (2024)
CF-VLA: Efficient Coarse-to-Fine Action Generation for Vision-Language-Action Policies
by: Du, Fan, et al.
Published: (2026)
by: Du, Fan, et al.
Published: (2026)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
by: Pu, Yifan, et al.
Published: (2025)
by: Pu, Yifan, et al.
Published: (2025)
PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
by: Liu, Kang, et al.
Published: (2025)
by: Liu, Kang, et al.
Published: (2025)
Uncertainty-Guided Coarse-to-Fine Tumor Segmentation with Anatomy-Aware Post-Processing
by: Isler, Ilkin Sevgi, et al.
Published: (2025)
by: Isler, Ilkin Sevgi, et al.
Published: (2025)
Progressive Fine-to-Coarse Reconstruction for Accurate Low-Bit Post-Training Quantization in Vision Transformers
by: Ding, Rui, et al.
Published: (2024)
by: Ding, Rui, et al.
Published: (2024)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
by: Tang, Tianci, et al.
Published: (2026)
by: Tang, Tianci, et al.
Published: (2026)
LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
by: Che, Chang, et al.
Published: (2025)
by: Che, Chang, et al.
Published: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Training-Free Text-Guided Image Editing with Visual Autoregressive Model
by: Wang, Yufei, et al.
Published: (2025)
by: Wang, Yufei, et al.
Published: (2025)
MambaScope: Coarse-to-Fine Scoping for Efficient Vision Mamba
by: Liu, Shanhui, et al.
Published: (2025)
by: Liu, Shanhui, et al.
Published: (2025)
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
by: Zhou, Jinxing, et al.
Published: (2025)
by: Zhou, Jinxing, et al.
Published: (2025)
VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping
by: Dong, Haotian, et al.
Published: (2025)
by: Dong, Haotian, et al.
Published: (2025)
Dynamic Pondering Sparsity-aware Mixture-of-Experts Transformer for Event Stream based Visual Object Tracking
by: Wang, Shiao, et al.
Published: (2026)
by: Wang, Shiao, et al.
Published: (2026)
Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning
by: Chen, Junkai, et al.
Published: (2026)
by: Chen, Junkai, et al.
Published: (2026)
DilateQuant: Accurate and Efficient Diffusion Quantization via Weight Dilation
by: Liu, Xuewen, et al.
Published: (2024)
by: Liu, Xuewen, et al.
Published: (2024)
RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy
by: Chen, Aiyue, et al.
Published: (2025)
by: Chen, Aiyue, et al.
Published: (2025)
Unsupervised Multi-View Visual Anomaly Detection via Progressive Homography-Guided Alignment
by: Chen, Xintao, et al.
Published: (2025)
by: Chen, Xintao, et al.
Published: (2025)
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Similar Items
-
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
by: Jiang, Ziqi, et al.
Published: (2024) -
Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation
by: Chen, Hongxu, et al.
Published: (2026) -
Exploring Cross-Modal Flows for Few-Shot Learning
by: Jiang, Ziqi, et al.
Published: (2025) -
Treat Stillness with Movement: Remote Sensing Change Detection via Coarse-grained Temporal Foregrounds Mining
by: Wang, Xixi, et al.
Published: (2024) -
Target-aware Image Editing via Cycle-consistent Constraints
by: Wang, Yanghao, et al.
Published: (2025)