Step by Step Network
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Dongchen, Ye, Tianzhu, Xia, Zhuofan, Chen, Kaiyi, Wang, Yulin, Chen, Hanting, Huang, Gao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Agent Attention: On the Integration of Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2023)
di: Han, Dongchen, et al.
Pubblicazione: (2023)
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
di: Pu, Yifan, et al.
Pubblicazione: (2024)
di: Pu, Yifan, et al.
Pubblicazione: (2024)
GSVA: Generalized Segmentation via Multimodal Large Language Models
di: Xia, Zhuofan, et al.
Pubblicazione: (2023)
di: Xia, Zhuofan, et al.
Pubblicazione: (2023)
Bridging the Divide: Reconsidering Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2024)
di: Han, Dongchen, et al.
Pubblicazione: (2024)
One Step Diffusion-based Super-Resolution with Time-Aware Distillation
di: He, Xiao, et al.
Pubblicazione: (2024)
di: He, Xiao, et al.
Pubblicazione: (2024)
Demystify Mamba in Vision: A Linear Attention Perspective
di: Han, Dongchen, et al.
Pubblicazione: (2024)
di: Han, Dongchen, et al.
Pubblicazione: (2024)
STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
di: Zhao, Zhonghan, et al.
Pubblicazione: (2024)
di: Zhao, Zhonghan, et al.
Pubblicazione: (2024)
Linear-Time Global Visual Modeling without Explicit Attention
di: He, Ruize, et al.
Pubblicazione: (2026)
di: He, Ruize, et al.
Pubblicazione: (2026)
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
di: Han, Zongyan, et al.
Pubblicazione: (2025)
di: Han, Zongyan, et al.
Pubblicazione: (2025)
Vision Transformers are Circulant Attention Learners
di: Han, Dongchen, et al.
Pubblicazione: (2025)
di: Han, Dongchen, et al.
Pubblicazione: (2025)
One Step Learning, One Step Review
di: Huang, Xiaolong, et al.
Pubblicazione: (2024)
di: Huang, Xiaolong, et al.
Pubblicazione: (2024)
Frequency Domain Modality-invariant Feature Learning for Visible-infrared Person Re-Identification
di: Li, Yulin, et al.
Pubblicazione: (2024)
di: Li, Yulin, et al.
Pubblicazione: (2024)
Denoising Diffusion Step-aware Models
di: Yang, Shuai, et al.
Pubblicazione: (2023)
di: Yang, Shuai, et al.
Pubblicazione: (2023)
Linear Differential Vision Transformer: Learning Visual Contrasts via Pairwise Differentials
di: Pu, Yifan, et al.
Pubblicazione: (2025)
di: Pu, Yifan, et al.
Pubblicazione: (2025)
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
di: Pu, Yifan, et al.
Pubblicazione: (2025)
di: Pu, Yifan, et al.
Pubblicazione: (2025)
Step-GUI Technical Report
di: Yan, Haolong, et al.
Pubblicazione: (2025)
di: Yan, Haolong, et al.
Pubblicazione: (2025)
Training an Open-Vocabulary Monocular 3D Object Detection Model without 3D Data
di: Huang, Rui, et al.
Pubblicazione: (2024)
di: Huang, Rui, et al.
Pubblicazione: (2024)
Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
di: Chen, Honghao, et al.
Pubblicazione: (2025)
di: Chen, Honghao, et al.
Pubblicazione: (2025)
Know Your Step: Faster and Better Alignment for Flow Matching Models via Step-aware Advantages
di: Yue, Zhixiong, et al.
Pubblicazione: (2026)
di: Yue, Zhixiong, et al.
Pubblicazione: (2026)
Self-Adversarial One Step Generation via Condition Shifting
di: Liu, Deyuan, et al.
Pubblicazione: (2026)
di: Liu, Deyuan, et al.
Pubblicazione: (2026)
VividFace: High-Quality and Efficient One-Step Diffusion For Video Face Enhancement
di: Zhang, Shulian, et al.
Pubblicazione: (2025)
di: Zhang, Shulian, et al.
Pubblicazione: (2025)
Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
di: Zhang, Jiahao, et al.
Pubblicazione: (2023)
di: Zhang, Jiahao, et al.
Pubblicazione: (2023)
Accelerating Diffusion Decoders via Multi-Scale Sampling and One-Step Distillation
di: Wang, Chuhan, et al.
Pubblicazione: (2026)
di: Wang, Chuhan, et al.
Pubblicazione: (2026)
A Step to Decouple Optimization in 3DGS
di: Ding, Renjie, et al.
Pubblicazione: (2026)
di: Ding, Renjie, et al.
Pubblicazione: (2026)
RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
di: Zhang, Ruoxuan, et al.
Pubblicazione: (2025)
One-Step Diffusion Model for Image Motion-Deblurring
di: Liu, Xiaoyang, et al.
Pubblicazione: (2025)
di: Liu, Xiaoyang, et al.
Pubblicazione: (2025)
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
di: Wang, Siting, et al.
Pubblicazione: (2026)
di: Wang, Siting, et al.
Pubblicazione: (2026)
StepAL: Step-aware Active Learning for Cataract Surgical Videos
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
di: Souček, Tomáš, et al.
Pubblicazione: (2024)
di: Souček, Tomáš, et al.
Pubblicazione: (2024)
MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
di: Xue, Haotian, et al.
Pubblicazione: (2025)
di: Xue, Haotian, et al.
Pubblicazione: (2025)
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
di: Lv, Zhengyao, et al.
Pubblicazione: (2026)
di: Lv, Zhengyao, et al.
Pubblicazione: (2026)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
One-Step Event-Driven High-Speed Autofocus
di: Bao, Yuhan, et al.
Pubblicazione: (2025)
di: Bao, Yuhan, et al.
Pubblicazione: (2025)
OSDFace: One-Step Diffusion Model for Face Restoration
di: Wang, Jingkai, et al.
Pubblicazione: (2024)
di: Wang, Jingkai, et al.
Pubblicazione: (2024)
FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction
di: Lin, Jiang, et al.
Pubblicazione: (2025)
di: Lin, Jiang, et al.
Pubblicazione: (2025)
Bridging Fidelity-Reality with Controllable One-Step Diffusion for Image Super-Resolution
di: Chen, Hao, et al.
Pubblicazione: (2025)
di: Chen, Hao, et al.
Pubblicazione: (2025)
DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing
di: Yang, Desong, et al.
Pubblicazione: (2026)
di: Yang, Desong, et al.
Pubblicazione: (2026)
An End-to-End Robust Point Cloud Semantic Segmentation Network with Single-Step Conditional Diffusion Models
di: Qu, Wentao, et al.
Pubblicazione: (2024)
di: Qu, Wentao, et al.
Pubblicazione: (2024)
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
di: Bhattacharyya, Apratim, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Apratim, et al.
Pubblicazione: (2025)
Step Saver: Predicting Minimum Denoising Steps for Diffusion Model Image Generation
di: Yu, Jean, et al.
Pubblicazione: (2024)
di: Yu, Jean, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Agent Attention: On the Integration of Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2023) -
Efficient Diffusion Transformer with Step-wise Dynamic Attention Mediators
di: Pu, Yifan, et al.
Pubblicazione: (2024) -
GSVA: Generalized Segmentation via Multimodal Large Language Models
di: Xia, Zhuofan, et al.
Pubblicazione: (2023) -
Bridging the Divide: Reconsidering Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2024) -
One Step Diffusion-based Super-Resolution with Time-Aware Distillation
di: He, Xiao, et al.
Pubblicazione: (2024)