Towards Transformer-Based Aligned Generation with Self-Coherence Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Shulei, Lin, Wang, Huang, Hai, Wang, Hanting, Cai, Sihang, Han, WenKang, Jin, Tao, Chen, Jingyuan, Sun, Jiacheng, Zhu, Jieming, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
by: Wang, Hanting, et al.
Published: (2025)
by: Wang, Hanting, et al.
Published: (2025)
MPCODER: Multi-user Personalized Code Generator with Explicit and Implicit Style Representation Learning
by: Dai, Zhenlong, et al.
Published: (2024)
by: Dai, Zhenlong, et al.
Published: (2024)
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Chat-Driven Text Generation and Interaction for Person Retrieval
by: Xie, Zequn, et al.
Published: (2025)
by: Xie, Zequn, et al.
Published: (2025)
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
by: Wang, Hanting, et al.
Published: (2025)
by: Wang, Hanting, et al.
Published: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Open-set Cross Modal Generalization via Multimodal Unified Representation
by: Huang, Hai, et al.
Published: (2025)
by: Huang, Hai, et al.
Published: (2025)
Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
by: Han, Wenkang, et al.
Published: (2025)
by: Han, Wenkang, et al.
Published: (2025)
On Linear Separation Capacity of Self-Supervised Representation Learning
by: Wang, Shulei
Published: (2023)
by: Wang, Shulei
Published: (2023)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
Implicit Guidance and Explicit Representation of Semantic Information in Points Cloud: A Survey
by: Tang, Jingyuan, et al.
Published: (2025)
by: Tang, Jingyuan, et al.
Published: (2025)
Towards Dynamic and Small Objects Refinement for Unsupervised Domain Adaptative Nighttime Semantic Segmentation
by: Pan, Jingyi, et al.
Published: (2023)
by: Pan, Jingyi, et al.
Published: (2023)
HART: Human Aligned Reconstruction Transformer
by: Chen, Xiyi, et al.
Published: (2025)
by: Chen, Xiyi, et al.
Published: (2025)
CARD: Channel Aligned Robust Blend Transformer for Time Series Forecasting
by: Xue, Wang, et al.
Published: (2023)
by: Xue, Wang, et al.
Published: (2023)
Efficient Controllable Diffusion via Optimal Classifier Guidance
by: Oertell, Owen, et al.
Published: (2025)
by: Oertell, Owen, et al.
Published: (2025)
Nexus: Higher-Order Attention Mechanisms in Transformers
by: Chen, Hanting, et al.
Published: (2025)
by: Chen, Hanting, et al.
Published: (2025)
EAGER-LLM: Enhancing Large Language Models as Recommenders through Exogenous Behavior-Semantic Integration
by: Hong, Minjie, et al.
Published: (2025)
by: Hong, Minjie, et al.
Published: (2025)
Learning to Align, Aligning to Learn: A Unified Approach for Self-Optimized Alignment
by: Wang, Haowen, et al.
Published: (2025)
by: Wang, Haowen, et al.
Published: (2025)
Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance
by: Li, Jihang, et al.
Published: (2026)
by: Li, Jihang, et al.
Published: (2026)
BrainDreamer: Reasoning-Coherent and Controllable Image Generation from EEG Brain Signals via Language Guidance
by: Wang, Ling, et al.
Published: (2024)
by: Wang, Ling, et al.
Published: (2024)
RAT: Retrieval-Augmented Transformer for Click-Through Rate Prediction
by: Li, Yushen, et al.
Published: (2024)
by: Li, Yushen, et al.
Published: (2024)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Autonomous Robotic System with Optical Coherence Tomography Guidance for Vascular Anastomosis
by: Haworth, Jesse, et al.
Published: (2024)
by: Haworth, Jesse, et al.
Published: (2024)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
U-REPA: Aligning Diffusion U-Nets to ViTs
by: Tian, Yuchuan, et al.
Published: (2025)
by: Tian, Yuchuan, et al.
Published: (2025)
Semantic Residual for Multimodal Unified Discrete Representation
by: Huang, Hai, et al.
Published: (2024)
by: Huang, Hai, et al.
Published: (2024)
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Dynamical Characterization of Quantum Coherence
by: Wang, Hai, et al.
Published: (2024)
by: Wang, Hai, et al.
Published: (2024)
Systemic risk measures with markets volatility
by: Sun, Fei, et al.
Published: (2018)
by: Sun, Fei, et al.
Published: (2018)
Efficient Prompting for Continual Adaptation to Missing Modalities
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
GalaxAlign: Mimicking Citizen Scientists' Multimodal Guidance for Galaxy Morphology Analysis
by: Wang, Ruoqi, et al.
Published: (2024)
by: Wang, Ruoqi, et al.
Published: (2024)
OSInsert: Towards High-authenticity and High-fidelity Image Composition
by: Wang, Jingyuan, et al.
Published: (2026)
by: Wang, Jingyuan, et al.
Published: (2026)
Chiral Organic Synapses Based on Self‐Assembled Charge‐Transfer Co‐Crystal Helix/TIPS‐Pentacene Heterojunctions
by: Shulei Chen, et al.
Published: (2026)
by: Shulei Chen, et al.
Published: (2026)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
by: Lin, Wang, et al.
Published: (2026)
by: Lin, Wang, et al.
Published: (2026)
HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Unified Diffusion-Based Rigid and Non-Rigid Editing with Text and Image Guidance
by: Wang, Jiacheng, et al.
Published: (2024)
by: Wang, Jiacheng, et al.
Published: (2024)
Towards Real-World Aerial Vision Guidance with Categorical 6D Pose Tracker
by: Sun, Jingtao, et al.
Published: (2024)
by: Sun, Jingtao, et al.
Published: (2024)
Augmentation Invariant Manifold Learning
by: Wang, Shulei
Published: (2022)
by: Wang, Shulei
Published: (2022)
Towards Instance Segmentation with Polygon Detection Transformers
by: Sun, Jiacheng, et al.
Published: (2026)
by: Sun, Jiacheng, et al.
Published: (2026)
Similar Items
-
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
by: Wang, Hanting, et al.
Published: (2025) -
MPCODER: Multi-user Personalized Code Generator with Explicit and Implicit Style Representation Learning
by: Dai, Zhenlong, et al.
Published: (2024) -
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
by: Xie, Zequn, et al.
Published: (2026) -
Chat-Driven Text Generation and Interaction for Person Retrieval
by: Xie, Zequn, et al.
Published: (2025) -
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
by: Wang, Hanting, et al.
Published: (2025)