Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bozhou, Guan, Yushuo, Li, Haolin, Zeng, Bohan, Ji, Yiyan, Ding, Yue, Wan, Pengfei, Gai, Kun, Zhang, Yuanxing, Zhang, Wentao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
by: Ding, Yue, et al.
Published: (2026)
by: Ding, Yue, et al.
Published: (2026)
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026)
by: Dai, Yifan, et al.
Published: (2026)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
by: Li, Bozhou, et al.
Published: (2025)
by: Li, Bozhou, et al.
Published: (2025)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026)
by: Chen, Xinlong, et al.
Published: (2026)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
SemanticGen: Video Generation in Semantic Space
by: Bai, Jianhong, et al.
Published: (2025)
by: Bai, Jianhong, et al.
Published: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
by: Ye, Zixuan, et al.
Published: (2025)
by: Ye, Zixuan, et al.
Published: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
by: Wang, Qixun, et al.
Published: (2025)
by: Wang, Qixun, et al.
Published: (2025)
Rethinking Cross-Layer Information Routing in Diffusion Transformers
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
Are Bigger Encoders Always Better in Vision Large Models?
by: Li, Bozhou, et al.
Published: (2024)
by: Li, Bozhou, et al.
Published: (2024)
Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
by: Huang, Yuzhou, et al.
Published: (2025)
by: Huang, Yuzhou, et al.
Published: (2025)
Towards Precise Scaling Laws for Video Diffusion Transformers
by: Yin, Yuanyang, et al.
Published: (2024)
by: Yin, Yuanyang, et al.
Published: (2024)
TFWT: Tabular Feature Weighting with Transformer
by: Zhang, Xinhao, et al.
Published: (2024)
by: Zhang, Xinhao, et al.
Published: (2024)
The First Prompt Counts the Most! An Evaluation of Large Language Models on Iterative Example-Based Code Generation
by: Fu, Yingjie, et al.
Published: (2024)
by: Fu, Yingjie, et al.
Published: (2024)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
MasRouter: Learning to Route LLMs for Multi-Agent Systems
by: Yue, Yanwei, et al.
Published: (2025)
by: Yue, Yanwei, et al.
Published: (2025)
Semantic Score Distillation Sampling for Compositional Text-to-3D Generation
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
by: Xu, Yiyan, et al.
Published: (2026)
by: Xu, Yiyan, et al.
Published: (2026)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
by: Tang, Yuqi, et al.
Published: (2026)
by: Tang, Yuqi, et al.
Published: (2026)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States
by: Zhang, Bozhou, et al.
Published: (2024)
by: Zhang, Bozhou, et al.
Published: (2024)
Refining the Understanding of Operator Size Dynamics in Open Quantum Systems
by: Jiang, Haolin, et al.
Published: (2025)
by: Jiang, Haolin, et al.
Published: (2025)
Restoration Adaptation for Semantic Segmentation on Low Quality Images
by: Guan, Kai, et al.
Published: (2026)
by: Guan, Kai, et al.
Published: (2026)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
by: Wu, Shengqiong, et al.
Published: (2025)
by: Wu, Shengqiong, et al.
Published: (2025)
Score Augmentation for Diffusion Models
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
by: Liu, Tengfei, et al.
Published: (2026)
by: Liu, Tengfei, et al.
Published: (2026)
MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
by: Shi, Yang, et al.
Published: (2025)
by: Shi, Yang, et al.
Published: (2025)
A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows
by: Han, Yuxuan, et al.
Published: (2026)
by: Han, Yuxuan, et al.
Published: (2026)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
by: Shi, Minglei, et al.
Published: (2025)
by: Shi, Minglei, et al.
Published: (2025)
Exploring the Key Features of Repeating Fast Radio Bursts with Machine Learning
by: Sun, Wan-Peng, et al.
Published: (2024)
by: Sun, Wan-Peng, et al.
Published: (2024)
VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
XFEM‐Based Analysis of Crack Propagation in Cold‐Region Railway Tunnels Under Coupled Train Loads and Localized Frost Heave
by: Wei Chen, et al.
Published: (2026)
by: Wei Chen, et al.
Published: (2026)
Similar Items
-
GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models
by: Li, Bozhou, et al.
Published: (2025) -
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
by: Ding, Yue, et al.
Published: (2026) -
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
by: Li, Bozhou, et al.
Published: (2025) -
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
by: Dai, Yifan, et al.
Published: (2026) -
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
by: Shi, Minglei, et al.
Published: (2025)