Not all tokens contribute equally to diffusion learning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Guoqing, Shi, Lu, Xu, Wanru, Zhang, Linna, Wang, Sen, Wang, Fangfang, Cen, Yigang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
by: Shen, Hao, et al.
Published: (2024)
by: Shen, Hao, et al.
Published: (2024)
Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
by: Zhang, Guoqing, et al.
Published: (2025)
by: Zhang, Guoqing, et al.
Published: (2025)
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024)
by: Kan, Shichao, et al.
Published: (2024)
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
by: Kan, Shichao, et al.
Published: (2026)
by: Kan, Shichao, et al.
Published: (2026)
CFIS-YOLO: A Lightweight Multi-Scale Fusion Network for Edge-Deployable Wood Defect Detection
by: Kang, Jincheng, et al.
Published: (2025)
by: Kang, Jincheng, et al.
Published: (2025)
MMA: Multimodal Memory Agent
by: Lu, Yihao, et al.
Published: (2026)
by: Lu, Yihao, et al.
Published: (2026)
Imagine with the Teacher: Complete Shape in a Multi-View Distillation Way
by: Luo, Zhanpeng, et al.
Published: (2025)
by: Luo, Zhanpeng, et al.
Published: (2025)
LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
by: Zhang, Zhenghao, et al.
Published: (2025)
by: Zhang, Zhenghao, et al.
Published: (2025)
3D representation in 512-Byte:Variational tokenizer is the key for autoregressive 3D generation
by: Zhang, Jinzhi, et al.
Published: (2024)
by: Zhang, Jinzhi, et al.
Published: (2024)
FIRE: Robust Detection of Diffusion-Generated Images via Frequency-Guided Reconstruction Error
by: Chu, Beilin, et al.
Published: (2024)
by: Chu, Beilin, et al.
Published: (2024)
Differential Contrastive Training for Gaze Estimation
by: Zhang, Lin, et al.
Published: (2025)
by: Zhang, Lin, et al.
Published: (2025)
Group Critical-token Policy Optimization for Autoregressive Image Generation
by: Zhang, Guohui, et al.
Published: (2025)
by: Zhang, Guohui, et al.
Published: (2025)
Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental Learning
by: Zhang, Haojie, et al.
Published: (2025)
by: Zhang, Haojie, et al.
Published: (2025)
Diffusion-Occ: 3D Point Cloud Completion via Occupancy Diffusion
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
Adapting Foundation Models for Few-Shot Medical Image Segmentation: Actively and Sequentially
by: Yang, Jingyun, et al.
Published: (2025)
by: Yang, Jingyun, et al.
Published: (2025)
Learning What is Worth Learning: Active and Sequential Domain Adaptation for Multi-modal Gross Tumor Volume Segmentation
by: Yang, Jingyun, et al.
Published: (2025)
by: Yang, Jingyun, et al.
Published: (2025)
A Simple Baseline for Unifying Understanding, Generation, and Editing via Vanilla Next-token Prediction
by: Zhu, Jie, et al.
Published: (2026)
by: Zhu, Jie, et al.
Published: (2026)
DAP: A Discrete-token Autoregressive Planner for Autonomous Driving
by: Ye, Bowen, et al.
Published: (2025)
by: Ye, Bowen, et al.
Published: (2025)
Identity Clue Refinement and Enhancement for Visible-Infrared Person Re-Identification
by: Zhang, Guoqing, et al.
Published: (2025)
by: Zhang, Guoqing, et al.
Published: (2025)
AQUA-SLAM: Tightly-Coupled Underwater Acoustic-Visual-Inertial SLAM with Sensor Calibration
by: Xu, Shida, et al.
Published: (2025)
by: Xu, Shida, et al.
Published: (2025)
A Geometric Algorithm for Blood Vessel Reconstruction from Skeletal Representation
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
Scaling Multi-Camera 3D Object Detection through Weak-to-Strong Eliciting
by: Lu, Hao, et al.
Published: (2024)
by: Lu, Hao, et al.
Published: (2024)
Reducing catastrophic forgetting of incremental learning in the absence of rehearsal memory with task-specific token
by: Choi, Young Jo, et al.
Published: (2024)
by: Choi, Young Jo, et al.
Published: (2024)
Reduced Spatial Dependency for More General Video-level Deepfake Detection
by: Chu, Beilin, et al.
Published: (2025)
by: Chu, Beilin, et al.
Published: (2025)
Zero-shot sketch-based remote sensing image retrieval based on multi-level and attention-guided tokenization
by: Yang, Bo, et al.
Published: (2024)
by: Yang, Bo, et al.
Published: (2024)
Hierarchical Feature Learning for Medical Point Clouds via State Space Model
by: Zhang, Guoqing, et al.
Published: (2025)
by: Zhang, Guoqing, et al.
Published: (2025)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
by: Cui, Chenhang, et al.
Published: (2024)
by: Cui, Chenhang, et al.
Published: (2024)
ImagePiece: Content-aware Re-tokenization for Efficient Image Recognition
by: Yoa, Seungdong, et al.
Published: (2024)
by: Yoa, Seungdong, et al.
Published: (2024)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
Comparison of Autoencoders for tokenization of ASL datasets
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
by: Praun-Petrovic, Vouk, et al.
Published: (2025)
Towards Real-time Video Compressive Sensing on Mobile Devices
by: Cao, Miao, et al.
Published: (2024)
by: Cao, Miao, et al.
Published: (2024)
CLIP-guided Prototype Modulating for Few-shot Action Recognition
by: Wang, Xiang, et al.
Published: (2023)
by: Wang, Xiang, et al.
Published: (2023)
Boosting Cross-Domain Point Classification via Distilling Relational Priors from 2D Transformers
by: Zou, Longkun, et al.
Published: (2024)
by: Zou, Longkun, et al.
Published: (2024)
Continual Learning for Segment Anything Model Adaptation
by: Yang, Jinglong, et al.
Published: (2024)
by: Yang, Jinglong, et al.
Published: (2024)
Beluga Whale Detection from Satellite Imagery with Point Labels
by: Zheng, Yijie, et al.
Published: (2025)
by: Zheng, Yijie, et al.
Published: (2025)
High-Fidelity Medical Shape Generation via Skeletal Latent Diffusion
by: Zhang, Guoqing, et al.
Published: (2026)
by: Zhang, Guoqing, et al.
Published: (2026)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
Vision Transformer based Random Walk for Group Re-Identification
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
REPS: Reconstruction-based Point Cloud Sampling
by: Zhang, Guoqing, et al.
Published: (2024)
by: Zhang, Guoqing, et al.
Published: (2024)
Similar Items
-
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
by: Shen, Hao, et al.
Published: (2024) -
Condition Weaving Meets Expert Modulation: Towards Universal and Controllable Image Generation
by: Zhang, Guoqing, et al.
Published: (2025) -
Object Retrieval for Visual Question Answering with Outside Knowledge
by: Kan, Shichao, et al.
Published: (2024) -
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
by: Kan, Shichao, et al.
Published: (2026) -
CFIS-YOLO: A Lightweight Multi-Scale Fusion Network for Edge-Deployable Wood Defect Detection
by: Kang, Jincheng, et al.
Published: (2025)