Saved in:
| Main Authors: | Wang, Guangrun, Li, Changlin, Yuan, Liuchun, Peng, Jiefeng, Xian, Xiaoyu, Liang, Xiaodan, Chang, Xiaojun, Lin, Liang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.01326 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning
by: Li, Changlin, et al.
Published: (2024)
by: Li, Changlin, et al.
Published: (2024)
MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation
by: Cai, Kaixin, et al.
Published: (2023)
by: Cai, Kaixin, et al.
Published: (2023)
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning
by: Chen, Linsheng, et al.
Published: (2024)
by: Chen, Linsheng, et al.
Published: (2024)
SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NAS
by: Peng, Yameng, et al.
Published: (2024)
by: Peng, Yameng, et al.
Published: (2024)
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
by: Tang, Tao, et al.
Published: (2024)
by: Tang, Tao, et al.
Published: (2024)
Contrastive Learning with Counterfactual Explanations for Radiology Report Generation
by: Li, Mingjie, et al.
Published: (2024)
by: Li, Mingjie, et al.
Published: (2024)
Making Large Language Models Better Planners with Reasoning-Decision Alignment
by: Huang, Zhijian, et al.
Published: (2024)
by: Huang, Zhijian, et al.
Published: (2024)
GS: Generative Segmentation via Label Diffusion
by: Chen, Yuhao, et al.
Published: (2025)
by: Chen, Yuhao, et al.
Published: (2025)
HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment
by: Jiang, Zutao, et al.
Published: (2023)
by: Jiang, Zutao, et al.
Published: (2023)
MLP Can Be A Good Transformer Learner
by: Lin, Sihao, et al.
Published: (2024)
by: Lin, Sihao, et al.
Published: (2024)
Knowledge Distillation via the Target-aware Transformer
by: Lin, Sihao, et al.
Published: (2022)
by: Lin, Sihao, et al.
Published: (2022)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
by: Chen, Jianqi, et al.
Published: (2024)
by: Chen, Jianqi, et al.
Published: (2024)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
One Model for All: Unified Try-On and Try-Off in Any Pose via LLM-Inspired Bidirectional Tweedie Diffusion
by: Liu, Jinxi, et al.
Published: (2025)
by: Liu, Jinxi, et al.
Published: (2025)
Predicting Genetic Mutation from Whole Slide Images via Biomedical-Linguistic Knowledge Enhanced Multi-label Classification
by: Huang, Gexin, et al.
Published: (2024)
by: Huang, Gexin, et al.
Published: (2024)
Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression Manipulation
by: Chen, Tianshui, et al.
Published: (2026)
by: Chen, Tianshui, et al.
Published: (2026)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
NavCoT: Boosting LLM-Based Vision-and-Language Navigation via Learning Disentangled Reasoning
by: Lin, Bingqian, et al.
Published: (2024)
by: Lin, Bingqian, et al.
Published: (2024)
CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
by: Hao, Haihong, et al.
Published: (2025)
by: Hao, Haihong, et al.
Published: (2025)
ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving Augmentation
by: Bian, Siyuan, et al.
Published: (2024)
by: Bian, Siyuan, et al.
Published: (2024)
DFVO: Learning Darkness-free Visible and Infrared Image Disentanglement and Fusion All at Once
by: Zhou, Qi, et al.
Published: (2025)
by: Zhou, Qi, et al.
Published: (2025)
MirrorDiffusion: Stabilizing Diffusion Process in Zero-shot Image Translation by Prompts Redescription and Beyond
by: Lin, Yupei, et al.
Published: (2024)
by: Lin, Yupei, et al.
Published: (2024)
Flexiffusion: Training-Free Segment-Wise Neural Architecture Search for Efficient Diffusion Models
by: Huang, Hongtao, et al.
Published: (2025)
by: Huang, Hongtao, et al.
Published: (2025)
ActionSink: Toward Precise Robot Manipulation with Dynamic Integration of Action Flow
by: Guo, Shanshan, et al.
Published: (2025)
by: Guo, Shanshan, et al.
Published: (2025)
WildVidFit: Video Virtual Try-On in the Wild via Image-Based Controlled Diffusion Models
by: He, Zijian, et al.
Published: (2024)
by: He, Zijian, et al.
Published: (2024)
TopoNAS: Boosting Search Efficiency of Gradient-based NAS via Topological Simplification
by: Zhao, Danpei, et al.
Published: (2024)
by: Zhao, Danpei, et al.
Published: (2024)
Multi-View People Detection in Large Scenes via Supervised View-Wise Contribution Weighting
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Aquarius: A Family of Industry-Level Video Generation Models for Marketing Scenarios
by: Shi, Huafeng, et al.
Published: (2025)
by: Shi, Huafeng, et al.
Published: (2025)
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
by: Ma, Liang, et al.
Published: (2025)
by: Ma, Liang, et al.
Published: (2025)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
by: Han, Mingfei, et al.
Published: (2026)
by: Han, Mingfei, et al.
Published: (2026)
Geometry aware 3D generation from in-the-wild images in ImageNet
by: Shen, Qijia, et al.
Published: (2024)
by: Shen, Qijia, et al.
Published: (2024)
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
by: Han, Mingfei, et al.
Published: (2024)
by: Han, Mingfei, et al.
Published: (2024)
AnimateZoo: Zero-shot Video Generation of Cross-Species Animation via Subject Alignment
by: Xu, Yuanfeng, et al.
Published: (2024)
by: Xu, Yuanfeng, et al.
Published: (2024)
VTON 360: High-Fidelity Virtual Try-On from Any Viewing Direction
by: He, Zijian, et al.
Published: (2025)
by: He, Zijian, et al.
Published: (2025)
Efficient Training for Human Video Generation with Entropy-Guided Prioritized Progressive Learning
by: Li, Changlin, et al.
Published: (2025)
by: Li, Changlin, et al.
Published: (2025)
SIRST-5K: Exploring Massive Negatives Synthesis with Self-supervised Learning for Robust Infrared Small Target Detection
by: Lu, Yahao, et al.
Published: (2024)
by: Lu, Yahao, et al.
Published: (2024)
Similar Items
-
Efficient Training of Large Vision Models via Advanced Automated Progressive Learning
by: Li, Changlin, et al.
Published: (2024) -
MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic Segmentation
by: Cai, Kaixin, et al.
Published: (2023) -
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
by: Li, Weiqi, et al.
Published: (2025) -
NeRF-VPT: Learning Novel View Representations with Neural Radiance Fields via View Prompt Tuning
by: Chen, Linsheng, et al.
Published: (2024) -
SWAP-NAS: Sample-Wise Activation Patterns for Ultra-fast NAS
by: Peng, Yameng, et al.
Published: (2024)