Saved in:
| Main Authors: | Jiang, Yuxuan, Yu, Chenwei, Lin, Zhi, Liu, Xiaolan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2507.19917 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Brain-like emergent properties in deep networks: impact of network architecture, datasets and training
by: Rajesh, Niranjan, et al.
Published: (2024)
by: Rajesh, Niranjan, et al.
Published: (2024)
Applications of deep generative models to DNA reaction kinetics and to cryogenic electron microscopy
by: Zhang, Chenwei
Published: (2026)
by: Zhang, Chenwei
Published: (2026)
Referential communication in heterogeneous communities of pre-trained visual deep networks
by: Mahaut, Matéo, et al.
Published: (2023)
by: Mahaut, Matéo, et al.
Published: (2023)
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
by: Han, Zhen, et al.
Published: (2024)
by: Han, Zhen, et al.
Published: (2024)
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition
by: Liu, Feng, et al.
Published: (2025)
by: Liu, Feng, et al.
Published: (2025)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
by: Jia, Chenwei, et al.
Published: (2026)
by: Jia, Chenwei, et al.
Published: (2026)
Edge-Enhanced Dilated Residual Attention Network for Multimodal Medical Image Fusion
by: Zhou, Meng, et al.
Published: (2024)
by: Zhou, Meng, et al.
Published: (2024)
Multimodal Medical Image Classification via Synergistic Learning Pre-training
by: Lin, Qinghua, et al.
Published: (2025)
by: Lin, Qinghua, et al.
Published: (2025)
GPT-NAS: Evolutionary Neural Architecture Search with the Generative Pre-Trained Model
by: Yu, Caiyang, et al.
Published: (2023)
by: Yu, Caiyang, et al.
Published: (2023)
Harnessing GPT-4V(ision) for Insurance: A Preliminary Exploration
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
INS-MMBench: A Comprehensive Benchmark for Evaluating LVLMs' Performance in Insurance
by: Lin, Chenwei, et al.
Published: (2024)
by: Lin, Chenwei, et al.
Published: (2024)
Improving action classification with brain-inspired deep networks
by: Aglinskas, Aidas, et al.
Published: (2025)
by: Aglinskas, Aidas, et al.
Published: (2025)
PH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative Dynamics
by: Luan, Xueyu, et al.
Published: (2026)
by: Luan, Xueyu, et al.
Published: (2026)
Precise localization of corneal reflections in eye images using deep learning trained on synthetic data
by: Byrne, Sean Anthony, et al.
Published: (2023)
by: Byrne, Sean Anthony, et al.
Published: (2023)
Configural processing as an optimized strategy for robust object recognition in neural networks
by: Jang, Hojin, et al.
Published: (2024)
by: Jang, Hojin, et al.
Published: (2024)
Dr. Seg: Revisiting GRPO Training for Visual Large Language Models through Perception-Oriented Design
by: Sun, Haoxiang, et al.
Published: (2026)
by: Sun, Haoxiang, et al.
Published: (2026)
Violence detection in videos using deep recurrent and convolutional neural networks
by: Traoré, Abdarahmane, et al.
Published: (2024)
by: Traoré, Abdarahmane, et al.
Published: (2024)
EfficientSAM3: Progressive Hierarchical Distillation for Video Concept Segmentation from SAM1, 2, and 3
by: Zeng, Chengxi, et al.
Published: (2025)
by: Zeng, Chengxi, et al.
Published: (2025)
Detecting and refurbishing ground truth errors during training of deep learning-based echocardiography segmentation models
by: Islam, Iman, et al.
Published: (2026)
by: Islam, Iman, et al.
Published: (2026)
Generative Pre-trained Autoregressive Diffusion Transformer
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
by: Luo, Jie, et al.
Published: (2025)
by: Luo, Jie, et al.
Published: (2025)
Downstream Transfer Attack: Adversarial Attacks on Downstream Models with Pre-trained Vision Transformers
by: Zheng, Weijie, et al.
Published: (2024)
by: Zheng, Weijie, et al.
Published: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Pack-PTQ: Advancing Post-training Quantization of Neural Networks by Pack-wise Reconstruction
by: Li, Changjun, et al.
Published: (2025)
by: Li, Changjun, et al.
Published: (2025)
CryoSAMU: Enhancing 3D Cryo-EM Density Maps of Protein Structures at Intermediate Resolution with Structure-Aware Multimodal U-Nets
by: Zhang, Chenwei, et al.
Published: (2025)
by: Zhang, Chenwei, et al.
Published: (2025)
Application of deep learning techniques in non-contrast computed tomography pulmonary angiogram for pulmonary embolism diagnosis
by: Ting, I-Hsien, et al.
Published: (2026)
by: Ting, I-Hsien, et al.
Published: (2026)
QVD: Post-training Quantization for Video Diffusion Models
by: Tian, Shilong, et al.
Published: (2024)
by: Tian, Shilong, et al.
Published: (2024)
CM3AE: A Unified RGB Frame and Event-Voxel/-Frame Pre-training Framework
by: Wu, Wentao, et al.
Published: (2025)
by: Wu, Wentao, et al.
Published: (2025)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI
by: Phitidis, Jesse, et al.
Published: (2026)
by: Phitidis, Jesse, et al.
Published: (2026)
High Performance Space Debris Tracking in Complex Skylight Backgrounds with a Large-Scale Dataset
by: Zhuang, Guohang, et al.
Published: (2025)
by: Zhuang, Guohang, et al.
Published: (2025)
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
by: Deng, Ziye, et al.
Published: (2025)
by: Deng, Ziye, et al.
Published: (2025)
Generative deep learning for foundational video translation in ultrasound
by: Tomic, Nikolina, et al.
Published: (2025)
by: Tomic, Nikolina, et al.
Published: (2025)
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
Prompting Lipschitz-constrained network for multiple-in-one sparse-view CT reconstruction
by: Shi, Baoshun, et al.
Published: (2025)
by: Shi, Baoshun, et al.
Published: (2025)
PhyPrompt: RL-based Prompt Refinement for Physically Plausible Text-to-Video Generation
by: Wu, Shang, et al.
Published: (2026)
by: Wu, Shang, et al.
Published: (2026)
A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training Models
by: Zheng, Haonan, et al.
Published: (2024)
by: Zheng, Haonan, et al.
Published: (2024)
On Data Synthesis and Post-training for Visual Abstract Reasoning
by: Zhu, Ke, et al.
Published: (2025)
by: Zhu, Ke, et al.
Published: (2025)
Can GPT-4o mini and Gemini 2.0 Flash Predict Fine-Grained Fashion Product Attributes? A Zero-Shot Analysis
by: Shukla, Shubham, et al.
Published: (2025)
by: Shukla, Shubham, et al.
Published: (2025)
Structure-aware World Model for Probe Guidance via Large-scale Self-supervised Pre-train
by: Jiang, Haojun, et al.
Published: (2024)
by: Jiang, Haojun, et al.
Published: (2024)
Similar Items
-
Brain-like emergent properties in deep networks: impact of network architecture, datasets and training
by: Rajesh, Niranjan, et al.
Published: (2024) -
Applications of deep generative models to DNA reaction kinetics and to cryogenic electron microscopy
by: Zhang, Chenwei
Published: (2026) -
Referential communication in heterogeneous communities of pre-trained visual deep networks
by: Mahaut, Matéo, et al.
Published: (2023) -
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
by: Han, Zhen, et al.
Published: (2024) -
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition
by: Liu, Feng, et al.
Published: (2025)