Saved in:
| Main Authors: | Liu, Jiaming, Kong, Linghe, Wu, Yue, Gong, Maoguo, Li, Hao, Miao, Qiguang, Ma, Wenping, Qin, Can |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.17547 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution Network
by: Wang, Zhaoyang, et al.
Published: (2024)
by: Wang, Zhaoyang, et al.
Published: (2024)
Multi-scale Information Sharing and Selection Network with Boundary Attention for Polyp Segmentation
by: Kang, Xiaolu, et al.
Published: (2024)
by: Kang, Xiaolu, et al.
Published: (2024)
CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
by: Qin, Guangshuo, et al.
Published: (2026)
by: Qin, Guangshuo, et al.
Published: (2026)
PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
by: Liu, Kang, et al.
Published: (2025)
by: Liu, Kang, et al.
Published: (2025)
Fose: Fusion of One-Step Diffusion and End-to-End Network for Pansharpening
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing
by: Feng, Guanwen, et al.
Published: (2025)
by: Feng, Guanwen, et al.
Published: (2025)
Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report Generation
by: Liu, Kang, et al.
Published: (2025)
by: Liu, Kang, et al.
Published: (2025)
Generative Adversarial Patches for Physical Attacks on Cross-Modal Pedestrian Re-Identification
by: Su, Yue, et al.
Published: (2024)
by: Su, Yue, et al.
Published: (2024)
Seeing Like Radiologists: Context- and Gaze-Guided Vision-Language Pretraining for Chest X-rays
by: Liu, Kang, et al.
Published: (2026)
by: Liu, Kang, et al.
Published: (2026)
EVOKE: Elevating Chest X-ray Report Generation via Multi-View Contrastive Learning and Patient-Specific Knowledge
by: Miao, Qiguang, et al.
Published: (2024)
by: Miao, Qiguang, et al.
Published: (2024)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
by: Qin, Libo, et al.
Published: (2024)
by: Qin, Libo, et al.
Published: (2024)
Dog-IQA: Standard-guided Zero-shot MLLM for Mix-grained Image Quality Assessment
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
AdaSVD: Adaptive Singular Value Decomposition for Large Language Models
by: Li, Zhiteng, et al.
Published: (2025)
by: Li, Zhiteng, et al.
Published: (2025)
Improving SAM for Camouflaged Object Detection via Dual Stream Adapters
by: Liu, Jiaming, et al.
Published: (2025)
by: Liu, Jiaming, et al.
Published: (2025)
RBF++: Quantifying and Optimizing Reasoning Boundaries across Measurable and Unmeasurable Capabilities for Chain-of-Thought Reasoning
by: Chen, Qiguang, et al.
Published: (2025)
by: Chen, Qiguang, et al.
Published: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
by: Zhang, Yongheng, et al.
Published: (2025)
by: Zhang, Yongheng, et al.
Published: (2025)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
IDER: IDempotent Experience Replay for Reliable Continual Learning
by: Liu, Zhanwang, et al.
Published: (2026)
by: Liu, Zhanwang, et al.
Published: (2026)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
CL-CaGAN: Capsule differential adversarial continuous learning for cross-domain hyperspectral anomaly detection
by: Wang, Jianing, et al.
Published: (2025)
by: Wang, Jianing, et al.
Published: (2025)
Factual Serialization Enhancement: A Key Innovation for Chest X-ray Report Generation
by: Liu, Kang, et al.
Published: (2024)
by: Liu, Kang, et al.
Published: (2024)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
Masked Latent Transformer with the Random Masking Ratio to Advance the Diagnosis of Dental Fluorosis
by: Wu, Yun, et al.
Published: (2024)
by: Wu, Yun, et al.
Published: (2024)
MacDiff: Unified Skeleton Modeling with Masked Conditional Diffusion
by: Wu, Lehong, et al.
Published: (2024)
by: Wu, Lehong, et al.
Published: (2024)
Instruction-Guided Visual Masking
by: Zheng, Jinliang, et al.
Published: (2024)
by: Zheng, Jinliang, et al.
Published: (2024)
EFTViT: Efficient Federated Training of Vision Transformers with Masked Images on Resource-Constrained Clients
by: Wu, Meihan, et al.
Published: (2024)
by: Wu, Meihan, et al.
Published: (2024)
2DQuant: Low-bit Post-Training Quantization for Image Super-Resolution
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
by: Chen, Qiguang, et al.
Published: (2024)
by: Chen, Qiguang, et al.
Published: (2024)
Unveiling and Mitigating Generalized Biases of DNNs through the Intrinsic Dimensions of Perceptual Manifolds
by: Ma, Yanbiao, et al.
Published: (2024)
by: Ma, Yanbiao, et al.
Published: (2024)
Masked Generative Extractor for Synergistic Representation and 3D Generation of Point Clouds
by: Zeng, Hongliang, et al.
Published: (2024)
by: Zeng, Hongliang, et al.
Published: (2024)
EADReg: Probabilistic Correspondence Generation with Efficient Autoregressive Diffusion Model for Outdoor Point Cloud Registration
by: Gong, Linrui, et al.
Published: (2024)
by: Gong, Linrui, et al.
Published: (2024)
Point-DETR3D: Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection
by: Gao, Hongzhi, et al.
Published: (2024)
by: Gao, Hongzhi, et al.
Published: (2024)
Open-Vocabulary Segmentation with Unpaired Mask-Text Supervision
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
Point Cloud Self-supervised Learning via 3D to Multi-view Masked Learner
by: Chen, Zhimin, et al.
Published: (2023)
by: Chen, Zhimin, et al.
Published: (2023)
Predicting and Enhancing the Fairness of DNNs with the Curvature of Perceptual Manifolds
by: Ma, Yanbiao, et al.
Published: (2023)
by: Ma, Yanbiao, et al.
Published: (2023)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
MASSeg : 2nd Technical Report for 4th PVUW MOSE Track
by: Cao, Xuqiang, et al.
Published: (2025)
by: Cao, Xuqiang, et al.
Published: (2025)
Similar Items
-
Enhancing Hyperspectral Images via Diffusion Model and Group-Autoencoder Super-resolution Network
by: Wang, Zhaoyang, et al.
Published: (2024) -
Multi-scale Information Sharing and Selection Network with Boundary Attention for Polyp Segmentation
by: Kang, Xiaolu, et al.
Published: (2024) -
CLQ: Cross-Layer Guided Orthogonal-based Quantization for Diffusion Transformers
by: Liu, Kai, et al.
Published: (2025) -
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
by: Qin, Guangshuo, et al.
Published: (2026) -
PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report Generation
by: Liu, Kang, et al.
Published: (2025)