Guardado en:
| Autores principales: | Duan, Jiawei, Hu, Haibo, Ye, Qingqing, Sun, Xinyue |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2504.05618 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
por: Liu, Qinghua, et al.
Publicado: (2025)
por: Liu, Qinghua, et al.
Publicado: (2025)
Technical Report: Quantifying and Analyzing the Generalization Power of a DNN
por: He, Yuxuan, et al.
Publicado: (2025)
por: He, Yuxuan, et al.
Publicado: (2025)
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
por: Xu, Qinwu
Publicado: (2026)
por: Xu, Qinwu
Publicado: (2026)
CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation
por: Eslami, Mohammad, et al.
Publicado: (2026)
por: Eslami, Mohammad, et al.
Publicado: (2026)
Ovis2.5 Technical Report
por: Lu, Shiyin, et al.
Publicado: (2025)
por: Lu, Shiyin, et al.
Publicado: (2025)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
por: Celona, Luigi, et al.
Publicado: (2023)
por: Celona, Luigi, et al.
Publicado: (2023)
NeuroLip: An Event-driven Spatiotemporal Learning Framework for Cross-Scene Lip-Motion-based Visual Speaker Recognition
por: Yao, Junguang, et al.
Publicado: (2026)
por: Yao, Junguang, et al.
Publicado: (2026)
Assessing the Impact of Image Dataset Features on Privacy-Preserving Machine Learning
por: Lange, Lucas, et al.
Publicado: (2024)
por: Lange, Lucas, et al.
Publicado: (2024)
Physics-Guided Abnormal Trajectory Gap Detection
por: Sharma, Arun, et al.
Publicado: (2024)
por: Sharma, Arun, et al.
Publicado: (2024)
Unveiling the Pitfalls of Knowledge Editing for Large Language Models
por: Li, Zhoubo, et al.
Publicado: (2023)
por: Li, Zhoubo, et al.
Publicado: (2023)
Improving Diagnostic Performance on Small and Imbalanced Datasets Using Class-Based Input Image Composition
por: Azzeddine, Hlali, et al.
Publicado: (2025)
por: Azzeddine, Hlali, et al.
Publicado: (2025)
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
por: Yu, Wenhan, et al.
Publicado: (2025)
por: Yu, Wenhan, et al.
Publicado: (2025)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
por: Lin, Ling, et al.
Publicado: (2026)
por: Lin, Ling, et al.
Publicado: (2026)
A large-scale multicenter breast cancer DCE-MRI benchmark dataset with expert segmentations
por: Garrucho, Lidia, et al.
Publicado: (2024)
por: Garrucho, Lidia, et al.
Publicado: (2024)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
por: Dai, Josef, et al.
Publicado: (2024)
por: Dai, Josef, et al.
Publicado: (2024)
3D Primitives are a Spatial Language for VLMs
por: Liu, Junze, et al.
Publicado: (2026)
por: Liu, Junze, et al.
Publicado: (2026)
Probabilistic Kernel Function for Fast Angle Testing
por: Lu, Kejing, et al.
Publicado: (2025)
por: Lu, Kejing, et al.
Publicado: (2025)
Probabilistic Routing for Graph-Based Approximate Nearest Neighbor Search
por: Lu, Kejing, et al.
Publicado: (2024)
por: Lu, Kejing, et al.
Publicado: (2024)
MotionCFG: Boosting Motion Dynamics via Stochastic Concept Perturbation
por: Kim, Byungjun, et al.
Publicado: (2026)
por: Kim, Byungjun, et al.
Publicado: (2026)
Divide, Weight, and Route: Difficulty-Aware Optimization with Dynamic Expert Fusion for Long-tailed Recognition
por: Wei, Xiaolei, et al.
Publicado: (2025)
por: Wei, Xiaolei, et al.
Publicado: (2025)
Ovis-U1 Technical Report
por: Wang, Guo-Hua, et al.
Publicado: (2025)
por: Wang, Guo-Hua, et al.
Publicado: (2025)
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
por: Yuan, Ruicheng, et al.
Publicado: (2026)
por: Yuan, Ruicheng, et al.
Publicado: (2026)
Geometric 4D Stitching for Grounded 4D Generation
por: Park, Sunwoo, et al.
Publicado: (2026)
por: Park, Sunwoo, et al.
Publicado: (2026)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
por: Lab, Shanghai AI, et al.
Publicado: (2025)
por: Lab, Shanghai AI, et al.
Publicado: (2025)
UI-Venus-1.5 Technical Report
por: Venus Team, et al.
Publicado: (2026)
por: Venus Team, et al.
Publicado: (2026)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
por: Zhang, Xinwei, et al.
Publicado: (2026)
por: Zhang, Xinwei, et al.
Publicado: (2026)
Where and How to Perturb: On the Design of Perturbation Guidance in Diffusion and Flow Models
por: Ahn, Donghoon, et al.
Publicado: (2025)
por: Ahn, Donghoon, et al.
Publicado: (2025)
Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization
por: Guan, Jiwei, et al.
Publicado: (2026)
por: Guan, Jiwei, et al.
Publicado: (2026)
On-Demand Multi-Task Sparsity for Efficient Large-Model Deployment on Edge Devices
por: Huang, Lianming, et al.
Publicado: (2025)
por: Huang, Lianming, et al.
Publicado: (2025)
HunyuanOCR Technical Report
por: Hunyuan Vision Team, et al.
Publicado: (2025)
por: Hunyuan Vision Team, et al.
Publicado: (2025)
Seed1.5-VL Technical Report
por: Guo, Dong, et al.
Publicado: (2025)
por: Guo, Dong, et al.
Publicado: (2025)
Qwen3-VL Technical Report
por: Bai, Shuai, et al.
Publicado: (2025)
por: Bai, Shuai, et al.
Publicado: (2025)
Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence
por: Chen, Ruoxi, et al.
Publicado: (2021)
por: Chen, Ruoxi, et al.
Publicado: (2021)
DP$^2$O-SR: Direct Perceptual Preference Optimization for Real-World Image Super-Resolution
por: Wu, Rongyuan, et al.
Publicado: (2025)
por: Wu, Rongyuan, et al.
Publicado: (2025)
SODIUM: From Open Web Data to Queryable Databases
por: Hu, Chuxuan, et al.
Publicado: (2026)
por: Hu, Chuxuan, et al.
Publicado: (2026)
CMRxRecon2024: A Multi-Modality, Multi-View K-Space Dataset Boosting Universal Machine Learning for Accelerated Cardiac MRI
por: Wang, Zi, et al.
Publicado: (2024)
por: Wang, Zi, et al.
Publicado: (2024)
Perturbing the Gradient for Alleviating Meta Overfitting
por: Gogoi, Manas, et al.
Publicado: (2024)
por: Gogoi, Manas, et al.
Publicado: (2024)
H2OVL-Mississippi Vision Language Models Technical Report
por: Galib, Shaikat, et al.
Publicado: (2024)
por: Galib, Shaikat, et al.
Publicado: (2024)
Ovis-Image Technical Report
por: Wang, Guo-Hua, et al.
Publicado: (2025)
por: Wang, Guo-Hua, et al.
Publicado: (2025)
GR-3 Technical Report
por: Cheang, Chilam, et al.
Publicado: (2025)
por: Cheang, Chilam, et al.
Publicado: (2025)
Ejemplares similares
-
MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
por: Liu, Qinghua, et al.
Publicado: (2025) -
Technical Report: Quantifying and Analyzing the Generalization Power of a DNN
por: He, Yuxuan, et al.
Publicado: (2025) -
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
por: Xu, Qinwu
Publicado: (2026) -
CataractSAM-2: A Domain-Adapted Model for Anterior Segment Surgery Segmentation and Scalable Ground-Truth Annotation
por: Eslami, Mohammad, et al.
Publicado: (2026) -
Ovis2.5 Technical Report
por: Lu, Shiyin, et al.
Publicado: (2025)