Enregistré dans:
| Auteurs principaux: | Du, Tianxiang, He, Hulingxiao, Peng, Yuxin |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2602.23980 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AesFormer: Transform Everyday Photos into Beautiful Memories
par: Du, Tianxiang, et autres
Publié: (2026)
par: Du, Tianxiang, et autres
Publié: (2026)
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
par: He, Hulingxiao, et autres
Publié: (2026)
par: He, Hulingxiao, et autres
Publié: (2026)
CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting
par: He, Hulingxiao, et autres
Publié: (2024)
par: He, Hulingxiao, et autres
Publié: (2024)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
par: He, Hulingxiao, et autres
Publié: (2026)
par: He, Hulingxiao, et autres
Publié: (2026)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
par: He, Hulingxiao, et autres
Publié: (2025)
par: He, Hulingxiao, et autres
Publié: (2025)
AesBench: An Expert Benchmark for Multimodal Large Language Models on Image Aesthetics Perception
par: Huang, Yipo, et autres
Publié: (2024)
par: Huang, Yipo, et autres
Publié: (2024)
AesCrop: Aesthetic-driven Cropping Guided by Composition
par: Wong, Yen-Hong, et autres
Publié: (2025)
par: Wong, Yen-Hong, et autres
Publié: (2025)
VideoAesBench: Benchmarking the Video Aesthetics Perception Capabilities of Large Multimodal Models
par: Li, Yunhao, et autres
Publié: (2026)
par: Li, Yunhao, et autres
Publié: (2026)
ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
par: Zhang, Ke, et autres
Publié: (2025)
par: Zhang, Ke, et autres
Publié: (2025)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
par: Qi, Daiqing, et autres
Publié: (2025)
par: Qi, Daiqing, et autres
Publié: (2025)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
par: Li, Mingxing, et autres
Publié: (2025)
par: Li, Mingxing, et autres
Publié: (2025)
VIP: Versatile Image Outpainting Empowered by Multimodal Large Language Model
par: Yang, Jinze, et autres
Publié: (2024)
par: Yang, Jinze, et autres
Publié: (2024)
EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models
par: Peng, Xiaomeng, et autres
Publié: (2026)
par: Peng, Xiaomeng, et autres
Publié: (2026)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
par: Yu, Hong-Tao, et autres
Publié: (2025)
par: Yu, Hong-Tao, et autres
Publié: (2025)
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
par: Qiu, Wenmo, et autres
Publié: (2024)
par: Qiu, Wenmo, et autres
Publié: (2024)
A Multimodal Benchmark Dataset and Model for Crop Disease Diagnosis
par: Liu, Xiang, et autres
Publié: (2025)
par: Liu, Xiang, et autres
Publié: (2025)
Diffusion-based Aesthetic QR Code Generation via Scanning-Robust Perceptual Guidance
par: Liao, Jia-Wei, et autres
Publié: (2024)
par: Liao, Jia-Wei, et autres
Publié: (2024)
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
par: Fu, Chaoyou, et autres
Publié: (2023)
par: Fu, Chaoyou, et autres
Publié: (2023)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
par: Zhu, Wenxin, et autres
Publié: (2025)
par: Zhu, Wenxin, et autres
Publié: (2025)
Empowering Segmentation Ability to Multi-modal Large Language Models
par: Yang, Yuqi, et autres
Publié: (2024)
par: Yang, Yuqi, et autres
Publié: (2024)
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
par: Liu, Yuansen, et autres
Publié: (2025)
par: Liu, Yuansen, et autres
Publié: (2025)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
par: Zhu, Muzhi, et autres
Publié: (2025)
par: Zhu, Muzhi, et autres
Publié: (2025)
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
par: He, Yiran, et autres
Publié: (2025)
par: He, Yiran, et autres
Publié: (2025)
Image Aesthetic Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
par: Hu, Zhiyuan, et autres
Publié: (2025)
par: Hu, Zhiyuan, et autres
Publié: (2025)
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
par: Peng, Tianfan, et autres
Publié: (2025)
par: Peng, Tianfan, et autres
Publié: (2025)
Empowering Large Language Models with 3D Situation Awareness
par: Yuan, Zhihao, et autres
Publié: (2025)
par: Yuan, Zhihao, et autres
Publié: (2025)
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
par: Li, Jiansheng, et autres
Publié: (2025)
par: Li, Jiansheng, et autres
Publié: (2025)
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
par: Ren, Xiyu, et autres
Publié: (2026)
par: Ren, Xiyu, et autres
Publié: (2026)
ProactiveBench: Benchmarking Proactiveness in Multimodal Large Language Models
par: De Min, Thomas, et autres
Publié: (2026)
par: De Min, Thomas, et autres
Publié: (2026)
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
par: Xu, Zhaopan, et autres
Publié: (2025)
par: Xu, Zhaopan, et autres
Publié: (2025)
Diffusion-based Facial Aesthetics Enhancement with 3D Structure Guidance
par: Li, Lisha, et autres
Publié: (2025)
par: Li, Lisha, et autres
Publié: (2025)
SHIELD : An Evaluation Benchmark for Face Spoofing and Forgery Detection with Multimodal Large Language Models
par: Shi, Yichen, et autres
Publié: (2024)
par: Shi, Yichen, et autres
Publié: (2024)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
par: Xu, Mingjie, et autres
Publié: (2025)
par: Xu, Mingjie, et autres
Publié: (2025)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
par: Jin, Zhe, et autres
Publié: (2025)
par: Jin, Zhe, et autres
Publié: (2025)
PromptLNet: Region-Adaptive Aesthetic Enhancement via Prompt Guidance in Low-Light Enhancement Net
par: Yin, Jun, et autres
Publié: (2025)
par: Yin, Jun, et autres
Publié: (2025)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
par: Ye, Qinghao, et autres
Publié: (2023)
par: Ye, Qinghao, et autres
Publié: (2023)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
par: Luo, Gen, et autres
Publié: (2024)
par: Luo, Gen, et autres
Publié: (2024)
TiFRe: Text-guided Video Frame Reduction for Efficient Video Multi-modal Large Language Models
par: Zheng, Xiangtian, et autres
Publié: (2026)
par: Zheng, Xiangtian, et autres
Publié: (2026)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
par: Luo, Fuwen, et autres
Publié: (2024)
par: Luo, Fuwen, et autres
Publié: (2024)
Dynamic Resolution Guidance for Facial Expression Recognition
par: Wang, Songpan, et autres
Publié: (2024)
par: Wang, Songpan, et autres
Publié: (2024)
Documents similaires
-
AesFormer: Transform Everyday Photos into Beautiful Memories
par: Du, Tianxiang, et autres
Publié: (2026) -
Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
par: He, Hulingxiao, et autres
Publié: (2026) -
CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting
par: He, Hulingxiao, et autres
Publié: (2024) -
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
par: He, Hulingxiao, et autres
Publié: (2026) -
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
par: He, Hulingxiao, et autres
Publié: (2025)