PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Dexiang, Che, Zhenning, Zhang, Haijun, Zhou, Dongliang, Zhang, Zhao, Han, Yahong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility Boosting
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
di: Liu, Che, et al.
Pubblicazione: (2025)
di: Liu, Che, et al.
Pubblicazione: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
di: Han, ZhaoYang, et al.
Pubblicazione: (2025)
di: Han, ZhaoYang, et al.
Pubblicazione: (2025)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
di: Wang, Sen, et al.
Pubblicazione: (2025)
di: Wang, Sen, et al.
Pubblicazione: (2025)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
di: Wu, Haoning, et al.
Pubblicazione: (2023)
di: Wu, Haoning, et al.
Pubblicazione: (2023)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
di: Ukai, Mahiro, et al.
Pubblicazione: (2025)
di: Ukai, Mahiro, et al.
Pubblicazione: (2025)
LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
di: Yu, Yanrui, et al.
Pubblicazione: (2025)
PaveSAM Segment Anything for Pavement Distress
di: Owor, Neema Jakisa, et al.
Pubblicazione: (2024)
di: Owor, Neema Jakisa, et al.
Pubblicazione: (2024)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
di: Hao, Jing, et al.
Pubblicazione: (2025)
di: Hao, Jing, et al.
Pubblicazione: (2025)
VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics
di: Yin, Zhiyu, et al.
Pubblicazione: (2026)
di: Yin, Zhiyu, et al.
Pubblicazione: (2026)
DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection
di: Zhao, Kangran, et al.
Pubblicazione: (2025)
di: Zhao, Kangran, et al.
Pubblicazione: (2025)
Enhancing Interactive Image Retrieval With Query Rewriting Using Large Language Models and Vision Language Models
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
di: Zhu, Hongyi, et al.
Pubblicazione: (2024)
Hierarchical Refinement of Universal Multimodal Attacks on Vision-Language Models
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2026)
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2026)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
di: Hao, Bowen, et al.
Pubblicazione: (2025)
di: Hao, Bowen, et al.
Pubblicazione: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
di: Du, Mengfei, et al.
Pubblicazione: (2024)
di: Du, Mengfei, et al.
Pubblicazione: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
di: Sun, Yanpeng, et al.
Pubblicazione: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
di: Feng, X., et al.
Pubblicazione: (2024)
di: Feng, X., et al.
Pubblicazione: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
di: Wang, Xinran, et al.
Pubblicazione: (2026)
di: Wang, Xinran, et al.
Pubblicazione: (2026)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
di: Zhang, Rui, et al.
Pubblicazione: (2024)
di: Zhang, Rui, et al.
Pubblicazione: (2024)
Selective Vision-Language Subspace Projection for Few-shot CLIP
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
di: Zhu, Xingyu, et al.
Pubblicazione: (2024)
A Preprocessing Framework for Video Machine Vision under Compression
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
PaveSync: A Unified and Comprehensive Dataset for Pavement Distress Analysis and Classification
di: Kyem, Blessing Agyei, et al.
Pubblicazione: (2025)
di: Kyem, Blessing Agyei, et al.
Pubblicazione: (2025)
FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
di: Wen, Haokun, et al.
Pubblicazione: (2026)
di: Wen, Haokun, et al.
Pubblicazione: (2026)
MAO: Efficient Model-Agnostic Optimization of Prompt Tuning for Vision-Language Models
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
di: Liu, Yating, et al.
Pubblicazione: (2025)
di: Liu, Yating, et al.
Pubblicazione: (2025)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
di: Fu, Jiyuan, et al.
Pubblicazione: (2024)
di: Fu, Jiyuan, et al.
Pubblicazione: (2024)
POINTS1.5: Building a Vision-Language Model towards Real World Applications
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
di: Tian, Huilin, et al.
Pubblicazione: (2024)
di: Tian, Huilin, et al.
Pubblicazione: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
di: Zhou, Yuchen, et al.
Pubblicazione: (2025)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
di: Xie, Jingjing, et al.
Pubblicazione: (2024)
Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
di: Ma, Jingtian, et al.
Pubblicazione: (2025)
di: Ma, Jingtian, et al.
Pubblicazione: (2025)
TAVGBench: Benchmarking Text to Audible-Video Generation
di: Mao, Yuxin, et al.
Pubblicazione: (2024)
di: Mao, Yuxin, et al.
Pubblicazione: (2024)
DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
di: Yang, Wei, et al.
Pubblicazione: (2025)
di: Yang, Wei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility Boosting
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
BC-GAN: A Generative Adversarial Network for Synthesizing a Batch of Collocated Clothing
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
COutfitGAN: Learning to Synthesize Compatible Outfits Supervised by Silhouette Masks and Fashion Styles
di: Zhou, Dongliang, et al.
Pubblicazione: (2025) -
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
di: Liu, Che, et al.
Pubblicazione: (2025) -
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
di: Han, ZhaoYang, et al.
Pubblicazione: (2025)