Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yulong, Liang, Tianyi, Huang, Xinyue, Cui, Erfei, Wang, Guoqing, Guo, Xu, Li, Chenhui, Liu, Gongshen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Harnessing Self-Supervised Features for Art Classification
por: Melis, Federico, et al.
Publicado: (2026)
por: Melis, Federico, et al.
Publicado: (2026)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
por: Wang, Tianyi, et al.
Publicado: (2025)
por: Wang, Tianyi, et al.
Publicado: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
por: Yang, Jianxuan, et al.
Publicado: (2025)
por: Yang, Jianxuan, et al.
Publicado: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
por: Tu, Jinzhe, et al.
Publicado: (2026)
por: Tu, Jinzhe, et al.
Publicado: (2026)
Self-supervised Photographic Image Layout Representation Learning
por: Zhao, Zhaoran, et al.
Publicado: (2024)
por: Zhao, Zhaoran, et al.
Publicado: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
por: Zhang, Zhenxing, et al.
Publicado: (2024)
por: Zhang, Zhenxing, et al.
Publicado: (2024)
Copy-Move Forgery Detection and Question Answering for Remote Sensing Image
por: Zhang, Ze, et al.
Publicado: (2024)
por: Zhang, Ze, et al.
Publicado: (2024)
Self-similarity Prior Distillation for Unsupervised Remote Physiological Measurement
por: Zhang, Xinyu, et al.
Publicado: (2023)
por: Zhang, Xinyu, et al.
Publicado: (2023)
HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
por: Zhang, Zijian, et al.
Publicado: (2025)
por: Zhang, Zijian, et al.
Publicado: (2025)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
por: Shan, Ziyu, et al.
Publicado: (2024)
por: Shan, Ziyu, et al.
Publicado: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
por: Liang, Zhengyang, et al.
Publicado: (2024)
por: Liang, Zhengyang, et al.
Publicado: (2024)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
por: Zhu, Rui, et al.
Publicado: (2024)
por: Zhu, Rui, et al.
Publicado: (2024)
Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy Labels
por: Pu, Ruitao, et al.
Publicado: (2025)
por: Pu, Ruitao, et al.
Publicado: (2025)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
por: Wu, Jianyu, et al.
Publicado: (2025)
por: Wu, Jianyu, et al.
Publicado: (2025)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
por: Lu, Renjie, et al.
Publicado: (2026)
por: Lu, Renjie, et al.
Publicado: (2026)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
por: Masui, Kento, et al.
Publicado: (2024)
por: Masui, Kento, et al.
Publicado: (2024)
DPC-VQA: Decoupling Quality Perception and Residual Calibration for Video Quality Assessment
por: Li, Xinyue, et al.
Publicado: (2026)
por: Li, Xinyue, et al.
Publicado: (2026)
MVBIND: Self-Supervised Music Recommendation For Videos Via Embedding Space Binding
por: Teng, Jiajie, et al.
Publicado: (2024)
por: Teng, Jiajie, et al.
Publicado: (2024)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
por: Wu, Yiming, et al.
Publicado: (2024)
por: Wu, Yiming, et al.
Publicado: (2024)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
por: Zhang, Wang, et al.
Publicado: (2024)
por: Zhang, Wang, et al.
Publicado: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
por: Fang, Ziye, et al.
Publicado: (2023)
por: Fang, Ziye, et al.
Publicado: (2023)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
por: Liu, Xiaofeng, et al.
Publicado: (2023)
por: Liu, Xiaofeng, et al.
Publicado: (2023)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
por: Dai, Guangyu, et al.
Publicado: (2025)
por: Dai, Guangyu, et al.
Publicado: (2025)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
por: Sun, Yanpeng, et al.
Publicado: (2024)
por: Sun, Yanpeng, et al.
Publicado: (2024)
DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
por: Yang, Wei, et al.
Publicado: (2025)
por: Yang, Wei, et al.
Publicado: (2025)
LookupForensics: A Large-Scale Multi-Task Dataset for Multi-Phase Image-Based Fact Verification
por: Cui, Shuhan, et al.
Publicado: (2024)
por: Cui, Shuhan, et al.
Publicado: (2024)
Context Guided Transformer Entropy Modeling for Video Compression
por: Tong, Junlong, et al.
Publicado: (2025)
por: Tong, Junlong, et al.
Publicado: (2025)
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
por: Wang, Zihua, et al.
Publicado: (2025)
por: Wang, Zihua, et al.
Publicado: (2025)
MM-Point: Multi-View Information-Enhanced Multi-Modal Self-Supervised 3D Point Cloud Understanding
por: Yu, Hai-Tao, et al.
Publicado: (2024)
por: Yu, Hai-Tao, et al.
Publicado: (2024)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
por: Guo, Xinyue, et al.
Publicado: (2025)
por: Guo, Xinyue, et al.
Publicado: (2025)
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
por: Chen, Rongjun, et al.
Publicado: (2025)
por: Chen, Rongjun, et al.
Publicado: (2025)
Spotlighting Partially Visible Cinematic Language for Video-to-Audio Generation via Self-distillation
por: Huang, Feizhen, et al.
Publicado: (2025)
por: Huang, Feizhen, et al.
Publicado: (2025)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
por: Liang, Hao, et al.
Publicado: (2024)
por: Liang, Hao, et al.
Publicado: (2024)
Quantifying and Enhancing Multi-modal Robustness with Modality Preference
por: Yang, Zequn, et al.
Publicado: (2024)
por: Yang, Zequn, et al.
Publicado: (2024)
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
por: Pan, Yiming, et al.
Publicado: (2026)
por: Pan, Yiming, et al.
Publicado: (2026)
GTPBD-MM: A Global Terraced Parcel and Boundary Dataset with Multi-Modality
por: Zhang, Zhiwei, et al.
Publicado: (2026)
por: Zhang, Zhiwei, et al.
Publicado: (2026)
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation
por: Zhang, Yingying, et al.
Publicado: (2024)
por: Zhang, Yingying, et al.
Publicado: (2024)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
por: Jiang, Xun, et al.
Publicado: (2026)
por: Jiang, Xun, et al.
Publicado: (2026)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
por: Zheng, Junjie, et al.
Publicado: (2025)
por: Zheng, Junjie, et al.
Publicado: (2025)
Ejemplares similares
-
Harnessing Self-Supervised Features for Art Classification
por: Melis, Federico, et al.
Publicado: (2026) -
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
por: Wang, Tianyi, et al.
Publicado: (2025) -
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
por: Yang, Jianxuan, et al.
Publicado: (2025) -
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
por: Tu, Jinzhe, et al.
Publicado: (2026) -
Self-supervised Photographic Image Layout Representation Learning
por: Zhao, Zhaoran, et al.
Publicado: (2024)