VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Xingyu, Wang, Jinpeng, Zhang, Yi-Fan, Yang, Yankai, Long, Yancheng, Fan, Yiyang, Zheng, Xuanyu, Fan, Haonan, Jiang, Kaiyu, Zhang, Tianke, Liu, Changyi, Wen, Bin, Yang, Fan, Gao, Tingting, Li, Han, Yuan, Chun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
di: Long, Yancheng, et al.
Pubblicazione: (2026)
di: Long, Yancheng, et al.
Pubblicazione: (2026)
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
di: Yang, Yankai, et al.
Pubblicazione: (2026)
di: Yang, Yankai, et al.
Pubblicazione: (2026)
Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
di: Nie, Fan, et al.
Pubblicazione: (2025)
di: Nie, Fan, et al.
Pubblicazione: (2025)
InstructEngine: Instruction-driven Text-to-Image Alignment
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
di: Lu, Xingyu, et al.
Pubblicazione: (2025)
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2025)
di: Zhang, Yi-Fan, et al.
Pubblicazione: (2025)
ContextRL: Enhancing MLLM's Knowledge Discovery Efficiency with Context-Augmented RL
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection
di: Nie, Fan, et al.
Pubblicazione: (2024)
di: Nie, Fan, et al.
Pubblicazione: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
di: Zi, Xing, et al.
Pubblicazione: (2025)
di: Zi, Xing, et al.
Pubblicazione: (2025)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
di: Hong, Yuyang, et al.
Pubblicazione: (2025)
di: Hong, Yuyang, et al.
Pubblicazione: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
di: Yang, Chenglin, et al.
Pubblicazione: (2023)
di: Yang, Chenglin, et al.
Pubblicazione: (2023)
ChartAdapter: Large Vision-Language Model for Chart Summarization
di: Xu, Peixin, et al.
Pubblicazione: (2024)
di: Xu, Peixin, et al.
Pubblicazione: (2024)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
di: Fan, Hector, et al.
Pubblicazione: (2026)
di: Fan, Hector, et al.
Pubblicazione: (2026)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
di: Zhao, Baoquan, et al.
Pubblicazione: (2025)
di: Zhao, Baoquan, et al.
Pubblicazione: (2025)
Unveiling the Visual Rhetoric of Persuasive Cartography: A Case Study of the Design of Octopus Maps
di: Lin, Daocheng, et al.
Pubblicazione: (2025)
di: Lin, Daocheng, et al.
Pubblicazione: (2025)
FCBoost-Net: A Generative Network for Synthesizing Multiple Collocated Outfits via Fashion Compatibility Boosting
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
di: Zhang, Bolin, et al.
Pubblicazione: (2026)
di: Zhang, Bolin, et al.
Pubblicazione: (2026)
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
di: Bao, Guangyin, et al.
Pubblicazione: (2024)
di: Bao, Guangyin, et al.
Pubblicazione: (2024)
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
di: Fan, Chenyang, et al.
Pubblicazione: (2026)
di: Fan, Chenyang, et al.
Pubblicazione: (2026)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
di: Wang, Yingna, et al.
Pubblicazione: (2025)
di: Wang, Yingna, et al.
Pubblicazione: (2025)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
Dependency Structure Augmented Contextual Scoping Framework for Multimodal Aspect-Based Sentiment Analysis
di: Liu, Hao, et al.
Pubblicazione: (2025)
di: Liu, Hao, et al.
Pubblicazione: (2025)
Mutual Information-based Representations Disentanglement for Unaligned Multimodal Language Sequences
di: Qian, Fan, et al.
Pubblicazione: (2024)
di: Qian, Fan, et al.
Pubblicazione: (2024)
DRFormer: A Dual-Regularized Bidirectional Transformer for Person Re-identification
di: Shu, Ying, et al.
Pubblicazione: (2026)
di: Shu, Ying, et al.
Pubblicazione: (2026)
MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering
di: Fan, Xinqi, et al.
Pubblicazione: (2026)
di: Fan, Xinqi, et al.
Pubblicazione: (2026)
Mitigating Image Captioning Hallucinations in Vision-Language Models
di: Zhao, Fei, et al.
Pubblicazione: (2025)
di: Zhao, Fei, et al.
Pubblicazione: (2025)
MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
di: Yang, Fan, et al.
Pubblicazione: (2025)
di: Yang, Fan, et al.
Pubblicazione: (2025)
Band-Attention Modulated RetNet for Face Forgery Detection
di: Zhang, Zhida, et al.
Pubblicazione: (2024)
di: Zhang, Zhida, et al.
Pubblicazione: (2024)
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
di: Chen, Yi-Chun
Pubblicazione: (2025)
di: Chen, Yi-Chun
Pubblicazione: (2025)
Learning Brain Representation with Hierarchical Visual Embeddings
di: Zheng, Jiawen, et al.
Pubblicazione: (2026)
di: Zheng, Jiawen, et al.
Pubblicazione: (2026)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
di: Fan, Linfeng, et al.
Pubblicazione: (2026)
SpikeMba: Multi-Modal Spiking Saliency Mamba for Temporal Video Grounding
di: Li, Wenrui, et al.
Pubblicazione: (2024)
di: Li, Wenrui, et al.
Pubblicazione: (2024)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
di: Wang, Yabing, et al.
Pubblicazione: (2023)
di: Wang, Yabing, et al.
Pubblicazione: (2023)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Scalable Diffusion Models with State Space Backbone
di: Fei, Zhengcong, et al.
Pubblicazione: (2024)
di: Fei, Zhengcong, et al.
Pubblicazione: (2024)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
di: Liu, Xiaofeng, et al.
Pubblicazione: (2023)
di: Liu, Xiaofeng, et al.
Pubblicazione: (2023)
Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
di: Luo, Tianci, et al.
Pubblicazione: (2026)
di: Luo, Tianci, et al.
Pubblicazione: (2026)
A Multimodal Transformer for Live Streaming Highlight Prediction
di: Deng, Jiaxin, et al.
Pubblicazione: (2024)
di: Deng, Jiaxin, et al.
Pubblicazione: (2024)
Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
di: Wang, Junze, et al.
Pubblicazione: (2025)
di: Wang, Junze, et al.
Pubblicazione: (2025)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
di: Li, Wenrui, et al.
Pubblicazione: (2024)
di: Li, Wenrui, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform
di: Lu, Xingyu, et al.
Pubblicazione: (2025) -
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
di: Long, Yancheng, et al.
Pubblicazione: (2026) -
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
di: Yang, Yankai, et al.
Pubblicazione: (2026) -
Towards Generalizable Deepfake Detection via Forgery-aware Audio-Visual Adaptation: A Variational Bayesian Approach
di: Nie, Fan, et al.
Pubblicazione: (2025) -
InstructEngine: Instruction-driven Text-to-Image Alignment
di: Lu, Xingyu, et al.
Pubblicazione: (2025)