Vision-Braille: A Curriculum Learning Toolkit and Braille-Chinese Corpus for Braille Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Alan, Yuan, Ye, Xiao, Zhiping, Zhang, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Braille to Text Translation for Bengali Language: A Geometric Approach
von: Kamal, Minhas, et al.
Veröffentlicht: (2020)
von: Kamal, Minhas, et al.
Veröffentlicht: (2020)
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025)
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
von: Wanchoo, Karan, et al.
Veröffentlicht: (2024)
von: Wanchoo, Karan, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
von: Wang, Xintong, et al.
Veröffentlicht: (2025)
TUBench: Benchmarking Large Vision-Language Models on Trustworthiness with Unanswerable Questions
von: He, Xingwei, et al.
Veröffentlicht: (2024)
von: He, Xingwei, et al.
Veröffentlicht: (2024)
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
CVLUE: A New Benchmark Dataset for Chinese Vision-Language Understanding Evaluation
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
von: Yuan, Ruifeng, et al.
Veröffentlicht: (2025)
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies
von: Acosta-Triana, José-M., et al.
Veröffentlicht: (2024)
von: Acosta-Triana, José-M., et al.
Veröffentlicht: (2024)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
von: Zhang, Yikai, et al.
Veröffentlicht: (2024)
ArtGPT-4: Towards Artistic-understanding Large Vision-Language Models with Enhanced Adapter
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
von: Yuan, Zhengqing, et al.
Veröffentlicht: (2023)
BiomedCoOp: Learning to Prompt for Biomedical Vision-Language Models
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
von: Koleilat, Taha, et al.
Veröffentlicht: (2024)
Learning Rate Curriculum
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
von: Croitoru, Florinel-Alin, et al.
Veröffentlicht: (2022)
VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
PyVision: Agentic Vision with Dynamic Tooling
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
von: Zhao, Shitian, et al.
Veröffentlicht: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
Pseudo-Prompt Generating in Pre-trained Vision-Language Models for Multi-Label Medical Image Classification
von: Ye, Yaoqin, et al.
Veröffentlicht: (2024)
von: Ye, Yaoqin, et al.
Veröffentlicht: (2024)
Autoregressive Models in Vision: A Survey
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
von: Xiong, Jing, et al.
Veröffentlicht: (2024)
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
HELPD: Mitigating Hallucination of LVLMs by Hierarchical Feedback Learning with Vision-enhanced Penalty Decoding
von: Yuan, Fan, et al.
Veröffentlicht: (2024)
von: Yuan, Fan, et al.
Veröffentlicht: (2024)
Evi-Steer: Learning to Steer Biomedical Vision-Language Models through Efficient and Generalizable Evidential Tuning
von: Koleilat, Taha, et al.
Veröffentlicht: (2026)
von: Koleilat, Taha, et al.
Veröffentlicht: (2026)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
Seeing the Poem: Image-Semantic Detection of AI-Generated Modern Chinese Poetry with MLLMs
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
von: Wang, Shanshan, et al.
Veröffentlicht: (2026)
TransLPRNet: Lite Vision-Language Network for Single/Dual-line Chinese License Plate Recognition
von: Xu, Guangzhu, et al.
Veröffentlicht: (2025)
von: Xu, Guangzhu, et al.
Veröffentlicht: (2025)
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
von: Futeral, Matthieu, et al.
Veröffentlicht: (2024)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
von: Li, Sifan, et al.
Veröffentlicht: (2025)
von: Li, Sifan, et al.
Veröffentlicht: (2025)
TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
von: Ye, Jinlun, et al.
Veröffentlicht: (2026)
von: Ye, Jinlun, et al.
Veröffentlicht: (2026)
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2025)
Direct Translation between Sign Languages
von: Wu, Zetian, et al.
Veröffentlicht: (2026)
von: Wu, Zetian, et al.
Veröffentlicht: (2026)
Play to Generalize: Learning to Reason Through Game Play
von: Xie, Yunfei, et al.
Veröffentlicht: (2025)
von: Xie, Yunfei, et al.
Veröffentlicht: (2025)
Granular Privacy Control for Geolocation with Vision Language Models
von: Mendes, Ethan, et al.
Veröffentlicht: (2024)
von: Mendes, Ethan, et al.
Veröffentlicht: (2024)
GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
von: Xia, Renqiu, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Few-Shot Adaptation with Negative Learning
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
von: Zhang, Ce, et al.
Veröffentlicht: (2024)
A Hong Kong Sign Language Corpus Collected from Sign-interpreted TV News
von: Niu, Zhe, et al.
Veröffentlicht: (2024)
von: Niu, Zhe, et al.
Veröffentlicht: (2024)
Dual-branch Prompting for Multimodal Machine Translation
von: Wang, Jie, et al.
Veröffentlicht: (2025)
von: Wang, Jie, et al.
Veröffentlicht: (2025)
Vision as LoRA
von: Wang, Han, et al.
Veröffentlicht: (2025)
von: Wang, Han, et al.
Veröffentlicht: (2025)
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models
von: Wang, Qidong, et al.
Veröffentlicht: (2026)
von: Wang, Qidong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Braille to Text Translation for Bengali Language: A Geometric Approach
von: Kamal, Minhas, et al.
Veröffentlicht: (2020) -
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025) -
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024) -
NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
von: Wanchoo, Karan, et al.
Veröffentlicht: (2024) -
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)