Harnessing Large Vision and Language Models in Agriculture: A Review
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Hongyan, Qin, Shuai, Su, Min, Lin, Chengzhi, Li, Anjie, Gao, Junfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Review of Hallucination Understanding in Large Language and Vision Models
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025)
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
von: Lin, Zihang, et al.
Veröffentlicht: (2026)
von: Lin, Zihang, et al.
Veröffentlicht: (2026)
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
von: Lin, Ling, et al.
Veröffentlicht: (2026)
von: Lin, Ling, et al.
Veröffentlicht: (2026)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
von: Song, Python, et al.
Veröffentlicht: (2025)
von: Song, Python, et al.
Veröffentlicht: (2025)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
AgroGPT: Efficient Agricultural Vision-Language Model with Expert Tuning
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
von: Awais, Muhammad, et al.
Veröffentlicht: (2024)
Vision Language Model-Empowered Contract Theory for AIGC Task Allocation in Teleoperation
von: Zhan, Zijun, et al.
Veröffentlicht: (2024)
von: Zhan, Zijun, et al.
Veröffentlicht: (2024)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
PLPHP: Per-Layer Per-Head Vision Token Pruning for Efficient Large Vision-Language Models
von: Meng, Yu, et al.
Veröffentlicht: (2025)
von: Meng, Yu, et al.
Veröffentlicht: (2025)
ArchiLense: A Framework for Quantitative Analysis of Architectural Styles Based on Vision Large Language Models
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
von: Zhong, Jing, et al.
Veröffentlicht: (2025)
UrbanSense:A Framework for Quantitative Analysis of Urban Streetscapes leveraging Vision Large Language Models
von: Yin, Jun, et al.
Veröffentlicht: (2025)
von: Yin, Jun, et al.
Veröffentlicht: (2025)
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
von: Li, Guanzhen, et al.
Veröffentlicht: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaobei, et al.
Veröffentlicht: (2025)
Breaking Modality Heterogeneity in Low-Bit Quantization for Large Vision-Language Models
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
von: Zhong, Yi, et al.
Veröffentlicht: (2026)
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
von: Yan, Siming, et al.
Veröffentlicht: (2024)
von: Yan, Siming, et al.
Veröffentlicht: (2024)
RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
ImgTrojan: Jailbreaking Vision-Language Models with ONE Image
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
von: Tao, Xijia, et al.
Veröffentlicht: (2024)
Probing Perceptual Constancy in Large Vision-Language Models
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
Large Vision-Language Models Get Lost in Attention
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
von: Xi, Gongli, et al.
Veröffentlicht: (2026)
Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation
von: Feng, Yongchao, et al.
Veröffentlicht: (2025)
von: Feng, Yongchao, et al.
Veröffentlicht: (2025)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
von: Yin, Jianghao, et al.
Veröffentlicht: (2026)
Cross-Cultural Value Awareness in Large Vision-Language Models
von: Howard, Phillip, et al.
Veröffentlicht: (2026)
von: Howard, Phillip, et al.
Veröffentlicht: (2026)
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
von: Woo, Sangmin, et al.
Veröffentlicht: (2025)
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
von: Zhu, Younan, et al.
Veröffentlicht: (2025)
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
von: Boudiaf, Abderrahmene, et al.
Veröffentlicht: (2026)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Delineating Knowledge Boundaries for Honest Large Vision-Language Models
von: Song, Junru, et al.
Veröffentlicht: (2026)
von: Song, Junru, et al.
Veröffentlicht: (2026)
KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
von: Wang, Yujin, et al.
Veröffentlicht: (2025)
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Review of Hallucination Understanding in Large Language and Vision Models
von: Ho, Zhengyi, et al.
Veröffentlicht: (2025) -
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
von: Lin, Zihang, et al.
Veröffentlicht: (2026) -
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
von: Lin, Ling, et al.
Veröffentlicht: (2026) -
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
von: Song, Python, et al.
Veröffentlicht: (2025) -
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)