LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Yi, Wang, Siqi, Li, Jing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
von: Mo, Shentong, et al.
Veröffentlicht: (2026)
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
von: Pan, Yiming, et al.
Veröffentlicht: (2026)
von: Pan, Yiming, et al.
Veröffentlicht: (2026)
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
Modeling Layout Reading Order as Ordering Relations for Visually-rich Document Understanding
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
von: Zhang, Chong, et al.
Veröffentlicht: (2024)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
von: Wang, Yining, et al.
Veröffentlicht: (2025)
von: Wang, Yining, et al.
Veröffentlicht: (2025)
MIND Your Reasoning: A Meta-Cognitive Intuitive-Reflective Network for Dual-Reasoning in Multimodal Stance Detection
von: Wang, Bingbing, et al.
Veröffentlicht: (2025)
von: Wang, Bingbing, et al.
Veröffentlicht: (2025)
Training Data Efficiency in Multimodal Process Reward Models
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
von: Li, Jinyuan, et al.
Veröffentlicht: (2026)
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
von: Chen, Yi-Chun, et al.
Veröffentlicht: (2025)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
von: Yi, Xuan, et al.
Veröffentlicht: (2024)
MIDI-LLM: Adapting Large Language Models for Text-to-MIDI Music Generation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2025)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2025)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
von: Wang, Ziteng, et al.
Veröffentlicht: (2026)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
Text2Sign Diffusion: A Generative Approach for Gloss-Free Sign Language Production
von: Feng, Liqian, et al.
Veröffentlicht: (2025)
von: Feng, Liqian, et al.
Veröffentlicht: (2025)
P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
von: Sun, Tao, et al.
Veröffentlicht: (2025)
von: Sun, Tao, et al.
Veröffentlicht: (2025)
Improving Generalization in Intent Detection: GRPO with Reward-Based Curriculum Sampling
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
von: Feng, Zihao, et al.
Veröffentlicht: (2025)
Towards Better Text-to-Image Generation Alignment via Attention Modulation
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
von: Wu, Yihang, et al.
Veröffentlicht: (2024)
Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction
von: Ali, Mai, et al.
Veröffentlicht: (2025)
von: Ali, Mai, et al.
Veröffentlicht: (2025)
Towards Event Extraction from Speech with Contextual Clues
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
von: Zhao, Fei, et al.
Veröffentlicht: (2025)
Distilling Implicit Multimodal Knowledge into Large Language Models for Zero-Resource Dialogue Generation
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
von: Zhang, Bo, et al.
Veröffentlicht: (2024)
ShieldVLM: Safeguarding the Multimodal Implicit Toxicity via Deliberative Reasoning with LVLMs
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
von: Cui, Shiyao, et al.
Veröffentlicht: (2025)
MIntRec2.0: A Large-scale Benchmark Dataset for Multimodal Intent Recognition and Out-of-scope Detection in Conversations
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2024)
Visual Set Program Synthesizer
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
von: Li, Zhu, et al.
Veröffentlicht: (2026)
von: Li, Zhu, et al.
Veröffentlicht: (2026)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
von: Li, Haoxuan, et al.
Veröffentlicht: (2025)
TCAN: Text-oriented Cross Attention Network for Multimodal Sentiment Analysis
von: Quan, Weize, et al.
Veröffentlicht: (2024)
von: Quan, Weize, et al.
Veröffentlicht: (2024)
Sentiment-enhanced Graph-based Sarcasm Explanation in Dialogue
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
von: Chen, Xiaolin, et al.
Veröffentlicht: (2022)
Retrieval-Augmented Multimodal Model for Fake News Detection
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
MLANet: Multi-Level Attention Network with Sub-instruction for Continuous Vision-and-Language Navigation
von: He, Zongtao, et al.
Veröffentlicht: (2023)
von: He, Zongtao, et al.
Veröffentlicht: (2023)
TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Improving Visual Representation Alignment Generation with GRPO
von: Mo, Shentong, et al.
Veröffentlicht: (2026) -
LVRPO: Language-Visual Alignment with GRPO for Multimodal Understanding and Generation
von: Mo, Shentong, et al.
Veröffentlicht: (2026) -
AeSlides: Incentivizing Aesthetic Layout in LLM-Based Slide Generation via Verifiable Rewards
von: Pan, Yiming, et al.
Veröffentlicht: (2026) -
Robust Symbolic Reasoning for Visual Narratives via Hierarchical and Semantically Normalized Knowledge Graphs
von: Chen, Yi-Chun
Veröffentlicht: (2025) -
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
von: Li, Jia, et al.
Veröffentlicht: (2025)