AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dao, Alan, Vu, Dinh Bach |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Jan-nano Technical Report
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
von: Dao, Alan, et al.
Veröffentlicht: (2024)
von: Dao, Alan, et al.
Veröffentlicht: (2024)
PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
VoxRep: Enhancing 3D Spatial Understanding in 2D Vision-Language Models via Voxel Representation
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
von: Wang, Jiaming, et al.
Veröffentlicht: (2025)
von: Wang, Jiaming, et al.
Veröffentlicht: (2025)
Enhancing Document Retrieval in COVID-19 Research: Leveraging Large Language Models for Hidden Relation Extraction
von: Trieu, Hoang-An, et al.
Veröffentlicht: (2025)
von: Trieu, Hoang-An, et al.
Veröffentlicht: (2025)
Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
von: Pimenta, Rui A., et al.
Veröffentlicht: (2025)
von: Pimenta, Rui A., et al.
Veröffentlicht: (2025)
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
von: Zheng, Jinquan, et al.
Veröffentlicht: (2026)
ReZero: Enhancing LLM search ability by trying one-more-time
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
From Slides to Chatbots: Enhancing Large Language Models with University Course Materials
von: Dinh, Tu Anh, et al.
Veröffentlicht: (2025)
von: Dinh, Tu Anh, et al.
Veröffentlicht: (2025)
Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2025)
Exploring the Maze of Multilingual Modeling
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2023)
von: Nezhad, Sina Bagheri, et al.
Veröffentlicht: (2023)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models
von: Xuan, Minh Chu, et al.
Veröffentlicht: (2026)
von: Xuan, Minh Chu, et al.
Veröffentlicht: (2026)
Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
von: Wang, Yining, et al.
Veröffentlicht: (2025)
von: Wang, Yining, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
What Makes a Maze Look Like a Maze?
von: Hsu, Joy, et al.
Veröffentlicht: (2024)
von: Hsu, Joy, et al.
Veröffentlicht: (2024)
Word Definitions from Large Language Models
von: Pham, Bach, et al.
Veröffentlicht: (2023)
von: Pham, Bach, et al.
Veröffentlicht: (2023)
Navigating the Shortcut Maze: A Comprehensive Analysis of Shortcut Learning in Text Classification by Language Models
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
von: Zhou, Yuqing, et al.
Veröffentlicht: (2024)
Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models
von: Hwang, Hyeontaek, et al.
Veröffentlicht: (2026)
von: Hwang, Hyeontaek, et al.
Veröffentlicht: (2026)
Artificial Intelligence and the Spatial Documentation of Languages
von: Ghanim, Hakam
Veröffentlicht: (2024)
von: Ghanim, Hakam
Veröffentlicht: (2024)
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
von: Zhao, Yi, et al.
Veröffentlicht: (2025)
From Emotion Classification to Emotional Reasoning: Enhancing Emotional Intelligence in Large Language Models
von: Sreedar, Arjhun, et al.
Veröffentlicht: (2026)
von: Sreedar, Arjhun, et al.
Veröffentlicht: (2026)
M-GRPO: Stabilizing Self-Supervised Reinforcement Learning for Large Language Models with Momentum-Anchored Policy Optimization
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
von: Bai, Bizhe, et al.
Veröffentlicht: (2025)
Multi-Reward GRPO Fine-Tuning for De-biasing Large Language Models: A Study Based on Chinese-Context Discrimination Data
von: Yixuan, Deng, et al.
Veröffentlicht: (2025)
von: Yixuan, Deng, et al.
Veröffentlicht: (2025)
Navigating the Metrics Maze: Reconciling Score Magnitudes and Accuracies
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
von: Kocmi, Tom, et al.
Veröffentlicht: (2024)
Neural Topic Modeling with Large Language Models in the Loop
von: Yang, Xiaohao, et al.
Veröffentlicht: (2024)
von: Yang, Xiaohao, et al.
Veröffentlicht: (2024)
Large Language Models for Explainable Threat Intelligence
von: Dinis, Tiago, et al.
Veröffentlicht: (2025)
von: Dinis, Tiago, et al.
Veröffentlicht: (2025)
Linguistic Intelligence in Large Language Models for Telecommunications
von: Ahmed, Tasnim, et al.
Veröffentlicht: (2024)
von: Ahmed, Tasnim, et al.
Veröffentlicht: (2024)
LLM Reading Tea Leaves: Automatically Evaluating Topic Models with Large Language Models
von: Yang, Xiaohao, et al.
Veröffentlicht: (2024)
von: Yang, Xiaohao, et al.
Veröffentlicht: (2024)
IgnitionInnovators at "Discharge Me!": Chain-of-Thought Instruction Finetuning Large Language Models for Discharge Summaries
von: Tang, An Quang, et al.
Veröffentlicht: (2024)
von: Tang, An Quang, et al.
Veröffentlicht: (2024)
Teaching a Multilingual Large Language Model to Understand Multilingual Speech via Multi-Instructional Training
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
von: Denisov, Pavel, et al.
Veröffentlicht: (2024)
Evaluating the Symbol Binding Ability of Large Language Models for Multiple-Choice Questions in Vietnamese General Education
von: Nguyen, Duc-Vu, et al.
Veröffentlicht: (2023)
von: Nguyen, Duc-Vu, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
von: Dao, Alan, et al.
Veröffentlicht: (2025) -
Jan-nano Technical Report
von: Dao, Alan, et al.
Veröffentlicht: (2025) -
Lucy: edgerunning agentic web search on mobile with machine generated task vectors
von: Dao, Alan, et al.
Veröffentlicht: (2025) -
Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
von: Dao, Alan, et al.
Veröffentlicht: (2024) -
PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM
von: Dao, Alan, et al.
Veröffentlicht: (2025)