Navigating Simply, Aligning Deeply: Winning Solutions for Mouse vs. AI 2025
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Phu-Hoa, Tran, Chi-Nguyen, Minh, Dao Sy Duy, Quy, Nguyen Lam Phu, Kiet, Huynh Trung |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
von: Minh, Dao Sy Duy, et al.
Veröffentlicht: (2025)
von: Minh, Dao Sy Duy, et al.
Veröffentlicht: (2025)
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
von: Quy, Nguyen Lam Phu, et al.
Veröffentlicht: (2025)
von: Quy, Nguyen Lam Phu, et al.
Veröffentlicht: (2025)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
More at Stake: How Payoff and Language Shape LLM Agent Strategies in Cooperation Dilemmas
von: Huynh, Trung-Kiet, et al.
Veröffentlicht: (2026)
von: Huynh, Trung-Kiet, et al.
Veröffentlicht: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
Towards Accurate and Efficient Waste Image Classification: A Hybrid Deep Learning and Machine Learning Approach
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
von: Nguyen, Ngoc-Bao-Quang, et al.
Veröffentlicht: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
von: Dai, Song, et al.
Veröffentlicht: (2025)
von: Dai, Song, et al.
Veröffentlicht: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
von: Tong, Jingqi, et al.
Veröffentlicht: (2025)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Bian, Zhipeng, et al.
Veröffentlicht: (2026)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
von: Endo, Masafumi, et al.
Veröffentlicht: (2024)
von: Endo, Masafumi, et al.
Veröffentlicht: (2024)
VietMed-MCQ: A Consistency-Filtered Data Synthesis Framework for Vietnamese Traditional Medicine Evaluation
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
von: Kiet, Huynh Trung, et al.
Veröffentlicht: (2026)
A Survey of Spatial Memory Representations for Efficient Robot Navigation
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
von: Pangaliman, Ma. Madecheen S., et al.
Veröffentlicht: (2026)
RAE-NWM: Navigation World Model in Dense Visual Representation Space
von: Zhang, Mingkun, et al.
Veröffentlicht: (2026)
von: Zhang, Mingkun, et al.
Veröffentlicht: (2026)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
von: Doan, Gia-Bao, et al.
Veröffentlicht: (2026)
von: Doan, Gia-Bao, et al.
Veröffentlicht: (2026)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
von: Tran, Duc Dang Trung, et al.
Veröffentlicht: (2024)
Zero-Shot Object Goal Visual Navigation With Class-Independent Relationship Network
von: Li, Xinting, et al.
Veröffentlicht: (2023)
von: Li, Xinting, et al.
Veröffentlicht: (2023)
PC-SNN: Predictive Coding-based Local Hebbian Plasticity Learning in Spiking Neural Networks
von: Wang, Haidong, et al.
Veröffentlicht: (2022)
von: Wang, Haidong, et al.
Veröffentlicht: (2022)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Accelerating Post-Tornado Disaster Assessment Using Advanced Deep Learning Models
von: Umeike, Robinson, et al.
Veröffentlicht: (2024)
von: Umeike, Robinson, et al.
Veröffentlicht: (2024)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
Learning from Watching: Scalable Extraction of Manipulation Trajectories from Human Videos
von: Hu, X., et al.
Veröffentlicht: (2025)
von: Hu, X., et al.
Veröffentlicht: (2025)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
Transformers for Image-Goal Navigation
von: Pelluri, Nikhilanj
Veröffentlicht: (2024)
von: Pelluri, Nikhilanj
Veröffentlicht: (2024)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
von: Alanazi, Ahmed, et al.
Veröffentlicht: (2025)
von: Alanazi, Ahmed, et al.
Veröffentlicht: (2025)
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
von: Shah, Nisarg A., et al.
Veröffentlicht: (2025)
Revisiting SVD and Wavelet Difference Reduction for Lossy Image Compression: A Reproducibility Study
von: Makarova, Alena
Veröffentlicht: (2025)
von: Makarova, Alena
Veröffentlicht: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
von: Bucher, Martin JJ., et al.
Veröffentlicht: (2025)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
von: Zhou, Fang, et al.
Veröffentlicht: (2025)
von: Zhou, Fang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leveraging Lightweight Entity Extraction for Scalable Event-Based Image Retrieval
von: Minh, Dao Sy Duy, et al.
Veröffentlicht: (2025) -
Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
von: Quy, Nguyen Lam Phu, et al.
Veröffentlicht: (2025) -
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
von: Anh, Duy Le Dinh, et al.
Veröffentlicht: (2024) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)