Socratic Planner: Self-QA-Based Zero-Shot Planning for Embodied Instruction Following
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shin, Suyeon, jeon, Sujin, Kim, Junghyun, Kang, Gi-Cheon, Zhang, Byoung-Tak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
von: Brusnicki, Roberto, et al.
Veröffentlicht: (2026)
von: Brusnicki, Roberto, et al.
Veröffentlicht: (2026)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
von: Ferdousi, Rahatara, et al.
Veröffentlicht: (2025)
von: Ferdousi, Rahatara, et al.
Veröffentlicht: (2025)
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
von: Štefánik, Michal, et al.
Veröffentlicht: (2025)
Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
von: Souverain, Thomas
Veröffentlicht: (2025)
von: Souverain, Thomas
Veröffentlicht: (2025)
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
von: Wu, Ren-Di, et al.
Veröffentlicht: (2025)
von: Wu, Ren-Di, et al.
Veröffentlicht: (2025)
Evaluating AI Grading on Real-World Handwritten College Mathematics: A Large-Scale Study Toward a Benchmark
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
From MTEB to MTOB: Retrieval-Augmented Classification for Descriptive Grammars
von: Kornilov, Albert, et al.
Veröffentlicht: (2024)
von: Kornilov, Albert, et al.
Veröffentlicht: (2024)
AI Assistants for Spaceflight Procedures: Combining Generative Pre-Trained Transformer and Retrieval-Augmented Generation on Knowledge Graphs With Augmented Reality Cues
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
Sequence Transferability and Task Order Selection in Continual Learning
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thinh, et al.
Veröffentlicht: (2025)
The Trap of Presumed Equivalence: Artificial General Intelligence Should Not Be Assessed on the Scale of Human Intelligence
von: Dolgikh, Serge
Veröffentlicht: (2024)
von: Dolgikh, Serge
Veröffentlicht: (2024)
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
von: Aksoy, Sinan G., et al.
Veröffentlicht: (2026)
The Shape of Math To Come
von: Kontorovich, Alex
Veröffentlicht: (2025)
von: Kontorovich, Alex
Veröffentlicht: (2025)
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
von: Salehi, Pegah, et al.
Veröffentlicht: (2024)
von: Salehi, Pegah, et al.
Veröffentlicht: (2024)
Biomedical Visual Instruction Tuning with Clinician Preference Alignment
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
von: Cui, Hejie, et al.
Veröffentlicht: (2024)
COCO-Urdu: A Large-Scale Urdu Image-Caption Dataset with Multimodal Quality Estimation
von: Hassan, Umair
Veröffentlicht: (2025)
von: Hassan, Umair
Veröffentlicht: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
von: Li, Mingda, et al.
Veröffentlicht: (2024)
von: Li, Mingda, et al.
Veröffentlicht: (2024)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
Towards a Reliable Offline Personal AI Assistant for Long Duration Spaceflight
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
von: Bensch, Oliver, et al.
Veröffentlicht: (2024)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2024)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
von: Arabov, Mullosharaf K.
Veröffentlicht: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
von: Yao, Huaiyuan, et al.
Veröffentlicht: (2026)
von: Yao, Huaiyuan, et al.
Veröffentlicht: (2026)
Separate Before You Compress: The WWHO Tokenization Architecture
von: Darshana, Kusal
Veröffentlicht: (2026)
von: Darshana, Kusal
Veröffentlicht: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
Attire-Based Anomaly Detection in Restricted Areas Using YOLOv8 for Enhanced CCTV Security
von: B, Abdul Aziz A., et al.
Veröffentlicht: (2024)
von: B, Abdul Aziz A., et al.
Veröffentlicht: (2024)
If you can describe it, they can see it: Cross-Modal Learning of Visual Concepts from Textual Descriptions
von: Barbano, Carlo Alberto, et al.
Veröffentlicht: (2024)
von: Barbano, Carlo Alberto, et al.
Veröffentlicht: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
Depth Priors in Removal Neural Radiance Fields
von: Guo, Zhihao, et al.
Veröffentlicht: (2024)
von: Guo, Zhihao, et al.
Veröffentlicht: (2024)
Does CLIP perceive art the same way we do?
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
von: Asperti, Andrea, et al.
Veröffentlicht: (2025)
A Human-In-The-Loop Approach for Improving Fairness in Predictive Business Process Monitoring
von: Käppel, Martin, et al.
Veröffentlicht: (2025)
von: Käppel, Martin, et al.
Veröffentlicht: (2025)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
Empowering Manufacturers with Privacy-Preserving AI Tools: A Case Study in Privacy-Preserving Machine Learning to Solve Real-World Problems
von: Ji, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Ji, Xiaoyu, et al.
Veröffentlicht: (2025)
Prompting Encoder Models for Zero-Shot Classification: A Cross-Domain Study in Italian
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
von: Auriemma, Serena, et al.
Veröffentlicht: (2024)
Low-Resource Neural Machine Translation Using Recurrent Neural Networks and Transfer Learning: A Case Study on English-to-Igbo
von: Ekle, Ocheme Anthony, et al.
Veröffentlicht: (2025)
von: Ekle, Ocheme Anthony, et al.
Veröffentlicht: (2025)
Surfing the modeling of PoS taggers in low-resource scenarios
von: Ferro, Manuel Vilares, et al.
Veröffentlicht: (2024)
von: Ferro, Manuel Vilares, et al.
Veröffentlicht: (2024)
Z-Order Transformer for Feed-Forward Gaussian Splatting
von: Wang, Can, et al.
Veröffentlicht: (2026)
von: Wang, Can, et al.
Veröffentlicht: (2026)
EatGAN: An Edge-Attention Guided Generative Adversarial Network for Single Image Super-Resolution
von: Rao, Penghao, et al.
Veröffentlicht: (2025)
von: Rao, Penghao, et al.
Veröffentlicht: (2025)
TWIG: Two-Step Image Generation using Segmentation Masks in Diffusion Models
von: Rakib, Mazharul Islam, et al.
Veröffentlicht: (2025)
von: Rakib, Mazharul Islam, et al.
Veröffentlicht: (2025)
Isolated Sign Language Recognition with Segmentation and Pose Estimation
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
von: Perkins, Daniel, et al.
Veröffentlicht: (2025)
Forging GEMs: Advancing Greek NLP through Quality-Based Corpus Curation
von: Apostolopoulou, Alexandra, et al.
Veröffentlicht: (2025)
von: Apostolopoulou, Alexandra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
von: Brusnicki, Roberto, et al.
Veröffentlicht: (2026) -
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
von: Ferdousi, Rahatara, et al.
Veröffentlicht: (2025) -
Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering
von: Štefánik, Michal, et al.
Veröffentlicht: (2025) -
Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology
von: Souverain, Thomas
Veröffentlicht: (2025) -
SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval
von: Wu, Ren-Di, et al.
Veröffentlicht: (2025)