Bongard-OpenWorld: Few-Shot Reasoning for Free-form Visual Concepts in the Real World
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Rujie, Ma, Xiaojian, Zhang, Zhenliang, Wang, Wei, Li, Qing, Zhu, Song-Chun, Wang, Yizhou |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
by: Pawlonka, Szymon, et al.
Published: (2025)
by: Pawlonka, Szymon, et al.
Published: (2025)
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025)
by: Wu, Rujie, et al.
Published: (2025)
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
by: Langenfeld, Cassidy, et al.
Published: (2026)
by: Langenfeld, Cassidy, et al.
Published: (2026)
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
Towards Few-Shot Learning in the Open World: A Review and Beyond
by: Xue, Hui, et al.
Published: (2024)
by: Xue, Hui, et al.
Published: (2024)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
by: Kim, Byeonggeun, et al.
Published: (2024)
by: Kim, Byeonggeun, et al.
Published: (2024)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols
by: Burghouts, Gertjan, et al.
Published: (2024)
by: Burghouts, Gertjan, et al.
Published: (2024)
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
by: Wu, Rujie, et al.
Published: (2026)
by: Wu, Rujie, et al.
Published: (2026)
BoostTaxo: Zero-Shot Taxonomy Induction via Boosting-Style Agentic Reasoning and Constraint-Aware Calibration
by: Ling, Yancheng, et al.
Published: (2026)
by: Ling, Yancheng, et al.
Published: (2026)
FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling
by: Ye, Hang, et al.
Published: (2024)
by: Ye, Hang, et al.
Published: (2024)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
by: Ji, Yuheng, et al.
Published: (2025)
by: Ji, Yuheng, et al.
Published: (2025)
Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?
by: Wüst, Antonia, et al.
Published: (2024)
by: Wüst, Antonia, et al.
Published: (2024)
Visual Analytics for Causal Reasoning from Real-World Health Data
by: Wang, Arran Zeyu, et al.
Published: (2025)
by: Wang, Arran Zeyu, et al.
Published: (2025)
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
by: Wu, Yihao, et al.
Published: (2025)
by: Wu, Yihao, et al.
Published: (2025)
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Generative Compositor for Few-Shot Visual Information Extraction
by: Yang, Zhibo, et al.
Published: (2025)
by: Yang, Zhibo, et al.
Published: (2025)
Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs
by: Hu, Zixuan, et al.
Published: (2024)
by: Hu, Zixuan, et al.
Published: (2024)
Mars: Situated Inductive Reasoning in an Open-World Environment
by: Tang, Xiaojuan, et al.
Published: (2024)
by: Tang, Xiaojuan, et al.
Published: (2024)
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios
by: Meng, Jinxiang, et al.
Published: (2026)
by: Meng, Jinxiang, et al.
Published: (2026)
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds
by: Cui, Christopher Z., et al.
Published: (2024)
by: Cui, Christopher Z., et al.
Published: (2024)
Visual Concept Connectome (VCC): Open World Concept Discovery and their Interlayer Connections in Deep Models
by: Kowal, Matthew, et al.
Published: (2024)
by: Kowal, Matthew, et al.
Published: (2024)
Support-Set Context Matters for Bongard Problems
by: Raghuraman, Nikhil, et al.
Published: (2023)
by: Raghuraman, Nikhil, et al.
Published: (2023)
FlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot Manipulation
by: Guo, Jun, et al.
Published: (2025)
by: Guo, Jun, et al.
Published: (2025)
OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents
by: Wang, Zihao, et al.
Published: (2024)
by: Wang, Zihao, et al.
Published: (2024)
YOLO-World: Real-Time Open-Vocabulary Object Detection
by: Cheng, Tianheng, et al.
Published: (2024)
by: Cheng, Tianheng, et al.
Published: (2024)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
by: Gao, Zhi, et al.
Published: (2023)
by: Gao, Zhi, et al.
Published: (2023)
Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
by: Wang, Zihao, et al.
Published: (2023)
by: Wang, Zihao, et al.
Published: (2023)
Reasoning Limitations of Multimodal Large Language Models. A Case Study of Bongard Problems
by: Małkiński, Mikołaj, et al.
Published: (2024)
by: Małkiński, Mikołaj, et al.
Published: (2024)
Language-Augmented Symbolic Planner for Open-World Task Planning
by: Chen, Guanqi, et al.
Published: (2024)
by: Chen, Guanqi, et al.
Published: (2024)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing
by: Wang, Dianyi, et al.
Published: (2026)
by: Wang, Dianyi, et al.
Published: (2026)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
by: Wei, Shaohang, et al.
Published: (2025)
by: Wei, Shaohang, et al.
Published: (2025)
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
by: Xu, Derek, et al.
Published: (2024)
by: Xu, Derek, et al.
Published: (2024)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
by: Mi, Yapeng, et al.
Published: (2025)
by: Mi, Yapeng, et al.
Published: (2025)
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
by: Lee, Andrew Jun, et al.
Published: (2025)
by: Lee, Andrew Jun, et al.
Published: (2025)
A Concept of Possibility for Real-World Events
by: Schwartz, Daniel G.
Published: (2025)
by: Schwartz, Daniel G.
Published: (2025)
Building Explicit World Model for Zero-Shot Open-World Object Manipulation
by: Li, Xiaotong, et al.
Published: (2026)
by: Li, Xiaotong, et al.
Published: (2026)
Similar Items
-
Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
by: Pawlonka, Szymon, et al.
Published: (2025) -
LongViTU: Instruction Tuning for Long-Form Video Understanding
by: Wu, Rujie, et al.
Published: (2025) -
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
by: Langenfeld, Cassidy, et al.
Published: (2026) -
ROCKET-1: Mastering Open-World Interaction with Visual-Temporal Context Prompting
by: Cai, Shaofei, et al.
Published: (2024) -
Towards Few-Shot Learning in the Open World: A Review and Beyond
by: Xue, Hui, et al.
Published: (2024)