Blueprint-Bench: Comparing spatial intelligence of LLMs, agents and image models
Fuente:
arXiv
Guardado en:
| Autores principales: | Petersson, Lukas, Backlund, Axel, Wennstöm, Axel, Petersson, Hanna, Sharrock, Callum, Dabiri, Arash |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence
por: Sharrock, Callum, et al.
Publicado: (2025)
por: Sharrock, Callum, et al.
Publicado: (2025)
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
por: Backlund, Axel, et al.
Publicado: (2025)
por: Backlund, Axel, et al.
Publicado: (2025)
Scalable Optimal Transport Methods in Machine Learning: A Contemporary Survey
por: Khamis, Abdelwahed, et al.
Publicado: (2023)
por: Khamis, Abdelwahed, et al.
Publicado: (2023)
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
por: Ruan, Shouwei, et al.
Publicado: (2025)
por: Ruan, Shouwei, et al.
Publicado: (2025)
Theories of synaptic memory consolidation and intelligent plasticity for continual learning
por: Zenke, Friedemann, et al.
Publicado: (2024)
por: Zenke, Friedemann, et al.
Publicado: (2024)
Topological Deep Learning: A Review of an Emerging Paradigm
por: Zia, Ali, et al.
Publicado: (2023)
por: Zia, Ali, et al.
Publicado: (2023)
Addressing Labelled Data Scarcity: Taxonomy-Agnostic Annotation of PII Values in HTTP Traffic using LLMs
por: Cory, Thomas, et al.
Publicado: (2026)
por: Cory, Thomas, et al.
Publicado: (2026)
Enhancing Programming Pair Workshops: The Case of Teacher Pre-Prompting
por: Petersson, Johan
Publicado: (2025)
por: Petersson, Johan
Publicado: (2025)
Fourier Multipliers on Quasi-Banach Orlicz Spaces and Orlicz Modulation Spaces
por: Petersson, Albin
Publicado: (2025)
por: Petersson, Albin
Publicado: (2025)
Deferred Semantic Binding Language: A Framework for Context-Dependent Semantics
por: Petersson, Joel
Publicado: (2025)
por: Petersson, Joel
Publicado: (2025)
Trace type Orlicz spaces and analysis of Orlicz spaces by Lebesgue exponents
por: Petersson, Albin
Publicado: (2025)
por: Petersson, Albin
Publicado: (2025)
Fourier characterizations and non-triviality of Gelfand-Shilov spaces, with applications to Toeplitz operators
por: Petersson, Albin
Publicado: (2022)
por: Petersson, Albin
Publicado: (2022)
An autonomous agent for auditing and improving the reliability of clinical AI models
por: Kuhn, Lukas, et al.
Publicado: (2025)
por: Kuhn, Lukas, et al.
Publicado: (2025)
BenchECG and xECG: a benchmark and baseline for ECG foundation models
por: Lunelli, Riccardo, et al.
Publicado: (2025)
por: Lunelli, Riccardo, et al.
Publicado: (2025)
RULER: Representation-Level Verification of Machine Unlearning
por: Cosma, Georgina, et al.
Publicado: (2026)
por: Cosma, Georgina, et al.
Publicado: (2026)
HKD-SHO: A hybrid smart home system based on knowledge-based and data-driven services
por: Qiu, Mingming, et al.
Publicado: (2024)
por: Qiu, Mingming, et al.
Publicado: (2024)
Artificial intelligence is algorithmic mimicry: why artificial "agents" are not (and won't be) proper agents
por: Jaeger, Johannes
Publicado: (2023)
por: Jaeger, Johannes
Publicado: (2023)
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
por: Marmoret, Axel
Publicado: (2026)
por: Marmoret, Axel
Publicado: (2026)
RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts
por: Weidener, Lukas, et al.
Publicado: (2026)
por: Weidener, Lukas, et al.
Publicado: (2026)
Towards user-centered interactive medical image segmentation in VR with an assistive AI agent
por: Spiegler, Pascal, et al.
Publicado: (2025)
por: Spiegler, Pascal, et al.
Publicado: (2025)
Who Followed the Blueprint? Analyzing the Responses of U.S. Federal Agencies to the Blueprint for an AI Bill of Rights
por: Lage, Darren, et al.
Publicado: (2024)
por: Lage, Darren, et al.
Publicado: (2024)
PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units
por: Deutel, Mark, et al.
Publicado: (2026)
por: Deutel, Mark, et al.
Publicado: (2026)
A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards
por: Patil, Avinash
Publicado: (2025)
por: Patil, Avinash
Publicado: (2025)
Early-Stage Product Line Validation Using LLMs: A Study on Semi-Formal Blueprint Analysis
por: Le, Viet-Man, et al.
Publicado: (2026)
por: Le, Viet-Man, et al.
Publicado: (2026)
Reasoning Language Models: A Blueprint
por: Besta, Maciej, et al.
Publicado: (2025)
por: Besta, Maciej, et al.
Publicado: (2025)
A Blueprint for Auditing Generative AI
por: Mokander, Jakob, et al.
Publicado: (2024)
por: Mokander, Jakob, et al.
Publicado: (2024)
AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
por: Kirichenko, Polina, et al.
Publicado: (2025)
por: Kirichenko, Polina, et al.
Publicado: (2025)
lmgame-Bench: How Good are LLMs at Playing Games?
por: Hu, Lanxiang, et al.
Publicado: (2025)
por: Hu, Lanxiang, et al.
Publicado: (2025)
ReEfBench: Quantifying the Reasoning Efficiency of LLMs
por: Fu, Zhizhang, et al.
Publicado: (2026)
por: Fu, Zhizhang, et al.
Publicado: (2026)
ConvexBench: Can LLMs Recognize Convex Functions?
por: Liu, Yepeng, et al.
Publicado: (2026)
por: Liu, Yepeng, et al.
Publicado: (2026)
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
por: Bai, Songlin, et al.
Publicado: (2026)
por: Bai, Songlin, et al.
Publicado: (2026)
Grid-Based Projection of Spatial Data into Knowledge Graphs
por: Anjomshoaa, Amin, et al.
Publicado: (2024)
por: Anjomshoaa, Amin, et al.
Publicado: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
por: Kil, Jihyung, et al.
Publicado: (2024)
por: Kil, Jihyung, et al.
Publicado: (2024)
Constructing coherent spatial memory in LLM agents through graph rectification
por: Zhang, Puzhen, et al.
Publicado: (2025)
por: Zhang, Puzhen, et al.
Publicado: (2025)
Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
por: Huh, Dom, et al.
Publicado: (2025)
por: Huh, Dom, et al.
Publicado: (2025)
An adaptive sampling algorithm for data-generation to build a data-manifold for physical problem surrogate modeling
por: Mang, Chetra, et al.
Publicado: (2025)
por: Mang, Chetra, et al.
Publicado: (2025)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
A Blueprint Architecture of Compound AI Systems for Enterprise
por: Kandogan, Eser, et al.
Publicado: (2024)
por: Kandogan, Eser, et al.
Publicado: (2024)
Empathic and agentic artificial intelligence in nursing: perspectives on a human-centered framework for cancer care navigation in the United States
por: Girdwood, Tyra, et al.
Publicado: (2026)
por: Girdwood, Tyra, et al.
Publicado: (2026)
Every Software as an Agent: Blueprint and Case Study
por: Xu, Mengwei
Publicado: (2025)
por: Xu, Mengwei
Publicado: (2025)
Ejemplares similares
-
Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence
por: Sharrock, Callum, et al.
Publicado: (2025) -
Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
por: Backlund, Axel, et al.
Publicado: (2025) -
Scalable Optimal Transport Methods in Machine Learning: A Contemporary Survey
por: Khamis, Abdelwahed, et al.
Publicado: (2023) -
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents
por: Ruan, Shouwei, et al.
Publicado: (2025) -
Theories of synaptic memory consolidation and intelligent plasticity for continual learning
por: Zenke, Friedemann, et al.
Publicado: (2024)