On measuring grounding and generalizing grounding problems
Fuente:
arXiv
Salvato in:
| Autori principali: | Quigley, Daniel, Maynard, Eric |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022)
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
di: Jia, Xiao
Pubblicazione: (2026)
di: Jia, Xiao
Pubblicazione: (2026)
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
di: Berman, Shmuel, et al.
Pubblicazione: (2024)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
di: Alpay, Faruk, et al.
Pubblicazione: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
di: Ashley, Dylan R., et al.
Pubblicazione: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
di: Kaiser, Daniel, et al.
Pubblicazione: (2025)
di: Kaiser, Daniel, et al.
Pubblicazione: (2025)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
di: Kalušev, Vladimir, et al.
Pubblicazione: (2026)
di: Kalušev, Vladimir, et al.
Pubblicazione: (2026)
Intersymbolic AI: Interlinking Symbolic AI and Subsymbolic AI
di: Platzer, André
Pubblicazione: (2024)
di: Platzer, André
Pubblicazione: (2024)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
di: Huo, Dongjie, et al.
Pubblicazione: (2026)
JAM: Controllable and Responsible Text Generation via Causal Reasoning and Latent Vector Manipulation
di: Huang, Yingbing, et al.
Pubblicazione: (2025)
di: Huang, Yingbing, et al.
Pubblicazione: (2025)
Approaches to Semantic Textual Similarity in Slovak Language: From Algorithms to Transformers
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
di: Radosky, Lukas, et al.
Pubblicazione: (2026)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
di: Consoli, Sergio, et al.
Pubblicazione: (2025)
di: Consoli, Sergio, et al.
Pubblicazione: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
di: Zhang, Xue
Pubblicazione: (2025)
di: Zhang, Xue
Pubblicazione: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025)
di: Lei, Xiang, et al.
Pubblicazione: (2025)
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data
di: Koursaris, A., et al.
Pubblicazione: (2026)
di: Koursaris, A., et al.
Pubblicazione: (2026)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
Agentic AI Systems Applied to tasks in Financial Services: Modeling and model risk management crews
di: Okpala, Izunna, et al.
Pubblicazione: (2025)
di: Okpala, Izunna, et al.
Pubblicazione: (2025)
Enhancing Mathematical Problem Solving in LLMs through Execution-Driven Reasoning Augmentation
di: Basarkar, Aditya, et al.
Pubblicazione: (2026)
di: Basarkar, Aditya, et al.
Pubblicazione: (2026)
Data and AI governance: Promoting equity, ethics, and fairness in large language models
di: Abhishek, Alok, et al.
Pubblicazione: (2025)
di: Abhishek, Alok, et al.
Pubblicazione: (2025)
SHARP: Social Harm Analysis via Risk Profiles for Measuring Inequities in Large Language Models
di: Abhishek, Alok, et al.
Pubblicazione: (2026)
di: Abhishek, Alok, et al.
Pubblicazione: (2026)
BEATS: Bias Evaluation and Assessment Test Suite for Large Language Models
di: Abhishek, Alok, et al.
Pubblicazione: (2025)
di: Abhishek, Alok, et al.
Pubblicazione: (2025)
Reasoning Promotes Robustness in Theory of Mind Tasks
di: de Haan, Ian B., et al.
Pubblicazione: (2026)
di: de Haan, Ian B., et al.
Pubblicazione: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
di: Wang, Zhen, et al.
Pubblicazione: (2025)
di: Wang, Zhen, et al.
Pubblicazione: (2025)
Judgment2vec: Apply Graph Analytics to Searching and Recommendation of Similar Judgments
di: Shao, Hsuan-Lei
Pubblicazione: (2024)
di: Shao, Hsuan-Lei
Pubblicazione: (2024)
CLEV: LLM-Based Evaluation Through Lightweight Efficient Voting for Free-Form Question-Answering
di: Badshah, Sher, et al.
Pubblicazione: (2025)
di: Badshah, Sher, et al.
Pubblicazione: (2025)
Multi-Agent Synergy-Driven Iterative Visual Narrative Synthesis
di: Xi, Wang, et al.
Pubblicazione: (2025)
di: Xi, Wang, et al.
Pubblicazione: (2025)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
di: Kalušev, Vladimir, et al.
Pubblicazione: (2025)
di: Kalušev, Vladimir, et al.
Pubblicazione: (2025)
Safe Distributed Control of Multi-Robot Systems with Communication Delays
di: Ballotta, Luca, et al.
Pubblicazione: (2024)
di: Ballotta, Luca, et al.
Pubblicazione: (2024)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
di: Da, Longchao, et al.
Pubblicazione: (2025)
di: Da, Longchao, et al.
Pubblicazione: (2025)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
di: Chen, Tiejin, et al.
Pubblicazione: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
di: Guo, Dongxin, et al.
Pubblicazione: (2026)
HumanMCP: A Human-Like Query Dataset for Evaluating MCP Tool Retrieval Performance
di: Laddha, Shubh, et al.
Pubblicazione: (2025)
di: Laddha, Shubh, et al.
Pubblicazione: (2025)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
di: Cao, Deyu, et al.
Pubblicazione: (2025)
di: Cao, Deyu, et al.
Pubblicazione: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
di: Qi, Dekang, et al.
Pubblicazione: (2026)
di: Qi, Dekang, et al.
Pubblicazione: (2026)
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
di: Zhang, Haoran, et al.
Pubblicazione: (2026)
di: Zhang, Haoran, et al.
Pubblicazione: (2026)
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
di: Zhang, Junwen, et al.
Pubblicazione: (2025)
Semantic Modeling for World-Centered Architectures
di: Mantsivoda, Andrei, et al.
Pubblicazione: (2026)
di: Mantsivoda, Andrei, et al.
Pubblicazione: (2026)
Generative AI and the Transformation of Software Development Practices
di: Acharya, Vivek
Pubblicazione: (2025)
di: Acharya, Vivek
Pubblicazione: (2025)
Modularity in Transformers: Investigating Neuron Separability & Specialization
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
di: Nwokocha, Caleb Princewill
Pubblicazione: (2022) -
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
di: Jia, Xiao
Pubblicazione: (2026) -
Project Synapse: A Hierarchical Multi-Agent Framework with Hybrid Memory for Autonomous Resolution of Last-Mile Delivery Disruptions
di: Yadav, Arin Gopalan, et al.
Pubblicazione: (2026) -
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
di: Berman, Shmuel, et al.
Pubblicazione: (2024) -
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
di: Alpay, Faruk, et al.
Pubblicazione: (2025)