Paladin-mini: A Compact and Efficient Grounding Model Excelling in Real-World Scenarios
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ivry, Dror, Nahum, Oran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sentinel: SOTA model to protect against prompt injections
von: Ivry, Dror, et al.
Veröffentlicht: (2025)
von: Ivry, Dror, et al.
Veröffentlicht: (2025)
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
von: Pang, Yan, et al.
Veröffentlicht: (2025)
von: Pang, Yan, et al.
Veröffentlicht: (2025)
Challenges in Grounding Language in the Real World
von: Lindes, Peter, et al.
Veröffentlicht: (2025)
von: Lindes, Peter, et al.
Veröffentlicht: (2025)
From Real-World Traffic Data to Relevant Critical Scenarios
von: Lüttner, Florian, et al.
Veröffentlicht: (2025)
von: Lüttner, Florian, et al.
Veröffentlicht: (2025)
Efficient and Versatile Model for Multilingual Information Retrieval of Islamic Text: Development and Deployment in Real-World Scenarios
von: Pavlova, Vera, et al.
Veröffentlicht: (2025)
von: Pavlova, Vera, et al.
Veröffentlicht: (2025)
Jenius Agent: Towards Experience-Driven Accuracy Optimization in Real-World Scenarios
von: Xia, Defei, et al.
Veröffentlicht: (2026)
von: Xia, Defei, et al.
Veröffentlicht: (2026)
MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
von: Song, Zhiheng, et al.
Veröffentlicht: (2026)
DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
von: Gao, Zeyu, et al.
Veröffentlicht: (2025)
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
von: Wu, Yao, et al.
Veröffentlicht: (2026)
von: Wu, Yao, et al.
Veröffentlicht: (2026)
3D-Anchored Lookahead Planning for Persistent Robotic Scene Memory via World-Model-Based MCTS
von: Sidik, Bronislav, et al.
Veröffentlicht: (2026)
von: Sidik, Bronislav, et al.
Veröffentlicht: (2026)
TIME: A Multi-level Benchmark for Temporal Reasoning of LLMs in Real-World Scenarios
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
von: Wei, Shaohang, et al.
Veröffentlicht: (2025)
The Pump Scheduling Problem: A Real-World Scenario for Reinforcement Learning
von: Donâncio, Henrique, et al.
Veröffentlicht: (2022)
von: Donâncio, Henrique, et al.
Veröffentlicht: (2022)
MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios
von: Li, Zhang, et al.
Veröffentlicht: (2026)
von: Li, Zhang, et al.
Veröffentlicht: (2026)
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Ruiwen, et al.
Veröffentlicht: (2024)
TRIP-Bench: A Benchmark for Long-Horizon Interactive Agents in Real-World Scenarios
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
von: Shen, Yuanzhe, et al.
Veröffentlicht: (2026)
Experimental Evaluation of ROS-Causal in Real-World Human-Robot Spatial Interaction Scenarios
von: Castri, Luca, et al.
Veröffentlicht: (2024)
von: Castri, Luca, et al.
Veröffentlicht: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
von: Wu, Junchao, et al.
Veröffentlicht: (2024)
SCoGen: Scenario-Centric Graph-Based Synthesis of Real-World Code Problems
von: Yao, Xifeng, et al.
Veröffentlicht: (2025)
von: Yao, Xifeng, et al.
Veröffentlicht: (2025)
Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?
von: Zeng, Qingyuan, et al.
Veröffentlicht: (2026)
von: Zeng, Qingyuan, et al.
Veröffentlicht: (2026)
World Models as an Intermediary between Agents and the Real World
von: Yang, Sherry
Veröffentlicht: (2026)
von: Yang, Sherry
Veröffentlicht: (2026)
From Understanding to Excelling: Template-Free Algorithm Design through Structural-Functional Co-Evolution
von: Zhao, Zhe, et al.
Veröffentlicht: (2025)
von: Zhao, Zhe, et al.
Veröffentlicht: (2025)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
Grounded World Model for Semantically Generalizable Planning
von: Li, Quanyi, et al.
Veröffentlicht: (2026)
von: Li, Quanyi, et al.
Veröffentlicht: (2026)
Neurosymbolic Grounding for Compositional World Models
von: Sehgal, Atharva, et al.
Veröffentlicht: (2023)
von: Sehgal, Atharva, et al.
Veröffentlicht: (2023)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
von: Rao, Delip, et al.
Veröffentlicht: (2026)
von: Rao, Delip, et al.
Veröffentlicht: (2026)
UniTalk: Towards Universal Active Speaker Detection in Real World Scenarios
von: Nguyen, Le Thien Phuc, et al.
Veröffentlicht: (2025)
von: Nguyen, Le Thien Phuc, et al.
Veröffentlicht: (2025)
Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
von: Corbière, Charles, et al.
Veröffentlicht: (2025)
von: Corbière, Charles, et al.
Veröffentlicht: (2025)
NeoRL-2: Near Real-World Benchmarks for Offline Reinforcement Learning with Extended Realistic Scenarios
von: Gao, Songyi, et al.
Veröffentlicht: (2025)
von: Gao, Songyi, et al.
Veröffentlicht: (2025)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios
von: Ding, Yihang, et al.
Veröffentlicht: (2026)
von: Ding, Yihang, et al.
Veröffentlicht: (2026)
From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models
von: Chen, Zhikang, et al.
Veröffentlicht: (2026)
von: Chen, Zhikang, et al.
Veröffentlicht: (2026)
AI Planning Framework for LLM-Based Web Agents
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
von: Shahnovsky, Orit, et al.
Veröffentlicht: (2026)
Best of mini-N in-loop Sampling: A Contextual Quality Reward Model for Reliable and Efficient Best-of-N Sampling
von: Rho, Hyung Gyu, et al.
Veröffentlicht: (2025)
von: Rho, Hyung Gyu, et al.
Veröffentlicht: (2025)
Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios
von: Lei, Liangqi, et al.
Veröffentlicht: (2025)
von: Lei, Liangqi, et al.
Veröffentlicht: (2025)
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
von: Chen, Zhi, et al.
Veröffentlicht: (2024)
MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
von: Xiong, Xuantang, et al.
Veröffentlicht: (2025)
von: Xiong, Xuantang, et al.
Veröffentlicht: (2025)
WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
von: Lu, Shuo, et al.
Veröffentlicht: (2026)
Interpretable Unsupervised Joint Denoising and Enhancement for Real-World low-light Scenarios
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
von: Li, Huaqiu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sentinel: SOTA model to protect against prompt injections
von: Ivry, Dror, et al.
Veröffentlicht: (2025) -
OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
von: Gao, Hong, et al.
Veröffentlicht: (2025) -
Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm
von: Pang, Yan, et al.
Veröffentlicht: (2025) -
Challenges in Grounding Language in the Real World
von: Lindes, Peter, et al.
Veröffentlicht: (2025) -
From Real-World Traffic Data to Relevant Critical Scenarios
von: Lüttner, Florian, et al.
Veröffentlicht: (2025)