Similar Items
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
by: Davide, Fabrizio, et al.
Published: (2024)
by: Davide, Fabrizio, et al.
Published: (2024)
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024)
by: Wen, Yuqiao, et al.
Published: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026)
by: Scantamburlo, Riccardo, et al.
Published: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
by: Costa, Rimom
Published: (2025)
by: Costa, Rimom
Published: (2025)
Exploring Model Invariance with Discrete Search for Ultra-Low-Bit Quantization
by: Wen, Yuqiao, et al.
Published: (2025)
by: Wen, Yuqiao, et al.
Published: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
A Review of the Applications of Deep Learning-Based Emergent Communication
by: Boldt, Brendon, et al.
Published: (2024)
by: Boldt, Brendon, et al.
Published: (2024)
MEMTIER: Tiered Memory Architecture and Retrieval Bottleneck Analysis for Long-Running Autonomous AI Agents
by: Sidik, Bronislav, et al.
Published: (2026)
by: Sidik, Bronislav, et al.
Published: (2026)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
by: Pathak, Gangesh, et al.
Published: (2025)
by: Pathak, Gangesh, et al.
Published: (2025)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Helping Johnny Make Sense of Privacy Policies with LLMs
by: Freiberger, Vincent, et al.
Published: (2025)
by: Freiberger, Vincent, et al.
Published: (2025)
PRISM: Prompt Reliability via Iterative Simulation and Monitoring for Enterprise Conversational AI
by: Chaitanya, Keshava, et al.
Published: (2026)
by: Chaitanya, Keshava, et al.
Published: (2026)
ChatGPT4PCG 2 Competition: Prompt Engineering for Science Birds Level Generation
by: Taveekitworachai, Pittawat, et al.
Published: (2024)
by: Taveekitworachai, Pittawat, et al.
Published: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
Knowing Isn't Understanding: Re-grounding Generative Proactivity with Epistemic and Behavioral Insight
by: Kaur, Kirandeep, et al.
Published: (2026)
by: Kaur, Kirandeep, et al.
Published: (2026)
Do We Always Need Query-Level Workflows? Rethinking Agentic Workflow Generation for Multi-Agent Systems
by: Wang, Zixu, et al.
Published: (2026)
by: Wang, Zixu, et al.
Published: (2026)
How to Evaluate Medical AI
by: Kopanichuk, Ilia, et al.
Published: (2025)
by: Kopanichuk, Ilia, et al.
Published: (2025)
Opinion Mining on Offshore Wind Energy for Environmental Engineering
by: Bittencourt, Isabele, et al.
Published: (2024)
by: Bittencourt, Isabele, et al.
Published: (2024)
ATANT: An Evaluation Framework for AI Continuity
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
Social Cooperation in Conversational AI Agents
by: Çelikok, Mustafa Mert, et al.
Published: (2025)
by: Çelikok, Mustafa Mert, et al.
Published: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
by: Preuveneers, Jack, et al.
Published: (2025)
by: Preuveneers, Jack, et al.
Published: (2025)
Learning Software Bug Reports: A Systematic Literature Review
by: Long, Guoming, et al.
Published: (2025)
by: Long, Guoming, et al.
Published: (2025)
Critical Insights into Leading Conversational AI Models
by: Kohli, Urja, et al.
Published: (2025)
by: Kohli, Urja, et al.
Published: (2025)
KemenkeuGPT: Leveraging a Large Language Model on Indonesia's Government Financial Data and Regulations to Enhance Decision Making
by: Febrian, Gilang Fajar, et al.
Published: (2024)
by: Febrian, Gilang Fajar, et al.
Published: (2024)
Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
by: Kaur, Rishemjit, et al.
Published: (2025)
by: Kaur, Rishemjit, et al.
Published: (2025)
The Paradox of Robustness: Decoupling Rule-Based Logic from Affective Noise in High-Stakes Decision-Making
by: Chun, Jon, et al.
Published: (2026)
by: Chun, Jon, et al.
Published: (2026)
Comparative Analysis of AI Agent Architectures for Entity Relationship Classification
by: Berijanian, Maryam, et al.
Published: (2025)
by: Berijanian, Maryam, et al.
Published: (2025)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
by: Raval, Shivam, et al.
Published: (2026)
by: Raval, Shivam, et al.
Published: (2026)
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
by: Hua, Yuncheng, et al.
Published: (2026)
by: Hua, Yuncheng, et al.
Published: (2026)
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
by: Kubica, Dominick, et al.
Published: (2025)
by: Kubica, Dominick, et al.
Published: (2025)
Active Context Compression: Autonomous Memory Management in LLM Agents
by: Verma, Nikhil
Published: (2026)
by: Verma, Nikhil
Published: (2026)
APP: Accelerated Path Patching with Task-Specific Pruning
by: Andersen, Frauke, et al.
Published: (2025)
by: Andersen, Frauke, et al.
Published: (2025)
Ask WhAI:Probing Belief Formation in Role-Primed LLM Agents
by: Moore, Keith, et al.
Published: (2025)
by: Moore, Keith, et al.
Published: (2025)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
by: Kohl, Jens, et al.
Published: (2024)
by: Kohl, Jens, et al.
Published: (2024)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
by: Bian, Zhipeng, et al.
Published: (2025)
by: Bian, Zhipeng, et al.
Published: (2025)
Review of Case-Based Reasoning for LLM Agents: Theoretical Foundations, Architectural Components, and Cognitive Integration
by: Hatalis, Kostas, et al.
Published: (2025)
by: Hatalis, Kostas, et al.
Published: (2025)
Heimdall: test-time scaling on the generative verification
by: Shi, Wenlei, et al.
Published: (2025)
by: Shi, Wenlei, et al.
Published: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
by: Blair-Stanek, Andrew, et al.
Published: (2023)
by: Blair-Stanek, Andrew, et al.
Published: (2023)
CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities
by: Zhu, Yuxuan, et al.
Published: (2025)
by: Zhu, Yuxuan, et al.
Published: (2025)
Similar Items
-
AI Predicts AGI: Leveraging AGI Forecasting and Peer Review to Explore LLMs' Complex Reasoning Capabilities
by: Davide, Fabrizio, et al.
Published: (2024) -
EBBS: An Ensemble with Bi-Level Beam Search for Zero-Shot Machine Translation
by: Wen, Yuqiao, et al.
Published: (2024) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025) -
Semantic Delta: An Interpretable Signal Differentiating Human and LLMs Dialogue
by: Scantamburlo, Riccardo, et al.
Published: (2026) -
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
by: Costa, Rimom
Published: (2025)