Heimdall: test-time scaling on the generative verification
Fuente:
arXiv
Salvato in:
| Autori principali: | Shi, Wenlei, Jin, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Process Supervision-Guided Policy Optimization for Code Generation
di: Dai, Ning, et al.
Pubblicazione: (2024)
di: Dai, Ning, et al.
Pubblicazione: (2024)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
di: Jin, Ming
Pubblicazione: (2026)
di: Jin, Ming
Pubblicazione: (2026)
Curveball Steering: The Right Direction To Steer Isn't Always Linear
di: Raval, Shivam, et al.
Pubblicazione: (2026)
di: Raval, Shivam, et al.
Pubblicazione: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Identifying Bias in Machine-generated Text Detection
di: Stowe, Kevin, et al.
Pubblicazione: (2025)
di: Stowe, Kevin, et al.
Pubblicazione: (2025)
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Evaluating Steering Techniques using Human Similarity Judgments
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
di: Studdiford, Zach, et al.
Pubblicazione: (2025)
Reasoning-Based AI for Startup Evaluation (R.A.I.S.E.): A Memory-Augmented, Multi-Step Decision Framework
di: Preuveneers, Jack, et al.
Pubblicazione: (2025)
di: Preuveneers, Jack, et al.
Pubblicazione: (2025)
AI-Powered Annotation Pipelines for Stabilizing Large Language Models: A Human-AI Synergy Approach
di: Pathak, Gangesh, et al.
Pubblicazione: (2025)
di: Pathak, Gangesh, et al.
Pubblicazione: (2025)
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance
di: Ivanov, Igor
Pubblicazione: (2025)
di: Ivanov, Igor
Pubblicazione: (2025)
From Extraction to Synthesis: Entangled Heuristics for Agent-Augmented Strategic Reasoning
di: Ghisellini, Renato, et al.
Pubblicazione: (2025)
di: Ghisellini, Renato, et al.
Pubblicazione: (2025)
Do Language Models Mirror Human Confidence? Exploring Psychological Insights to Address Overconfidence in LLMs
di: Xu, Chenjun, et al.
Pubblicazione: (2025)
di: Xu, Chenjun, et al.
Pubblicazione: (2025)
LLM-based Automated Theorem Proving Hinges on Scalable Synthetic Data Generation
di: Lai, Junyu, et al.
Pubblicazione: (2025)
di: Lai, Junyu, et al.
Pubblicazione: (2025)
The Unified Cognitive Consciousness Theory for Language Models: Anchoring Semantics, Thresholds of Activation, and Emergent Reasoning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
Learning Temporal Abstractions via Variational Homomorphisms in Option-Induced Abstract MDPs
di: Li, Chang, et al.
Pubblicazione: (2025)
di: Li, Chang, et al.
Pubblicazione: (2025)
SOCIA-$\nabla$: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
di: Hua, Yuncheng, et al.
Pubblicazione: (2025)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
di: Danilevsky, Marina, et al.
Pubblicazione: (2025)
Evaluating LLM Metrics Through Real-World Capabilities
di: Miller, Justin K, et al.
Pubblicazione: (2025)
di: Miller, Justin K, et al.
Pubblicazione: (2025)
A Fuzzy Logic Prompting Framework for Large Language Models in Adaptive and Uncertain Tasks
di: Figueiredo, Vanessa
Pubblicazione: (2025)
di: Figueiredo, Vanessa
Pubblicazione: (2025)
Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Report
di: Chen, Yan, et al.
Pubblicazione: (2025)
di: Chen, Yan, et al.
Pubblicazione: (2025)
PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data
di: Xiong, Kai, et al.
Pubblicazione: (2025)
di: Xiong, Kai, et al.
Pubblicazione: (2025)
Argumentatively Coherent Judgmental Forecasting
di: Gorur, Deniz, et al.
Pubblicazione: (2025)
di: Gorur, Deniz, et al.
Pubblicazione: (2025)
SagaLLM: Context Management, Validation, and Transaction Guarantees for Multi-Agent LLM Planning
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
di: Chang, Edward Y., et al.
Pubblicazione: (2025)
MapAgent: A Hierarchical Agent for Geospatial Reasoning with Dynamic Map Tool Integration
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2025)
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2025)
Learning Efficient Guardrails for Compliance
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
di: Wen, Xiaofei, et al.
Pubblicazione: (2025)
Unlocking the Wisdom of Large Language Models: An Introduction to The Path to Artificial General Intelligence
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
Understanding LLM Evaluator Behavior: A Structured Multi-Evaluator Framework for Merchant Risk Assessment
di: Wang, Liang, et al.
Pubblicazione: (2026)
di: Wang, Liang, et al.
Pubblicazione: (2026)
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
di: Niu, Guanglin, et al.
Pubblicazione: (2026)
di: Niu, Guanglin, et al.
Pubblicazione: (2026)
Quantifying Self-Preservation Bias in Large Language Models
di: Migliarini, Matteo, et al.
Pubblicazione: (2026)
di: Migliarini, Matteo, et al.
Pubblicazione: (2026)
Evaluating Relational Reasoning in LLMs with REL
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
di: Fesser, Lukas, et al.
Pubblicazione: (2026)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
di: Le, Linh, et al.
Pubblicazione: (2026)
di: Le, Linh, et al.
Pubblicazione: (2026)
KACE: Knowledge-Adaptive Context Engineering for Mathematical Reasoning
di: Parashar, Jayant, et al.
Pubblicazione: (2026)
di: Parashar, Jayant, et al.
Pubblicazione: (2026)
RAudit: A Blind Auditing Protocol for Large Language Model Reasoning
di: Chang, Edward Y., et al.
Pubblicazione: (2026)
di: Chang, Edward Y., et al.
Pubblicazione: (2026)
SciDMT: A Large-Scale Corpus for Detecting Scientific Mentions
di: Pan, Huitong, et al.
Pubblicazione: (2024)
di: Pan, Huitong, et al.
Pubblicazione: (2024)
Towards ethical multimodal systems
di: Roger, Alexis, et al.
Pubblicazione: (2023)
di: Roger, Alexis, et al.
Pubblicazione: (2023)
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
di: Fang, Bowen, et al.
Pubblicazione: (2026)
di: Fang, Bowen, et al.
Pubblicazione: (2026)
PersistBench: When Should Long-Term Memories Be Forgotten by LLMs?
di: Pulipaka, Sidharth, et al.
Pubblicazione: (2026)
di: Pulipaka, Sidharth, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Process Supervision-Guided Policy Optimization for Code Generation
di: Dai, Ning, et al.
Pubblicazione: (2024) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025) -
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025) -
What Makes a Good AI Review? Concern-Level Diagnostics for AI Peer Review
di: Jin, Ming
Pubblicazione: (2026) -
Curveball Steering: The Right Direction To Steer Isn't Always Linear
di: Raval, Shivam, et al.
Pubblicazione: (2026)