VerAs: Verify then Assess STEM Lab Reports
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Atil, Berk, Karizaki, Mahsa Sheikhi, Passonneau, Rebecca J. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Something Just Like TRuST : Toxicity Recognition of Span and Target
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
von: Karizaki, Mahsa Sheikhi, et al.
Veröffentlicht: (2024)
von: Karizaki, Mahsa Sheikhi, et al.
Veröffentlicht: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
von: Atil, Berk, et al.
Veröffentlicht: (2026)
von: Atil, Berk, et al.
Veröffentlicht: (2026)
Model Unlearning Objectives Vary for Distinct Language Functions
von: Atil, Berk, et al.
Veröffentlicht: (2026)
von: Atil, Berk, et al.
Veröffentlicht: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
von: Sheikhi, Saeid
Veröffentlicht: (2026)
von: Sheikhi, Saeid
Veröffentlicht: (2026)
Non-Determinism of "Deterministic" LLM Settings
von: Atil, Berk, et al.
Veröffentlicht: (2024)
von: Atil, Berk, et al.
Veröffentlicht: (2024)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
von: Gupta, Vipul, et al.
Veröffentlicht: (2023)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
von: Gupta, Vipul, et al.
Veröffentlicht: (2024)
von: Gupta, Vipul, et al.
Veröffentlicht: (2024)
CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
von: Pan, Teng, et al.
Veröffentlicht: (2026)
von: Pan, Teng, et al.
Veröffentlicht: (2026)
Joint Training for Selective Prediction
von: Li, Zhaohui, et al.
Veröffentlicht: (2024)
von: Li, Zhaohui, et al.
Veröffentlicht: (2024)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
SCI-Verifier: Scientific Verifier with Thinking
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
Non-Halting Queries: Exploiting Fixed Points in LLMs
von: Hammouri, Ghaith, et al.
Veröffentlicht: (2024)
von: Hammouri, Ghaith, et al.
Veröffentlicht: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
Measuring Vision-Language STEM Skills of Neural Models
von: Shen, Jianhao, et al.
Veröffentlicht: (2024)
von: Shen, Jianhao, et al.
Veröffentlicht: (2024)
On the Ability of Transformers to Verify Plans
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
von: Sarrof, Yash, et al.
Veröffentlicht: (2026)
Mathematical Derivation Graphs: A Relation Extraction Task in STEM Manuscripts
von: Prasad, Vishesh, et al.
Veröffentlicht: (2024)
von: Prasad, Vishesh, et al.
Veröffentlicht: (2024)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
von: Guan, Xinyan, et al.
Veröffentlicht: (2024)
SLaNC: Static LayerNorm Calibration
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
von: Salmani, Mahsa, et al.
Veröffentlicht: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
Towards Verifiable Text Generation with Symbolic References
von: Hennigen, Lucas Torroba, et al.
Veröffentlicht: (2023)
von: Hennigen, Lucas Torroba, et al.
Veröffentlicht: (2023)
Heuristics and Biases in AI Decision-Making: Implications for Responsible AGI
von: Saeedi, Payam, et al.
Veröffentlicht: (2024)
von: Saeedi, Payam, et al.
Veröffentlicht: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
von: Lai, Yuhang, et al.
Veröffentlicht: (2026)
References Improve LLM Alignment in Non-Verifiable Domains
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
von: Shi, Kejian, et al.
Veröffentlicht: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
von: Zhou, Jin Peng, et al.
Veröffentlicht: (2024)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zha, Kaiwen, et al.
Veröffentlicht: (2025)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
von: Pavlenko, Kirill, et al.
Veröffentlicht: (2026)
von: Pavlenko, Kirill, et al.
Veröffentlicht: (2026)
ProofSketch: Efficient Verified Reasoning for Large Language Models
von: Sheshanarayana, Disha, et al.
Veröffentlicht: (2025)
von: Sheshanarayana, Disha, et al.
Veröffentlicht: (2025)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zheng, et al.
Veröffentlicht: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
von: Naik, Aaditya, et al.
Veröffentlicht: (2026)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
von: Liu, Yixin, et al.
Veröffentlicht: (2026)
Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
von: Wang, Zhen, et al.
Veröffentlicht: (2025)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
von: Shen, Yiran, et al.
Veröffentlicht: (2025)
von: Shen, Yiran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Something Just Like TRuST : Toxicity Recognition of Span and Target
von: Atil, Berk, et al.
Veröffentlicht: (2025) -
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
von: Karizaki, Mahsa Sheikhi, et al.
Veröffentlicht: (2024) -
Sociodemographic Bias in Language Models: A Survey and Forward Path
von: Gupta, Vipul, et al.
Veröffentlicht: (2023) -
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
von: Atil, Berk, et al.
Veröffentlicht: (2026) -
Model Unlearning Objectives Vary for Distinct Language Functions
von: Atil, Berk, et al.
Veröffentlicht: (2026)