VerAs: Verify then Assess STEM Lab Reports
Fuente:
arXiv
Guardado en:
| Autores principales: | Atil, Berk, Karizaki, Mahsa Sheikhi, Passonneau, Rebecca J. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Something Just Like TRuST : Toxicity Recognition of Span and Target
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
por: Karizaki, Mahsa Sheikhi, et al.
Publicado: (2024)
por: Karizaki, Mahsa Sheikhi, et al.
Publicado: (2024)
Sociodemographic Bias in Language Models: A Survey and Forward Path
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026)
por: Atil, Berk, et al.
Publicado: (2026)
Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Chain of Simulation: A Dual-Mode Reasoning Framework for Large Language Models with Dynamic Problem Routing
por: Sheikhi, Saeid
Publicado: (2026)
por: Sheikhi, Saeid
Publicado: (2026)
Non-Determinism of "Deterministic" LLM Settings
por: Atil, Berk, et al.
Publicado: (2024)
por: Atil, Berk, et al.
Publicado: (2024)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
por: Gupta, Vipul, et al.
Publicado: (2024)
por: Gupta, Vipul, et al.
Publicado: (2024)
CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
por: Pan, Teng, et al.
Publicado: (2026)
por: Pan, Teng, et al.
Publicado: (2026)
Joint Training for Selective Prediction
por: Li, Zhaohui, et al.
Publicado: (2024)
por: Li, Zhaohui, et al.
Publicado: (2024)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
por: Seo, Wooseok, et al.
Publicado: (2025)
por: Seo, Wooseok, et al.
Publicado: (2025)
SCI-Verifier: Scientific Verifier with Thinking
por: Zheng, Shenghe, et al.
Publicado: (2025)
por: Zheng, Shenghe, et al.
Publicado: (2025)
Non-Halting Queries: Exploiting Fixed Points in LLMs
por: Hammouri, Ghaith, et al.
Publicado: (2024)
por: Hammouri, Ghaith, et al.
Publicado: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
por: Atil, Berk, et al.
Publicado: (2025)
por: Atil, Berk, et al.
Publicado: (2025)
Measuring Vision-Language STEM Skills of Neural Models
por: Shen, Jianhao, et al.
Publicado: (2024)
por: Shen, Jianhao, et al.
Publicado: (2024)
On the Ability of Transformers to Verify Plans
por: Sarrof, Yash, et al.
Publicado: (2026)
por: Sarrof, Yash, et al.
Publicado: (2026)
Mathematical Derivation Graphs: A Relation Extraction Task in STEM Manuscripts
por: Prasad, Vishesh, et al.
Publicado: (2024)
por: Prasad, Vishesh, et al.
Publicado: (2024)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
por: Guan, Xinyan, et al.
Publicado: (2024)
por: Guan, Xinyan, et al.
Publicado: (2024)
SLaNC: Static LayerNorm Calibration
por: Salmani, Mahsa, et al.
Publicado: (2024)
por: Salmani, Mahsa, et al.
Publicado: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
por: Hu, Haiquan, et al.
Publicado: (2025)
por: Hu, Haiquan, et al.
Publicado: (2025)
Towards Verifiable Text Generation with Symbolic References
por: Hennigen, Lucas Torroba, et al.
Publicado: (2023)
por: Hennigen, Lucas Torroba, et al.
Publicado: (2023)
Heuristics and Biases in AI Decision-Making: Implications for Responsible AGI
por: Saeedi, Payam, et al.
Publicado: (2024)
por: Saeedi, Payam, et al.
Publicado: (2024)
RLPR: Extrapolating RLVR to General Domains without Verifiers
por: Yu, Tianyu, et al.
Publicado: (2025)
por: Yu, Tianyu, et al.
Publicado: (2025)
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
por: Lai, Yuhang, et al.
Publicado: (2026)
por: Lai, Yuhang, et al.
Publicado: (2026)
References Improve LLM Alignment in Non-Verifiable Domains
por: Shi, Kejian, et al.
Publicado: (2026)
por: Shi, Kejian, et al.
Publicado: (2026)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
V-STaR: Training Verifiers for Self-Taught Reasoners
por: Hosseini, Arian, et al.
Publicado: (2024)
por: Hosseini, Arian, et al.
Publicado: (2024)
Don't Trust: Verify -- Grounding LLM Quantitative Reasoning with Autoformalization
por: Zhou, Jin Peng, et al.
Publicado: (2024)
por: Zhou, Jin Peng, et al.
Publicado: (2024)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
por: Nguyen, Hieu Trung, et al.
Publicado: (2026)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
por: Zha, Kaiwen, et al.
Publicado: (2025)
por: Zha, Kaiwen, et al.
Publicado: (2025)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
por: Pavlenko, Kirill, et al.
Publicado: (2026)
por: Pavlenko, Kirill, et al.
Publicado: (2026)
ProofSketch: Efficient Verified Reasoning for Large Language Models
por: Sheshanarayana, Disha, et al.
Publicado: (2025)
por: Sheshanarayana, Disha, et al.
Publicado: (2025)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
por: Zhao, Zheng, et al.
Publicado: (2025)
por: Zhao, Zheng, et al.
Publicado: (2025)
Do We Need Frontier Models to Verify Mathematical Proofs?
por: Naik, Aaditya, et al.
Publicado: (2026)
por: Naik, Aaditya, et al.
Publicado: (2026)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
por: Stojanovski, Zafir, et al.
Publicado: (2025)
por: Stojanovski, Zafir, et al.
Publicado: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
por: Liu, Yixin, et al.
Publicado: (2026)
por: Liu, Yixin, et al.
Publicado: (2026)
Masked-and-Reordered Self-Supervision for Reinforcement Learning from Verifiable Rewards
por: Wang, Zhen, et al.
Publicado: (2025)
por: Wang, Zhen, et al.
Publicado: (2025)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
por: Shen, Yiran, et al.
Publicado: (2025)
por: Shen, Yiran, et al.
Publicado: (2025)
Ejemplares similares
-
Something Just Like TRuST : Toxicity Recognition of Span and Target
por: Atil, Berk, et al.
Publicado: (2025) -
How Well Can You Articulate that Idea? Insights from Automated Formative Assessment
por: Karizaki, Mahsa Sheikhi, et al.
Publicado: (2024) -
Sociodemographic Bias in Language Models: A Survey and Forward Path
por: Gupta, Vipul, et al.
Publicado: (2023) -
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
por: Atil, Berk, et al.
Publicado: (2026) -
Model Unlearning Objectives Vary for Distinct Language Functions
por: Atil, Berk, et al.
Publicado: (2026)