ChemPro: A Progressive Chemistry Benchmark for Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Baranwal, Aaditya, Vyas, Shruti |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
di: Dhole, Kaustubh D.
Pubblicazione: (2026)
di: Dhole, Kaustubh D.
Pubblicazione: (2026)
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
di: Venkatasubramanian, Venkat, et al.
Pubblicazione: (2024)
di: Venkatasubramanian, Venkat, et al.
Pubblicazione: (2024)
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
di: Du, Bangde, et al.
Pubblicazione: (2025)
di: Du, Bangde, et al.
Pubblicazione: (2025)
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
di: Deng, Minghang, et al.
Pubblicazione: (2025)
di: Deng, Minghang, et al.
Pubblicazione: (2025)
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
di: Williamson, Dane, et al.
Pubblicazione: (2025)
di: Williamson, Dane, et al.
Pubblicazione: (2025)
LLMs and the Human Condition
di: Wallis, Peter
Pubblicazione: (2024)
di: Wallis, Peter
Pubblicazione: (2024)
MATH-PT: A Math Reasoning Benchmark for European and Brazilian Portuguese
di: Teixeira, Tiago, et al.
Pubblicazione: (2026)
di: Teixeira, Tiago, et al.
Pubblicazione: (2026)
Next Token Prediction Is a Dead End for Creativity
di: Olatunji, Ibukun, et al.
Pubblicazione: (2025)
di: Olatunji, Ibukun, et al.
Pubblicazione: (2025)
Automated Circuit Interpretation via Probe Prompting
di: Birardi, Giuseppe
Pubblicazione: (2025)
di: Birardi, Giuseppe
Pubblicazione: (2025)
QuickSilver -- Speeding up LLM Inference through Dynamic Token Halting, KV Skipping, Contextual Token Fusion, and Adaptive Matryoshka Quantization
di: Khanna, Danush, et al.
Pubblicazione: (2025)
di: Khanna, Danush, et al.
Pubblicazione: (2025)
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax?
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
di: Blair-Stanek, Andrew, et al.
Pubblicazione: (2023)
Graph Language Models
di: Plenz, Moritz, et al.
Pubblicazione: (2024)
di: Plenz, Moritz, et al.
Pubblicazione: (2024)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
di: Taskooh, Hamidreza Kazemi, et al.
Pubblicazione: (2026)
di: Taskooh, Hamidreza Kazemi, et al.
Pubblicazione: (2026)
Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
di: Baxi, Rahul
Pubblicazione: (2025)
di: Baxi, Rahul
Pubblicazione: (2025)
Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
di: Santos, Matheus L. O., et al.
Pubblicazione: (2023)
di: Santos, Matheus L. O., et al.
Pubblicazione: (2023)
A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
di: Adjovi, Mahounan Pericles, et al.
Pubblicazione: (2026)
The Democratic Paradox in Large Language Models' Underestimation of Press Freedom
di: Loaiza, I., et al.
Pubblicazione: (2025)
di: Loaiza, I., et al.
Pubblicazione: (2025)
The Company You Keep: How LLMs Respond to Dark Triad Traits
di: Lu, Zeyi, et al.
Pubblicazione: (2026)
di: Lu, Zeyi, et al.
Pubblicazione: (2026)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
Large Language Models Report Subjective Experience Under Self-Referential Processing
di: Berg, Cameron, et al.
Pubblicazione: (2025)
di: Berg, Cameron, et al.
Pubblicazione: (2025)
Discovering Differences in Strategic Behavior Between Humans and LLMs
di: Wang, Caroline, et al.
Pubblicazione: (2026)
di: Wang, Caroline, et al.
Pubblicazione: (2026)
Behavioural vs. Representational Systematicity in End-to-End Models: An Opinionated Survey
di: Vegner, Ivan, et al.
Pubblicazione: (2025)
di: Vegner, Ivan, et al.
Pubblicazione: (2025)
Self-Supervised Borrowing Detection on Multilingual Wordlists
di: Wientzek, Tim
Pubblicazione: (2025)
di: Wientzek, Tim
Pubblicazione: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
Chatbots put to the test in math and logic problems: A preliminary comparison and assessment of ChatGPT-3.5, ChatGPT-4, and Google Bard
di: Plevris, Vagelis, et al.
Pubblicazione: (2023)
di: Plevris, Vagelis, et al.
Pubblicazione: (2023)
Incentives or Ontology? A Structural Rebuttal to OpenAI's Hallucination Thesis
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
di: Ackermann, Richard, et al.
Pubblicazione: (2025)
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inference
di: Saito, Kei
Pubblicazione: (2026)
di: Saito, Kei
Pubblicazione: (2026)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation
di: Edin, Joakim, et al.
Pubblicazione: (2025)
di: Edin, Joakim, et al.
Pubblicazione: (2025)
Prompt Tuned Embedding Classification for Multi-Label Industry Sector Allocation
di: Buchner, Valentin Leonhard, et al.
Pubblicazione: (2023)
di: Buchner, Valentin Leonhard, et al.
Pubblicazione: (2023)
NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation
di: Saito, Kei
Pubblicazione: (2025)
di: Saito, Kei
Pubblicazione: (2025)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
di: Xing, Eric, et al.
Pubblicazione: (2024)
di: Xing, Eric, et al.
Pubblicazione: (2024)
CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2025)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
di: Peng, Bo, et al.
Pubblicazione: (2025)
di: Peng, Bo, et al.
Pubblicazione: (2025)
ACCORD: Closing the Commonsense Measurability Gap
di: Roewer-Després, François, et al.
Pubblicazione: (2024)
di: Roewer-Després, François, et al.
Pubblicazione: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
di: Dumitru, Razvan-Gabriel, et al.
Pubblicazione: (2024)
Predictive Simultaneous Interpretation: Harnessing Large Language Models for Democratizing Real-Time Multilingual Communication
di: Iida, Kurando, et al.
Pubblicazione: (2024)
di: Iida, Kurando, et al.
Pubblicazione: (2024)
Documenti analoghi
-
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
di: Dhole, Kaustubh D.
Pubblicazione: (2026) -
Quo Vadis ChatGPT? From Large Language Models to Large Knowledge Models
di: Venkatasubramanian, Venkat, et al.
Pubblicazione: (2024) -
TwinVoice: A Multi-dimensional Benchmark Towards Digital Twins via LLM Persona Simulation
di: Du, Bangde, et al.
Pubblicazione: (2025) -
ReFoRCE: A Text-to-SQL Agent with Self-Refinement, Consensus Enforcement, and Column Exploration
di: Deng, Minghang, et al.
Pubblicazione: (2025) -
Syntactic Blind Spots: How Misalignment Leads to LLMs Mathematical Errors
di: Williamson, Dane, et al.
Pubblicazione: (2025)