scientific-llm-benchmarks

v2026.09.24

A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery.

GitHub
Install command
npx skhub add akillness/scientific-llm-benchmarks
Markdown
SKILL.md

Scientific LLM Benchmarks

This skill provides references to benchmarks used for evaluating large language models on scientific reasoning and discovery. The data comes from the Awesome-Scientific-LLM-Benchmarks repository.

Contents

The complete benchmark list is stored locally within this skill:

  • References List: references/benchmarks.md
  • Data (YAML format): data/benchmarks.yaml

Benchmark Domains Covered

  • General / Multi-domain Science: Cross-disciplinary STEM reasoning benchmarks.
  • Mathematics: Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.
  • Physics and Astronomy: Physics olympiad, graduate physics, computational physics, and astronomy.
  • Chemistry: Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.
  • Materials Science: Crystals, materials property prediction, and materials-science knowledge.
  • Biology and Life Sciences: Genomics, proteins, bioinformatics agents, protocols, and research biology.
  • Agentic Science and AI Research: LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.

Helper Scripts

Also included are python scripts inside scripts/:

  • generate_readme.py: Regenerates the markdown tables and list from data/benchmarks.yaml.
  • fetch_examples.py: Fetches real sample rows from HuggingFace dataset URLs specified in the dataset metadata.
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Not specified

Source path

.agent-skills/scientific-llm-benchmarks

Default branch

main

Latest commit

f579bfe

Tree SHA

34a09b3