scientific-llm-benchmarks

v2026.09.24

A comprehensive reference of benchmarks for evaluating large language models on scientific reasoning and discovery.

GitHub
安装命令
npx skhub add akillness/scientific-llm-benchmarks
Markdown
SKILL.md

Scientific LLM Benchmarks

This skill provides references to benchmarks used for evaluating large language models on scientific reasoning and discovery. The data comes from the Awesome-Scientific-LLM-Benchmarks repository.

Contents

The complete benchmark list is stored locally within this skill:

  • References List: references/benchmarks.md
  • Data (YAML format): data/benchmarks.yaml

Benchmark Domains Covered

  • General / Multi-domain Science: Cross-disciplinary STEM reasoning benchmarks.
  • Mathematics: Arithmetic, competition, olympiad, and frontier / formal-proof mathematics.
  • Physics and Astronomy: Physics olympiad, graduate physics, computational physics, and astronomy.
  • Chemistry: Molecular property, reaction, retrosynthesis, safety, and chemical knowledge.
  • Materials Science: Crystals, materials property prediction, and materials-science knowledge.
  • Biology and Life Sciences: Genomics, proteins, bioinformatics agents, protocols, and research biology.
  • Agentic Science and AI Research: LLM agents that write research code, run data analyses, attempt autonomous discovery, and conduct ML/AI research.

Helper Scripts

Also included are python scripts inside scripts/:

  • generate_readme.py: Regenerates the markdown tables and list from data/benchmarks.yaml.
  • fetch_examples.py: Fetches real sample rows from HuggingFace dataset URLs specified in the dataset metadata.
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

未指定

源路径

.agent-skills/scientific-llm-benchmarks

默认分支

main

最新提交

f579bfe

Tree SHA

34a09b3