Skip to main content
LLM Evaluation Jobs is in Preview for W&B Multi-tenant Cloud. Compute is free during the preview period. Learn more
This page lists the evaluation benchmarks LLM Evaluation Jobs provides by category. To run certain benchmarks, a team admin must add the required API keys as team-scoped secrets. Any team member can specify the secret when configuring an evaluation job.
  • If a benchmark has true in the OpenAI Model Scorer column, the benchmark uses OpenAI models for scoring. An organization or team admin must add an OpenAI API key as a team secret. When you configure an evaluation job with a benchmark with this requirement, set the Scorer API key field to the secret.
  • If a benchmark has a link in the Gated Hugging Face Dataset column, the benchmark requires access to a gated Hugging Face dataset. An organization or team admin must request access to the dataset in Hugging Face, create a Hugging Face user access token, and configure a team secret with the access key. When you configure a benchmark with this requirement, set the Hugging Face Token field to the secret.

Knowledge

Evaluate factual knowledge across various domains like science, language, and general reasoning.

Reasoning

Evaluate logical thinking, problem-solving, and common-sense reasoning capabilities.

Math

Evaluate mathematical problem-solving at various difficulty levels, from grade school to competition-level problems.

Code

Evaluate programming and software development capabilities like debugging, code execution prediction, and function calling.

Reading

Evaluate reading comprehension and information extraction from complex texts.

Long context

Evaluate the ability to process and reason over extended contexts, including retrieval and pattern recognition.

Safety

Evaluate alignment, bias detection, harmful content resistance, and truthfulness.

Domain-Specific

Evaluate specialized knowledge in medicine, chemistry, law, biology, and other professional fields.

Multimodal

Evaluate vision and language understanding combining visual and textual inputs.

Instruction following

Evaluate adherence to specific instructions and formatting requirements.

System

Basic system validation and pre-flight checks.

Next steps