LLM Evaluation Jobs is in Preview for W&B Multi-tenant Cloud. Compute is free during the preview period. Learn more
- If a benchmark has
truein the OpenAI Model Scorer column, the benchmark uses OpenAI models for scoring. An organization or team admin must add an OpenAI API key as a team secret. When you configure an evaluation job with a benchmark with this requirement, set the Scorer API key field to the secret. - If a benchmark has a link in the Gated Hugging Face Dataset column, the benchmark requires access to a gated Hugging Face dataset. An organization or team admin must request access to the dataset in Hugging Face, create a Hugging Face user access token, and configure a team secret with the access key. When you configure a benchmark with this requirement, set the Hugging Face Token field to the secret.
Knowledge
Evaluate factual knowledge across various domains like science, language, and general reasoning.Reasoning
Evaluate logical thinking, problem-solving, and common-sense reasoning capabilities.Math
Evaluate mathematical problem-solving at various difficulty levels, from grade school to competition-level problems.Code
Evaluate programming and software development capabilities like debugging, code execution prediction, and function calling.Reading
Evaluate reading comprehension and information extraction from complex texts.Long context
Evaluate the ability to process and reason over extended contexts, including retrieval and pattern recognition.Safety
Evaluate alignment, bias detection, harmful content resistance, and truthfulness.Domain-Specific
Evaluate specialized knowledge in medicine, chemistry, law, biology, and other professional fields.Multimodal
Evaluate vision and language understanding combining visual and textual inputs.Instruction following
Evaluate adherence to specific instructions and formatting requirements.System
Basic system validation and pre-flight checks.Next steps
- Evaluate a model checkpoint
- Evaluate a hosted API model
- View details about specific benchmarks at AISI Inspect Evals