EleutherAI LM Evaluation Harness | Generated by AI

Home 2026.07

Question: intro to EleutherAI/lm-evaluation-harness

Answer:

EleutherAI’s lm-evaluation-harness is one of the standard tools for measuring LLM quality. Think of it as the pytest + benchmark runner for language models. It gives researchers a reproducible way to answer:

“Is my model actually better, or does it just feel better?”

(GitHub)

What problem does it solve?

Before tools like this, every paper had custom evaluation scripts:

model.py
eval.py
prompt_template.py
score.py

Hard to compare.

lm-evaluation-harness standardizes:

Model
  ↓
Tokenizer / prompt format
  ↓
Benchmark task
  ↓
Inference
  ↓
Metric calculation
  ↓
JSON result

Example:

Qwen3-8B
    |
    v
MMLU
    |
    v
accuracy = 78.5%

Now another researcher can run the same setup and compare.

It powers many public LLM leaderboards, including the Hugging Face Open LLM Leaderboard workflow. (GitHub)


Supported benchmarks

Common tasks:

It has dozens of benchmark families and hundreds of subtasks. (GitHub)


Basic usage

Install:

git clone https://github.com/EleutherAI/lm-evaluation-harness
cd lm-evaluation-harness

pip install -e .
pip install "lm_eval[hf]"

Run GPT-2 on HellaSwag:

lm_eval run \
  --model hf \
  --model_args pretrained=gpt2 \
  --tasks hellaswag

(GitHub)

Run a modern local model:

lm_eval run \
  --model hf \
  --model_args pretrained=Qwen/Qwen3-8B \
  --tasks mmlu,gsm8k \
  --device cuda:0 \
  --batch_size auto

For your RTX 4070 workflow

This is very relevant for your experiments:

You trained:

GPT-2 124M
SEC dataset
RTX 4070
loss ~2.28

The next step is not only loss. You want downstream evaluation:

checkpoint
    |
    v
lm-evaluation-harness
    |
    +--> lambada
    +--> hellaswag
    +--> arc_easy
    +--> piqa
    +--> winogrande

Example:

lm_eval run \
  --model hf \
  --model_args pretrained=./my-gpt2-checkpoint \
  --tasks hellaswag,arc_easy \
  --device cuda:0

This tells you whether training improved actual language ability.


Architecture idea

The code structure is roughly:

lm_eval/

├── models/
│   ├── hf.py          # HuggingFace models
│   ├── vllm.py        # vLLM backend
│   └── api_models.py  # OpenAI compatible APIs
│
├── tasks/
│   ├── mmlu/
│   ├── gsm8k/
│   └── hellaswag/
│
├── evaluator.py
└── metrics/

The key abstraction:

class LM:
    def loglikelihood()
    def generate_until()

Every backend implements this interface.

Then benchmarks do not care whether the model is:

transformers
vLLM
OpenAI API
DeepSpeed
Megatron

They only call:

model.generate_until(prompt)

Custom benchmark

You can create your own task:

tasks/
  my_company_eval/
      task.yaml

Example:

task: my_eval

dataset_path: json
dataset_kwargs:
  data_files:
    test: test.json

output_type: multiple_choice

doc_to_text: "{{question}}"

doc_to_choice:
  - A
  - B
  - C
  - D

metric_list:
  - metric: acc

The YAML-driven design makes experiments reproducible. (GitHub)


Difference from other tools

Tool Purpose
lm-evaluation-harness Academic/general LLM benchmark
vLLM benchmark Throughput/latency
OpenCompass Large evaluation platform
DeepEval Application-level testing
SWE-bench Coding agent evaluation

For model training research, lm-eval is probably the first tool to add after training.

Your current stack could become:

nanoGPT / PyTorch
        |
        v
checkpoint
        |
        v
lm-evaluation-harness
        |
        v
benchmark report
        |
        v
W&B / blog / paper

References:


Generated by AI. Curating and sharing still takes effort. If you find it useful, feel free to donate. WeChat: @lzwjavaWeChat QR · X: @lzwjava · Say hi 👋

Back Donate