text model

OpenAI logo

GPT-4

OpenAI

The generational jump that made LLMs viable for real work.

GPT-4 was the first model most companies trusted with production tasks. It is included here for the capability curve rather than as a recommendation — everything it does, cheaper models now do better.

CapabilitiesTool useVision (Turbo)JSON mode (Turbo)

Variants

3 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Base (8K)

Balanced
gpt-4

Original 8K context release.

rel. cost
58
quality

Turbo (128K)

High
gpt-4-turbo

Cheaper, faster, 128K context, vision.

rel. cost
64
quality

32K

Max
gpt-4-32k

Extended-context variant of the original.

rel. cost
58
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

48/100

Recruitment fit

Historic interest only — cheaper current models beat it on both quality and cost.

Best for

  • Nothing new. Migrate existing pipelines.

Strengths

  • Extremely well documented
  • Predictable behaviour

Watch out for

  • Costs more than GPT-5.6 while performing worse

The numbers

Measured, not asserted.

Published benchmarks

Vendor and independent figures.

MMLU
86.4%Published

Compare

Worth putting side by side.