text model

OpenAI logo

GPT-6

OpenAI

Frontier reasoning with an effort dial — from instant to deliberate.

GPT-6 is the current frontier of the GPT line. The headline change over GPT-5 is not raw benchmark score but controllability: a single effort parameter moves the same model from sub-second replies to multi-minute deliberate reasoning, so one integration covers both a live chat widget and an overnight batch job. For hiring teams that matters because screening and structured interview scoring have opposite latency budgets, and you no longer need two models to serve them.

Announced model. All pricing, benchmarks and evals are estimates — replace before this page is indexed.

CapabilitiesReasoning effort controlTool useStructured outputsVisionStreamingPrompt caching90+ languages

Variants

5 ways to run it.

Same family, different quality/cost rung. Picking the wrong one is the most common way teams overpay.

Max effort

Max
gpt-6-max

Longest deliberation. Reserve it for work where a wrong answer is expensive — final scoring, offer-risk analysis, contract review.

Final interview scoring · Complex multi-step analysis

42.0s
latency
24×
rel. cost
98
quality

High

High
gpt-6-high

The default for anything a human will act on. Strong multi-step reasoning without the max-effort latency tax.

Structured interview scoring · Long-document analysis

9.0s
latency
rel. cost
94
quality

Medium

Balanced
gpt-6-medium

The workhorse. Good judgement at a price that survives volume — this is the rung most production traffic should sit on.

CV parsing at volume · Drafting and summarisation

3.2s
latency
rel. cost
88
quality

Low

Fast
gpt-6-low

Near-instant. Skips deliberation entirely — right for classification, routing and anything a user is waiting on.

Live chat · Intent routing · Tagging

900ms
latency
1.4×
rel. cost
79
quality

Mini

Mini
gpt-6-mini

Smaller sibling. The cheapest way to run GPT-6-shaped prompts when the task is mechanical rather than judgemental.

Bulk extraction · Deduplication · Cheap pre-filters

600ms
latency
rel. cost
71
quality

Fit

Two audiences, one model.

The same model is a different proposition depending on what you point it at.

93/100

Recruitment fit

The strongest general model for hiring work — but only worth its price on the judgement-heavy steps, not on parsing.

Best for

  • Structured interview scoring against a rubric
  • Comparing a shortlist against a job spec
  • Drafting candidate feedback that a human will edit

Strengths

  • Holds a full interview transcript and a job spec in context without chunking
  • Effort dial lets live screening and overnight scoring share one integration
  • Structured outputs are reliable enough to write straight into an ATS record

Watch out for

  • Max effort is roughly 24× the cost of Mini — do not point bulk CV parsing at it
  • A frontier model is not a fairness guarantee: you still need audited rubrics and a human on the shortlist
  • Pricing is unpublished, so total cost of ownership is unmodellable today

The numbers

Measured, not asserted.

HumanLike evals

Run on our own harness.

CV field extraction (F1)
0.94Estimate

Projection. Not yet run against GPT-6.

Time to first token (medium effort)
640 msEstimate

Published benchmarks

Vendor and independent figures.

MMLU-Pro
TBDEstimate
SWE-bench Verified
TBDEstimate