Model guide
Measured results and catalogue overview for the 20-model roster. Benchmark closed 12 June 2026 (credits-only; the Regional Standard probe ran and was refused for all three candidate models). Sources: models-regions.json · benchmark-results.json · model-benchmark-report.md.
Headline findings
On Microsoft Fabric notebook and pipeline authoring tasks, our benchmark found the mid-tier gpt-4.1-mini matched flagship-class alternatives at roughly one-sixth the cost — £0.0010 per accepted PR against £0.0063 for gpt-4o (2024-11-20). Both models reached 100% acceptance on the 10-task held-out suite (5 notebooks, 5 pipelines; simple, medium and complex). For format-constrained, structured-output generation, the cost premium of a larger model is not evidenced in this task shape.
The gpt-4o supervisor reached a 60% catch-rate on a harder 10-defect mix (secret strings, wrong workspace references, missing error handling, schema mismatches) with a 0% false-block rate on 10 clean items. Security defects and structural errors were all caught. Schema and error-handling defects were the gap: schema mismatches and missing error-handling patterns were missed in this run. Candidate improvements include richer task-spec grounding in the supervisor prompt.
Supervisor defect-type gap map — gpt-4o, n=10 defects
- Demonstration environment only. This benchmark ran against synthetic, internal Fabric authoring tasks (Gate 5 — no client data). Results are directional only and are not production performance guarantees.
- Small sample (n=10). Both sweeps used a 10-task held-out suite. Variance is high; confidence intervals are wide. Rankings are directional, not definitive.
- Two models measured — benchmark closed. gpt-4.1-mini (authoring) and gpt-4o 2024-11-20 (authoring + supervisor) were the only models deployable on the Onyx02 Azure Sponsorship subscription. A second probe (12 June 2026) attempted Regional Standard deployments for gpt-4.1-nano, o4-mini and o3-mini — all three were refused by the Azure control plane (
InvalidResourceProperties). Per Leon's instruction (2026-06-12), the benchmark is closed on this measured set; PAYG and MCPP routes are cancelled. - Remaining models deferred to client-tenant benchmark. gpt-4.1, gpt-4.1-nano, gpt-5 family, o4-mini, o3-mini and others are deferred to the first client-tenant deployment, where normal quota applies and spend qualifies as ACR via PAL-linkage.
- Pricing [VERIFY]. All indicative price bands are drawn from the Azure Retail Prices REST API (productName = 'Azure OpenAI' / 'Azure OpenAI GPT5' / 'Azure OpenAI Reasoning', Accessed: 12 June 2026). Rates are pay-as-you-go GlobalStandard tier; USD/GBP at 0.79. Source: Microsoft Azure (2026) Azure OpenAI Service pricing. Available at: azure.microsoft.com/pricing/details/cognitive-services/openai-service (Accessed: 12 June 2026).
- Region availability [VERIFY]. Regional Standard model availability in uksouth drawn from Microsoft Learn (2026) Azure AI Foundry models sold directly by Azure — region availability. Available at: learn.microsoft.com (Accessed: 12 June 2026). Region lists may change; verify before deployment.
Full catalogue — 20 models
Model catalogue & benchmark results
| Model | Family / tier | GA status | Indicative price band | Benchmark | Regions | Notes |
|---|---|---|---|---|---|---|
| gpt-4.1 2025-04-14 | gpt-4.1 / Flagship | GA | $0.002 / $0.008 per 1k [VERIFY] | Deferred — client-tenant | Global + Regional (not uksouth Regional) | 1M context; Global Standard only from uksouth. Provisioned Managed Regional supports uksouth. |
| gpt-4.1-mini 2025-04-14 | gpt-4.1 / Cost-optimised | GA | $0.0004 / $0.0016 per 1k [VERIFY] | MeasuredAuthoring 100% acceptance · £0.0010/PR · mean latency 1.8s | UK-resident (uksouth Regional Std) | Only gpt-4.1 family model with uksouth Regional Standard support. Best-value authoring on this task shape. |
| gpt-4.1-nano 2025-04-14 | gpt-4.1 / Fastest, cheapest | GA | $0.0001 / $0.0004 per 1k [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Highest-throughput, lowest-cost option. Regional Standard probe refused on Onyx02 Sponsorship (InvalidResourceProperties). Deferred to client-tenant. |
| gpt-4o 2024-11-20 | gpt-4o / Flagship | GA | $0.0025 / $0.010 per 1k [VERIFY] | MeasuredAuthoring 100% acceptance · £0.0063/PR · supervisor 60% catch / 0% false-block | UK-resident (uksouth Regional Std) | Current pod default supervisor. Security defects fully caught; schema/error-handling gaps (see gap map). Approaching end-of-stable — prefer gpt-4.1 family for new authoring. |
| gpt-4o 2024-08-06 | gpt-4o / Flagship | GA | $0.0025 / $0.010 per 1k [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Earlier GA version of gpt-4o; prefer 2024-11-20 or newer. |
| gpt-4o 2024-05-13 | gpt-4o / Flagship | GA | [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Original GA release; prefer newer versions. |
| gpt-4o-mini 2024-07-18 | gpt-4o / Lightweight | Retiring 2025-09-15 | [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Retiring — use gpt-4.1-nano for the equivalent role in new deployments. |
| gpt-5 2025-08-07 | gpt-5 / Flagship | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | 400K context. Registration required. Global Standard only from uksouth resource. |
| gpt-5-mini 2025-08-07 | gpt-5 / Cost-efficient | GA | $0.00025 / $0.002 per 1k [VERIFY] | Deferred — client-tenant | Global Standard only | 400K context. Quota blocked on Onyx02 at benchmark time. |
| gpt-5-nano 2025-08-07 | gpt-5 / Fastest | GA | $0.00005 / $0.0004 per 1k [VERIFY] | Deferred — client-tenant | Global Standard only | 400K context; cheapest GPT-5 option. |
| gpt-5.1 2025-11-13 | gpt-5 / Iterative | GA | $0.00125 / $0.010 per 1k [VERIFY] | Deferred — client-tenant | Global Standard only | Iterative improvement on gpt-5. 400K context. |
| gpt-5.2 2025-12-11 | gpt-5 / Reliability | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | Focused on reliability and supervision tasks. 400K context. |
| gpt-5.3-codex 2026-02-24 | gpt-5 / Code-specialised | GA | $0.00175 / $0.014 per 1k [VERIFY] | Deferred — client-tenant | Global Standard only | Code-specialised variant. 400K context. |
| gpt-5.4 2026-03-05 | gpt-5 / Extended-context flagship | GA | $0.0025 / $0.015 per 1k [VERIFY] | Deferred — client-tenant | Global Standard only | 1M+ context; highest capability in the gpt-5 family. |
| gpt-5.4-mini 2026-03-17 | gpt-5 / Mini | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | Mini variant of gpt-5.4. 400K context. |
| gpt-5.4-nano 2026-03-17 | gpt-5 / Nano | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | Nano variant of gpt-5.4. Fastest/cheapest in the 5.4 sub-family. |
| o4-mini 2025-04-16 | o-series / Reasoning | GA | $0.0011 / $0.0044 per 1k [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Fast reasoning; strong on code and logic. Recommended o-series entry point. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Highest-priority candidate for client-tenant supervisor benchmark. |
| o3 2025-04-16 | o-series / Full reasoning | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | Highest o-series capability; registration required. Best for complex review and supervisor roles. |
| o3-mini 2025-01-31 | o-series / Lightweight reasoning | GA | [VERIFY] | Deferred — client-tenant | UK-resident (uksouth Regional Std) | Balanced speed/cost for reasoning tasks. |
| o1 2024-12-17 | o-series / Original full reasoning | GA | [VERIFY] | Deferred — client-tenant | Global Standard only | Predecessor to o3; 200K context. Global Standard only. |
| model-router 2025-05-19 | model-router / Routing | GA | [VERIFY] | Deferred — client-tenant | Global Std only — not uksouth | Routes prompts to an underlying model. Not available in uksouth. Limited to eastus, eastus2, westus3, swedencentral. |
| claude-opus-4 | claude-anthropic / Flagship | Preview (partner) | [VERIFY] | Deferred — client-tenant | eastus2 / swedencentral only | Not deployable from uksouth. Requires a Foundry hub in eastus2 or swedencentral; pay-as-you-go billing (Azure Sponsorship may not qualify). |
| claude-sonnet-4 | claude-anthropic / Balanced | Preview (partner) | [VERIFY] | Deferred — client-tenant | eastus2 / swedencentral only | Not deployable from uksouth. Same regional restriction as claude-opus-4. |
| DeepSeek-R1 | deepseek / Reasoning | GA (partner serverless) | [VERIFY] | Deferred — client-tenant | Global Standard (uksouth supported) | Available via Azure Marketplace serverless; uksouth resource supported. Open-weights provenance. |
| Phi-4 | phi-microsoft / SLM | GA (serverless) | [VERIFY] | Deferred — client-tenant | US regions + swedencentral only | Not available in uksouth. Serverless only. |
Benchmark run IDs: 32caf9d8b04b (initial) + credits-regional-2026-06-12 (probe) · Benchmark closed 12 June 2026 · n=10 tasks per sweep · Pricing: Azure Retail Prices REST API (Accessed: 12 June 2026) · Region: Microsoft Learn — Azure AI Foundry models sold directly by Azure — region availability (Accessed: 12 June 2026) · [VERIFY] = confirm current rates before deployment.
Regional models — Regional Standard, uksouth
When you select Regional (single-region processing preference) in the configurator, the model menu narrows to models confirmed available in uksouth Regional Standard deployment.
Regional model set — 6 models, uksouth Regional Standard (preference, not portal-enforced)
| Model | Version | Family | Deployment type | Notes |
|---|---|---|---|---|
| gpt-4.1-mini | 2025-04-14 | gpt-4.1 | Regional Standard — uksouth | Benchmark run: 100% authoring acceptance, £0.0010/PR. |
| gpt-4.1-nano | 2025-04-14 | gpt-4.1 | Regional Standard — uksouth | Highest throughput, lowest cost. Deferred to client-tenant benchmark. |
| gpt-4o | 2024-11-20 | gpt-4o | Regional Standard — uksouth | Benchmark run: supervisor 60% catch / 0% false-block. |
| gpt-4o | 2024-08-06 | gpt-4o | Regional Standard — uksouth | Earlier version; prefer 2024-11-20. Deferred to client-tenant benchmark. |
| o3-mini | 2025-01-31 | o-series | Regional Standard — uksouth | Reasoning model; balanced speed/cost. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Deferred to client-tenant benchmark. |
| o4-mini | 2025-04-16 | o-series | Regional Standard — uksouth | Fast reasoning; recommended o-series entry point. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Highest-priority supervisor candidate for client-tenant benchmark. |
Note: gpt-4o-mini (2024-07-18) is also available in uksouth Regional Standard but is retiring 2025-09-15 — excluded from new deployment recommendations. Region data: Microsoft Learn (Accessed: 12 June 2026).