Claim level Exploring — demonstration environment; indicative findings only.
V1
This is a configuration wizard, not a self-service deployment. Completing the wizard produces a priced enquiry the Onyx team reviews — it is not a binding quote or contract. Regulated estates and strict data-residency requirements are handled directly by Onyx at deployment scoping. Go to the configurator.
Back to configurator
ACC-09 Swarm · Benchmark

Model guide

Measured results and catalogue overview for the 20-model roster. Benchmark closed 12 June 2026 (credits-only; the Regional Standard probe ran and was refused for all three candidate models). Sources: models-regions.json · benchmark-results.json · model-benchmark-report.md.

Headline findings

Authoring cost & quality

On Microsoft Fabric notebook and pipeline authoring tasks, our benchmark found the mid-tier gpt-4.1-mini matched flagship-class alternatives at roughly one-sixth the cost — £0.0010 per accepted PR against £0.0063 for gpt-4o (2024-11-20). Both models reached 100% acceptance on the 10-task held-out suite (5 notebooks, 5 pipelines; simple, medium and complex). For format-constrained, structured-output generation, the cost premium of a larger model is not evidenced in this task shape.

Supervisor catch-rate & false blocks

The gpt-4o supervisor reached a 60% catch-rate on a harder 10-defect mix (secret strings, wrong workspace references, missing error handling, schema mismatches) with a 0% false-block rate on 10 clean items. Security defects and structural errors were all caught. Schema and error-handling defects were the gap: schema mismatches and missing error-handling patterns were missed in this run. Candidate improvements include richer task-spec grounding in the supervisor prompt.

Supervisor defect-type gap map — gpt-4o, n=10 defects

secret_string Caught 3/3 All three injected secret strings identified and rejected.
wrong_workspace Caught 2/2 Both incorrect workspace/notebook GUID references flagged.
missing_error_handling Missed 3/3 Missing error-handling patterns not flagged; supervisor accepted all three.
schema_mismatch Missed 2/2 Schema-level mismatches accepted; structural validation needs strengthening in the supervisor prompt.
Demonstration environment — read before relying on these figures

Full catalogue — 20 models

Model catalogue & benchmark results

ModelFamily / tierGA statusIndicative price bandBenchmarkRegionsNotes
gpt-4.1 2025-04-14 gpt-4.1 / FlagshipGA $0.002 / $0.008 per 1k [VERIFY] Deferred — client-tenant Global + Regional (not uksouth Regional) 1M context; Global Standard only from uksouth. Provisioned Managed Regional supports uksouth.
gpt-4.1-nano 2025-04-14 gpt-4.1 / Fastest, cheapestGA $0.0001 / $0.0004 per 1k [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Highest-throughput, lowest-cost option. Regional Standard probe refused on Onyx02 Sponsorship (InvalidResourceProperties). Deferred to client-tenant.
gpt-4o 2024-11-20 gpt-4o / FlagshipGA $0.0025 / $0.010 per 1k [VERIFY] MeasuredAuthoring 100% acceptance · £0.0063/PR · supervisor 60% catch / 0% false-block UK-resident (uksouth Regional Std) Current pod default supervisor. Security defects fully caught; schema/error-handling gaps (see gap map). Approaching end-of-stable — prefer gpt-4.1 family for new authoring.
gpt-4o 2024-08-06 gpt-4o / FlagshipGA $0.0025 / $0.010 per 1k [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Earlier GA version of gpt-4o; prefer 2024-11-20 or newer.
gpt-4o 2024-05-13 gpt-4o / FlagshipGA [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Original GA release; prefer newer versions.
gpt-4o-mini 2024-07-18 gpt-4o / LightweightRetiring 2025-09-15 [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Retiring — use gpt-4.1-nano for the equivalent role in new deployments.
gpt-5 2025-08-07 gpt-5 / FlagshipGA [VERIFY] Deferred — client-tenant Global Standard only 400K context. Registration required. Global Standard only from uksouth resource.
gpt-5-mini 2025-08-07 gpt-5 / Cost-efficientGA $0.00025 / $0.002 per 1k [VERIFY] Deferred — client-tenant Global Standard only 400K context. Quota blocked on Onyx02 at benchmark time.
gpt-5-nano 2025-08-07 gpt-5 / FastestGA $0.00005 / $0.0004 per 1k [VERIFY] Deferred — client-tenant Global Standard only 400K context; cheapest GPT-5 option.
gpt-5.1 2025-11-13 gpt-5 / IterativeGA $0.00125 / $0.010 per 1k [VERIFY] Deferred — client-tenant Global Standard only Iterative improvement on gpt-5. 400K context.
gpt-5.2 2025-12-11 gpt-5 / ReliabilityGA [VERIFY] Deferred — client-tenant Global Standard only Focused on reliability and supervision tasks. 400K context.
gpt-5.3-codex 2026-02-24 gpt-5 / Code-specialisedGA $0.00175 / $0.014 per 1k [VERIFY] Deferred — client-tenant Global Standard only Code-specialised variant. 400K context.
gpt-5.4 2026-03-05 gpt-5 / Extended-context flagshipGA $0.0025 / $0.015 per 1k [VERIFY] Deferred — client-tenant Global Standard only 1M+ context; highest capability in the gpt-5 family.
gpt-5.4-mini 2026-03-17 gpt-5 / MiniGA [VERIFY] Deferred — client-tenant Global Standard only Mini variant of gpt-5.4. 400K context.
gpt-5.4-nano 2026-03-17 gpt-5 / NanoGA [VERIFY] Deferred — client-tenant Global Standard only Nano variant of gpt-5.4. Fastest/cheapest in the 5.4 sub-family.
o4-mini 2025-04-16 o-series / ReasoningGA $0.0011 / $0.0044 per 1k [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Fast reasoning; strong on code and logic. Recommended o-series entry point. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Highest-priority candidate for client-tenant supervisor benchmark.
o3 2025-04-16 o-series / Full reasoningGA [VERIFY] Deferred — client-tenant Global Standard only Highest o-series capability; registration required. Best for complex review and supervisor roles.
o3-mini 2025-01-31 o-series / Lightweight reasoningGA [VERIFY] Deferred — client-tenant UK-resident (uksouth Regional Std) Balanced speed/cost for reasoning tasks.
o1 2024-12-17 o-series / Original full reasoningGA [VERIFY] Deferred — client-tenant Global Standard only Predecessor to o3; 200K context. Global Standard only.
model-router 2025-05-19 model-router / RoutingGA [VERIFY] Deferred — client-tenant Global Std only — not uksouth Routes prompts to an underlying model. Not available in uksouth. Limited to eastus, eastus2, westus3, swedencentral.
claude-opus-4 claude-anthropic / FlagshipPreview (partner) [VERIFY] Deferred — client-tenant eastus2 / swedencentral only Not deployable from uksouth. Requires a Foundry hub in eastus2 or swedencentral; pay-as-you-go billing (Azure Sponsorship may not qualify).
claude-sonnet-4 claude-anthropic / BalancedPreview (partner) [VERIFY] Deferred — client-tenant eastus2 / swedencentral only Not deployable from uksouth. Same regional restriction as claude-opus-4.
DeepSeek-R1 deepseek / ReasoningGA (partner serverless) [VERIFY] Deferred — client-tenant Global Standard (uksouth supported) Available via Azure Marketplace serverless; uksouth resource supported. Open-weights provenance.
Phi-4 phi-microsoft / SLMGA (serverless) [VERIFY] Deferred — client-tenant US regions + swedencentral only Not available in uksouth. Serverless only.

Benchmark run IDs: 32caf9d8b04b (initial) + credits-regional-2026-06-12 (probe) · Benchmark closed 12 June 2026 · n=10 tasks per sweep · Pricing: Azure Retail Prices REST API (Accessed: 12 June 2026) · Region: Microsoft Learn — Azure AI Foundry models sold directly by Azure — region availability (Accessed: 12 June 2026) · [VERIFY] = confirm current rates before deployment.

Regional models — Regional Standard, uksouth

When you select Regional (single-region processing preference) in the configurator, the model menu narrows to models confirmed available in uksouth Regional Standard deployment.

Residency preference — not a portal-level guarantee. Selecting “Regional” records your data-residency preference in the deployment order. Residency is confirmed and enforced by Onyx at deployment scoping, not by this configurator. Azure Regional Standard deployment constrains model inference to the nominated region; Onyx must confirm this configuration end-to-end before it can be represented as a binding residency guarantee. Source: Microsoft Learn (2026) Azure AI Foundry models sold directly by Azure — region availability (Accessed: 12 June 2026) — link.

Regional model set — 6 models, uksouth Regional Standard (preference, not portal-enforced)

ModelVersionFamilyDeployment typeNotes
gpt-4.1-mini2025-04-14gpt-4.1Regional Standard — uksouthBenchmark run: 100% authoring acceptance, £0.0010/PR.
gpt-4.1-nano2025-04-14gpt-4.1Regional Standard — uksouthHighest throughput, lowest cost. Deferred to client-tenant benchmark.
gpt-4o2024-11-20gpt-4oRegional Standard — uksouthBenchmark run: supervisor 60% catch / 0% false-block.
gpt-4o2024-08-06gpt-4oRegional Standard — uksouthEarlier version; prefer 2024-11-20. Deferred to client-tenant benchmark.
o3-mini2025-01-31o-seriesRegional Standard — uksouthReasoning model; balanced speed/cost. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Deferred to client-tenant benchmark.
o4-mini2025-04-16o-seriesRegional Standard — uksouthFast reasoning; recommended o-series entry point. Regional Standard probe refused on Onyx02 (InvalidResourceProperties). Highest-priority supervisor candidate for client-tenant benchmark.

Note: gpt-4o-mini (2024-07-18) is also available in uksouth Regional Standard but is retiring 2025-09-15 — excluded from new deployment recommendations. Region data: Microsoft Learn (Accessed: 12 June 2026).

Commission a pod