DRAFT FOR APPROVAL · not indexed, not live · content pending approval by Green Racks
AI Hardware LabSovereign AI · field guides Take the audit →

Architecture guide

Local LLM vs Azure OpenAI for Australian firms: data residency, cost and control

A sober comparison of Azure OpenAI (Foundry Models) and self-hosted open-weight LLMs for Australian organisations: where prompts are processed, what is stored, abuse monitoring, GPU sizing and when each makes sense.

Draft dated 1 October 2026 · 11 min read · Published by Green Racks Pty Ltd

Abstract wave of teal light points on black
IllustrativePhoto: jonakoh _ / Unsplash

Draft pending approval. General information only, not legal, financial or electrical advice. Facts were checked against the linked primary sources on the draft date; verify before relying on them.

For most Australian organisations, the first serious generative AI project starts on a hosted API. Azure OpenAI is a common choice because it sits inside an existing Microsoft agreement, has Australian regions and comes with enterprise privacy commitments. The second project often raises a harder question. Should the sensitive workloads (legal matters, patient records, engineering IP, customer support transcripts) run on a model you host yourself?

This guide compares the two options on the dimensions that matter to an Australian risk owner: where data is processed, what is stored, who can see it, what it costs and what it takes to operate. We have tried to quote Microsoft’s own documentation rather than paraphrase it. Product names change often, so check the linked pages before you decide.

A note on naming

In 2025 Microsoft folded Azure OpenAI into Microsoft Foundry. Its documentation now describes “Foundry Models sold by Azure”, a category that includes the Azure OpenAI models. This guide uses “Azure OpenAI” because that is still the name most people search for. The governing document is Data, privacy, and security for Foundry Models sold by Azure.

What Microsoft commits to

The documentation is clear on the headline points. Your prompts, completions, embeddings and training data:

  • are not available to other customers;
  • are not available to OpenAI or other model providers;
  • are not used by model providers to improve their models or services; and
  • are not used to train generative AI foundation models without your permission or instruction.

It also states that the models are hosted in Microsoft’s Azure environment and “do NOT interact with any services operated by” model providers such as OpenAI. For many organisations these commitments, together with the Microsoft Products and Services Data Protection Addendum, are enough.

Where your prompts are processed: deployment types

This is where the detail matters. Azure OpenAI offers several deployment types, and they have different data-processing locations. Microsoft’s privacy documentation says:

Prompts and responses are processed within the customer-specified geography (unless you are using a Global or DataZone deployment type), but may be processed between regions within the geography for operational purposes.

and

For any deployment type labeled ‘Global,’ prompts and responses may be processed in any geography where the relevant model sold by Azure is deployed.

“DataZone” deployments are currently defined for the United States and the European Union. Microsoft also notes that Batch processing is a Global deployment type. Data at rest stays in your geography, but processing may happen anywhere the model is deployed.

In practice:

Deployment typeData at restProcessing locationTypical reason to choose it
Standard (regional), Australian regionAustraliaAustralian geographyResidency requirements
Provisioned (regional)AustraliaAustralian geographyPredictable throughput
Global Standard / Global ProvisionedYour geographyAny geography where the model is deployedHighest quotas, newest models first
Global BatchYour geographyAny geography where the model is deployedCheap asynchronous jobs
Data Zone (US/EU)Within zoneWithin zoneNot applicable to Australia

The trap is operational. Newer models often appear on Global deployment types before regional ones, and Global types typically carry higher default quotas. A developer under deadline pressure can switch a workload from a regional deployment to a Global one without changing a line of application code. If processing location matters to you, enforce it with Azure Policy, and audit the deployment types in each resource.

Model availability per region changes frequently. Check Microsoft’s model and region availability table for Australia East and any other region you rely on before you commit to an architecture.

What is stored, and who can see it

Some features persist data inside your resource, in the same geography, encrypted at rest with AES-256 and optionally with a customer-managed key. These include files and vector stores, the Responses API, Assistants threads, stored completions, batch inputs and fine-tuning data. You can delete them.

The less obvious store is abuse monitoring. Microsoft’s abuse monitoring documentation explains that classifiers score prompts and completions for harmful content, and that pattern detection looks for potential misuse. Flagged content may first be checked by automated means, including LLMs. Where needed, authorised Microsoft employees inspect it using Secure Access Workstations with just-in-time approval. The flagged prompts and completions are kept in a separate data store located in the geography of your resource.

Customers whose use case involves highly sensitive data can apply for modified abuse monitoring under Microsoft’s Limited Access process. Once approved, that storage and human inspection are switched off, and you can confirm it by checking that the resource’s ContentLogging capability reads false. Approval is not automatic, and Microsoft notes that automated checks may still run.

For an Australian entity, the question to put to your lawyer is how this data flow sits with APP 8 and with the CLOUD Act exposure discussed in our CLOUD Act guide. Microsoft Corporation is a US company. The data may never leave Australia, yet it remains within a US provider’s possession, custody or control.

The self-hosted alternative

A self-hosted LLM means running an open-weight model on GPUs you control. Examples include Meta’s Llama family, Mistral, Qwen, Google’s Gemma and others released under various licences. The usual software stack is an inference server such as vLLM, llama.cpp or NVIDIA’s Triton/TensorRT-LLM, plus an OpenAI-compatible API layer so your applications barely notice the switch.

Done properly, this gives you properties no hosted service can:

  • Prompts never leave hardware you own. No abuse-monitoring store, no global routing, no provider staff access.
  • You pick the model version and keep it. No deprecation schedule forces a migration mid-project.
  • Predictable cost at steady load. After the hardware is paid for, the marginal cost of a token is mostly electricity.
  • Air-gap options. For the most sensitive work, the inference host need not touch the internet at all.

It also has real costs: capital outlay, someone to operate it, and models that may trail the best hosted models on some tasks.

Sizing GPU memory

The first constraint is VRAM. A useful approximation:

weights_GB ≈ parameters (billions) × bytes_per_parameter
  16-bit (FP16/BF16): 2.0 bytes
  8-bit (FP8/INT8):   1.0 byte
  4-bit (e.g. AWQ, GPTQ, Q4 GGUF): ~0.5–0.6 bytes
total_GB ≈ weights_GB + KV cache + runtime overhead (allow 15–30%)

The KV cache grows with context length and with the number of concurrent requests, so a chat assistant serving a whole office needs more headroom than a single user.

Model size16-bit weights4-bit weightsExample GPU fit (single card)
7–8B~16 GB~5 GB24 GB card (e.g. RTX PRO 4000 Blackwell) at 16-bit
13–14B~28 GB~8 GB32 GB card at 16-bit is tight; 24 GB at 8-bit
30–34B~68 GB~18–20 GB24–32 GB card at 4-bit; 96 GB card at 16-bit
70B~140 GB~35–40 GB48 GB card at 4-bit; 96 GB card at 8-bit

Figures are approximate and exclude KV cache. Memory sizes are from NVIDIA’s spec sheets: the RTX PRO 4000 Blackwell has 24 GB of GDDR7 with ECC, the RTX PRO 4500 Blackwell has 32 GB, the L40S has 48 GB, and the RTX PRO 6000 Blackwell has 96 GB. Our enterprise vs consumer GPU guide compares them in more depth.

Throughput is a memory-bandwidth problem

For single-stream text generation, each new token requires reading roughly the whole set of active weights from VRAM. Tokens per second are therefore bounded by memory bandwidth ÷ model size in bytes. An RTX PRO 4000 Blackwell has 672 GB/s of bandwidth. With an 8B model at 16-bit (~16 GB), that puts the theoretical ceiling around 40 tokens per second for a single stream, before real-world overheads. Batching many concurrent requests raises total throughput substantially, which is why inference servers such as vLLM matter. Treat any vendor tokens-per-second figure as workload-specific, and benchmark your own prompts.

Cost: a framework, not a verdict

We are not going to publish a “local is X% cheaper” figure. It depends entirely on your volume, model choice and utilisation. Instead, compare like with like:

Hosted API, monthly cost ≈ (input tokens × input price) + (output tokens × output price) + any provisioned throughput commitment + storage and networking.

Self-hosted, monthly cost ≈ (hardware cost ÷ useful life in months) + power + cooling + hosting fees + operator time + software support.

Some rules of thumb:

  • Bursty, low-volume, frontier-quality work usually favours a hosted API. You pay only when you use it, and you get the largest models.
  • Steady, high-volume, mid-sized-model work, such as classification, extraction, RAG over internal documents or coding assistants for a team, is where owned GPUs earn their keep.
  • Power is not free. A single 300 W GPU running around the clock uses about 2,630 kWh a year before cooling overhead. Our partner site has a running-cost calculator where you enter your own tariff.

Operating a local LLM well

If you go local, the work does not stop at docker run:

  1. Pin model and server versions, and keep checksums of the weights you deploy.
  2. Put an authenticating gateway in front of the inference server. Many inference servers ship with no authentication enabled by default.
  3. Log deliberately. Decide whether you need prompt logs at all. If you keep them, treat them as the most sensitive data in the system.
  4. Evaluate before you switch. Build a small test set from your real tasks and compare the local model against your current hosted one.
  5. Plan for power and heat. High-TDP GPUs in an office comms cupboard throttle, trip circuits and fail early. See our PSU sizing guide.
  6. Mind the licence. Open-weight licences vary. Some restrict commercial use above certain thresholds or require attribution.

A hybrid pattern that works

Many teams end up here:

  • Use Azure OpenAI on a regional Standard deployment in an Australian region for general productivity and non-sensitive work, with Azure Policy blocking Global deployment types.
  • Apply for modified abuse monitoring where your use case qualifies.
  • Run a self-hosted model on owned GPUs in an Australian facility for anything involving personal, privileged or commercially sensitive information, behind the same OpenAI-compatible interface.
  • Route by data classification at the gateway, so developers do not have to choose.

Sources

FAQ

Does Azure OpenAI use my prompts to train models?

Microsoft's data, privacy and security documentation states that prompts, completions, embeddings and training data are not available to other customers or to OpenAI, and are not used to train generative AI foundation models without your permission or instruction.

If I deploy Azure OpenAI in an Australian region, are prompts always processed in Australia?

Only for the regional (Standard) deployment types. Microsoft's documentation says that for any 'Global' deployment type, prompts and responses may be processed in any geography where the model is deployed, and that Batch is a Global deployment type. Data stored at rest stays in your resource's geography.

Does Microsoft store prompts for abuse monitoring?

By default, prompts and completions flagged by the abuse monitoring system may be stored in a separate data store in your resource's geography so that authorised Microsoft employees can inspect them. Customers who meet Limited Access criteria can apply for modified abuse monitoring, which turns that storage and human inspection off.

How much GPU memory do I need to run a model locally?

As a rule of thumb, weights need about 2 bytes per parameter at 16-bit precision and about 0.5 to 0.6 bytes per parameter at 4-bit quantisation, plus headroom for the KV cache and runtime. An 8-billion-parameter model at 16-bit needs roughly 16 GB for weights; a 70-billion-parameter model at 4-bit needs roughly 35 to 40 GB.

Sovereignty Audit

How exposed is your own AI stack?

Ten questions, scored in your browser.