For most Australian organisations, the first serious generative AI project starts on a hosted API. Azure OpenAI is a common choice because it sits inside an existing Microsoft agreement, has Australian regions and comes with enterprise privacy commitments. The second project often raises a harder question. Should the sensitive workloads (legal matters, patient records, engineering IP, customer support transcripts) run on a model you host yourself?
This guide compares the two options on the dimensions that matter to an Australian risk owner: where data is processed, what is stored, who can see it, what it costs and what it takes to operate. We have tried to quote Microsoft’s own documentation rather than paraphrase it. Product names change often, so check the linked pages before you decide.
A note on naming
In 2025 Microsoft folded Azure OpenAI into Microsoft Foundry. Its documentation now describes “Foundry Models sold by Azure”, a category that includes the Azure OpenAI models. This guide uses “Azure OpenAI” because that is still the name most people search for. The governing document is Data, privacy, and security for Foundry Models sold by Azure.
What Microsoft commits to
The documentation is clear on the headline points. Your prompts, completions, embeddings and training data:
- are not available to other customers;
- are not available to OpenAI or other model providers;
- are not used by model providers to improve their models or services; and
- are not used to train generative AI foundation models without your permission or instruction.
It also states that the models are hosted in Microsoft’s Azure environment and “do NOT interact with any services operated by” model providers such as OpenAI. For many organisations these commitments, together with the Microsoft Products and Services Data Protection Addendum, are enough.
Where your prompts are processed: deployment types
This is where the detail matters. Azure OpenAI offers several deployment types, and they have different data-processing locations. Microsoft’s privacy documentation says:
Prompts and responses are processed within the customer-specified geography (unless you are using a Global or DataZone deployment type), but may be processed between regions within the geography for operational purposes.
and
For any deployment type labeled ‘Global,’ prompts and responses may be processed in any geography where the relevant model sold by Azure is deployed.
“DataZone” deployments are currently defined for the United States and the European Union. Microsoft also notes that Batch processing is a Global deployment type. Data at rest stays in your geography, but processing may happen anywhere the model is deployed.
In practice:
| Deployment type | Data at rest | Processing location | Typical reason to choose it |
|---|---|---|---|
| Standard (regional), Australian region | Australia | Australian geography | Residency requirements |
| Provisioned (regional) | Australia | Australian geography | Predictable throughput |
| Global Standard / Global Provisioned | Your geography | Any geography where the model is deployed | Highest quotas, newest models first |
| Global Batch | Your geography | Any geography where the model is deployed | Cheap asynchronous jobs |
| Data Zone (US/EU) | Within zone | Within zone | Not applicable to Australia |
The trap is operational. Newer models often appear on Global deployment types before regional ones, and Global types typically carry higher default quotas. A developer under deadline pressure can switch a workload from a regional deployment to a Global one without changing a line of application code. If processing location matters to you, enforce it with Azure Policy, and audit the deployment types in each resource.
Model availability per region changes frequently. Check Microsoft’s model and region availability table for Australia East and any other region you rely on before you commit to an architecture.
What is stored, and who can see it
Some features persist data inside your resource, in the same geography, encrypted at rest with AES-256 and optionally with a customer-managed key. These include files and vector stores, the Responses API, Assistants threads, stored completions, batch inputs and fine-tuning data. You can delete them.
The less obvious store is abuse monitoring. Microsoft’s abuse monitoring documentation explains that classifiers score prompts and completions for harmful content, and that pattern detection looks for potential misuse. Flagged content may first be checked by automated means, including LLMs. Where needed, authorised Microsoft employees inspect it using Secure Access Workstations with just-in-time approval. The flagged prompts and completions are kept in a separate data store located in the geography of your resource.
Customers whose use case involves highly sensitive data can apply for modified abuse monitoring under Microsoft’s Limited Access process. Once approved, that storage and human inspection are switched off, and you can confirm it by checking that the resource’s ContentLogging capability reads false. Approval is not automatic, and Microsoft notes that automated checks may still run.
For an Australian entity, the question to put to your lawyer is how this data flow sits with APP 8 and with the CLOUD Act exposure discussed in our CLOUD Act guide. Microsoft Corporation is a US company. The data may never leave Australia, yet it remains within a US provider’s possession, custody or control.
The self-hosted alternative
A self-hosted LLM means running an open-weight model on GPUs you control. Examples include Meta’s Llama family, Mistral, Qwen, Google’s Gemma and others released under various licences. The usual software stack is an inference server such as vLLM, llama.cpp or NVIDIA’s Triton/TensorRT-LLM, plus an OpenAI-compatible API layer so your applications barely notice the switch.
Done properly, this gives you properties no hosted service can:
- Prompts never leave hardware you own. No abuse-monitoring store, no global routing, no provider staff access.
- You pick the model version and keep it. No deprecation schedule forces a migration mid-project.
- Predictable cost at steady load. After the hardware is paid for, the marginal cost of a token is mostly electricity.
- Air-gap options. For the most sensitive work, the inference host need not touch the internet at all.
It also has real costs: capital outlay, someone to operate it, and models that may trail the best hosted models on some tasks.
Sizing GPU memory
The first constraint is VRAM. A useful approximation:
weights_GB ≈ parameters (billions) × bytes_per_parameter
16-bit (FP16/BF16): 2.0 bytes
8-bit (FP8/INT8): 1.0 byte
4-bit (e.g. AWQ, GPTQ, Q4 GGUF): ~0.5–0.6 bytes
total_GB ≈ weights_GB + KV cache + runtime overhead (allow 15–30%)
The KV cache grows with context length and with the number of concurrent requests, so a chat assistant serving a whole office needs more headroom than a single user.
| Model size | 16-bit weights | 4-bit weights | Example GPU fit (single card) |
|---|---|---|---|
| 7–8B | ~16 GB | ~5 GB | 24 GB card (e.g. RTX PRO 4000 Blackwell) at 16-bit |
| 13–14B | ~28 GB | ~8 GB | 32 GB card at 16-bit is tight; 24 GB at 8-bit |
| 30–34B | ~68 GB | ~18–20 GB | 24–32 GB card at 4-bit; 96 GB card at 16-bit |
| 70B | ~140 GB | ~35–40 GB | 48 GB card at 4-bit; 96 GB card at 8-bit |
Figures are approximate and exclude KV cache. Memory sizes are from NVIDIA’s spec sheets: the RTX PRO 4000 Blackwell has 24 GB of GDDR7 with ECC, the RTX PRO 4500 Blackwell has 32 GB, the L40S has 48 GB, and the RTX PRO 6000 Blackwell has 96 GB. Our enterprise vs consumer GPU guide compares them in more depth.
Throughput is a memory-bandwidth problem
For single-stream text generation, each new token requires reading roughly the whole set of active weights from VRAM. Tokens per second are therefore bounded by memory bandwidth ÷ model size in bytes. An RTX PRO 4000 Blackwell has 672 GB/s of bandwidth. With an 8B model at 16-bit (~16 GB), that puts the theoretical ceiling around 40 tokens per second for a single stream, before real-world overheads. Batching many concurrent requests raises total throughput substantially, which is why inference servers such as vLLM matter. Treat any vendor tokens-per-second figure as workload-specific, and benchmark your own prompts.
Cost: a framework, not a verdict
We are not going to publish a “local is X% cheaper” figure. It depends entirely on your volume, model choice and utilisation. Instead, compare like with like:
Hosted API, monthly cost ≈ (input tokens × input price) + (output tokens × output price) + any provisioned throughput commitment + storage and networking.
Self-hosted, monthly cost ≈ (hardware cost ÷ useful life in months) + power + cooling + hosting fees + operator time + software support.
Some rules of thumb:
- Bursty, low-volume, frontier-quality work usually favours a hosted API. You pay only when you use it, and you get the largest models.
- Steady, high-volume, mid-sized-model work, such as classification, extraction, RAG over internal documents or coding assistants for a team, is where owned GPUs earn their keep.
- Power is not free. A single 300 W GPU running around the clock uses about 2,630 kWh a year before cooling overhead. Our partner site has a running-cost calculator where you enter your own tariff.
Operating a local LLM well
If you go local, the work does not stop at docker run:
- Pin model and server versions, and keep checksums of the weights you deploy.
- Put an authenticating gateway in front of the inference server. Many inference servers ship with no authentication enabled by default.
- Log deliberately. Decide whether you need prompt logs at all. If you keep them, treat them as the most sensitive data in the system.
- Evaluate before you switch. Build a small test set from your real tasks and compare the local model against your current hosted one.
- Plan for power and heat. High-TDP GPUs in an office comms cupboard throttle, trip circuits and fail early. See our PSU sizing guide.
- Mind the licence. Open-weight licences vary. Some restrict commercial use above certain thresholds or require attribution.
A hybrid pattern that works
Many teams end up here:
- Use Azure OpenAI on a regional Standard deployment in an Australian region for general productivity and non-sensitive work, with Azure Policy blocking Global deployment types.
- Apply for modified abuse monitoring where your use case qualifies.
- Run a self-hosted model on owned GPUs in an Australian facility for anything involving personal, privileged or commercially sensitive information, behind the same OpenAI-compatible interface.
- Route by data classification at the gateway, so developers do not have to choose.
Sources
- Data, privacy, and security for Foundry Models sold by Azure (Microsoft Learn)
- Foundry Models sold by Azure abuse monitoring (Microsoft Learn)
- Deployment types for Azure OpenAI / Foundry Models (Microsoft Learn)
- Model and region availability (Microsoft Learn)
- NVIDIA spec pages: RTX PRO 4000 Blackwell, RTX PRO 4500 Blackwell, L40S, RTX PRO 6000 Blackwell Server Edition
- APP Guidelines Chapter 8 (OAIC)