Own your AI. Don't rent it forever.
Private LLM infrastructure means the AI runs on hardware you own, inside your building, instead of a vendor's cloud. Your data never leaves, the token meter never runs, and the model you train on your own work is yours to keep.

*Local share varies by workload; the hybrid pattern routes the hard frontier work to the cloud. The 32B figure is at 4-bit quantisation, before conversation context.
So what is local AI?
Most AI today is rented. You send your words to a company like OpenAI or Google, their computer answers, and you pay for every message. Your data goes with it. Local AI flips that. The model runs on a computer you own, inside your own building. Nothing leaves, and once you have bought the machine, every answer after that is free.
- Your data leaves the building
- You pay per message, forever
- The vendor can change the terms or the price
- Your data never leaves your building
- You pay once for the hardware
- Train it on your work, and the model is yours
Local for the daily 80 percent, cloud for the frontier
Most of your work does not need a frontier model. That everyday 80 percent runs faster, cheaper and entirely private on hardware you own, and for a law firm, a health provider or a bank under APRA, it should never have left the building. It is how we run it for a medical network of 15 practices, whose patient data never leaves the region.
You start behind, because you buy the hardware first. Then their meter runs and yours does not, and the gap that opens up is money you keep.
Three steps to a model that is yours
No lengthy procurement, no black box. We come to you in person, build it on your hardware, and train it on your own work.
We come to you
We sit with your team on-site and watch the real work, the tools, the data, the bottlenecks. We find where private AI earns its keep in person, not from a form.
We build it on your hardware
We size the machine to your actual workloads, stand up open models in your building or your tenancy, and wire them into the tools your team already uses every day.
We train it on your work
We fine-tune the model on your own data so it gets sharper at your business, then keep it current. You own an asset that compounds, instead of renting one that never does.
The number that decides everything is memory, not speed
Most teams spec a machine like a gaming PC: faster chip, bigger card. For running models the constraint is different.
A model has to sit in the graphics card's memory, the VRAM, while it works. If it fits, it runs at full speed. The moment it spills over, the machine falls back to ordinary memory and throughput collapses, from a comfortable forty words a second to two or three. So the real question is never how fast the chip is. It is how much memory you need to hold the model plus a long conversation.
Quantisation, compressing a model to roughly a quarter of its size for a small loss of precision, is how we fit a capable model into affordable hardware. The upshot: a 32-billion-parameter model at 4-bit quantisation needs about 20 GB of VRAM, and handles most everyday business work at a quality your team will not distinguish from the cloud.
Read the full build guide: the bare-minimum setup for local AIMemory a model needs, before the conversation starts
Their cloud, your cloud, or your building
Private AI is not all-or-nothing. There are three ways to run it, and most organisations move down this ladder as trust and workloads grow.
Enterprise AI, pinned in-region
Claude and other frontier models through Bedrock, Azure or Vertex, locked to Sydney, with zero data retention and no-training agreements. Nothing to host.
- Live in days, not months
- Data stays in-region
- No hardware to buy
Local for the daily eighty per cent
Open models on your hardware for the everyday work that touches sensitive data, cloud frontier models for the hard twenty per cent that does not.
- Sensitive work stays home
- Frontier quality on tap
- Sovereignty priced per workload
Nothing leaves the building
The full stack on hardware you own, sized to your workloads, trained on your data. No token meter, no third party in the loop. The answer for health, government and finance.
- Your hardware, your model
- No metering, predictable cost
- The strongest residency answer
Real deployments, real numbers
From our production AI builds across ANZ. The same team, the same standard, on infrastructure you own.
A Claude agent fleet returning 40+ hours a week
A voice agent answering nine calls in ten
An AI SDR booking qualified calls on autopilot
Frequently asked questions
What is a private LLM?
An open-weight language model running on infrastructure you control, either hardware in your building or a tenancy you own, instead of a vendor's API. Your prompts and data never leave your boundary, and no third party trains on your secrets.
What hardware do we actually need to run models locally?
The number that matters is memory, the VRAM, not raw chip speed. A quantised model plus a working conversation has to fit in that memory or throughput collapses. A 32-billion-parameter model at 4-bit quantisation needs about 20 GB and rivals cloud quality for most everyday work.
Is a private LLM cheaper than paying for API calls?
It depends entirely on volume. Hardware is a fixed cost paid once; API pricing is metered forever. Below a certain usage threshold the API is cheaper and you should just use it. We work out which side of that line you are on before recommending anything.
When is running a local model the wrong call?
When your volume is low, when your work genuinely needs frontier reasoning, or when nobody in the business will own the hardware. If none of your data is sensitive and your usage is modest, keep using the cloud and spend the money elsewhere.
Does this help with the Privacy Act 2020 or the Australian Privacy Principles?
It removes the offshore data flow question entirely, because the data never leaves your environment. That answers cross-border disclosure under New Zealand's Privacy Act 2020 and Australian Privacy Principle 8, and it is the shortest route through APRA CPS 234 for Australian financial services. Not automatic compliance, but it takes the hardest problem off the table.
Do you deploy in Australia?
Yes. We are Auckland based and deploy across both New Zealand and Australia, on-site or remotely. Hardware in your building has no geography problem: the deployment is wherever your business is, whether that is Sydney, Melbourne, Auckland or a regional site.
Do we have to choose between local and cloud?
No, and almost nobody should. Most businesses land on a hybrid: local models handle the daily 80 percent where privacy and cost matter, and frontier cloud models handle the hardest reasoning. The value lives in the harness, not the model: when a model is deprecated or repriced, we swap the model underneath and the workflow keeps running.
Can we use open models instead of GPT or Claude?
Yes. Open-weight models in the 32-billion class now handle most everyday business work at a quality most teams cannot distinguish from a frontier model. The gap shows up on the hardest reasoning, which is what the cloud half of a hybrid setup is for.
Find out if local is right for you
The free AI Opportunity Audit models a private LLM setup against your actual usage, and gives you a straight answer either way.