Own your AI, Don't rent it forever.

Private LLM infrastructure means the AI runs on hardware you own, inside your building, instead of a vendor's cloud. Your data never leaves, the token meter never runs, and the model you train on your own work is yours to keep.

Private LLM infrastructure: open models running on GPU hardware you own
The stack we build on
NVIDIA
Meta Llama
Mistral
Ollama
Hugging Face

So what is local AI?

Most AI today is rented. You send your words to a company like OpenAI or Google, their computer answers, and you pay for every message. Your data goes with it. Local AI flips that. The model runs on a computer you own, inside your own building. Nothing leaves, and once you have bought the machine, every answer after that is free.

Rented, the cloud
  • Your data leaves the building
  • You pay per message, forever
  • The vendor can change the terms or the price
Owned, local
  • Your data never leaves your building
  • You pay once for the hardware
  • Train it on your work, and the model is yours

Local for the daily 80 percent, cloud for the frontier

Most of your work does not need a frontier model. That everyday 80 percent runs faster, cheaper and entirely private on hardware you own, and for a law firm, a health provider or a bank under APRA, it should never have left the building.

You start behind, because you buy the hardware first. Then their meter runs and yours does not, and the gap that opens up is money you keep.

Local stackFrontier API
Month 9 to break even. $55,920 saved by year three.

Three steps to a model that is yours

No lengthy procurement, no black box. We come to you in person, build it on your hardware, and train it on your own work.

01

We come to you

We sit with your team on-site and watch the real work, the tools, the data, the bottlenecks. We find where private AI earns its keep in person, not from a form.

02

We build it on your hardware

We size the machine to your actual workloads, stand up open models in your building or your tenancy, and wire them into the tools your team already uses every day.

03

We train it on your work

We fine-tune the model on your own data so it gets sharper at your business, then keep it current. You own an asset that compounds, instead of renting one that never does.

The number that decides everything is memory, not speed

Most teams spec a machine like a gaming PC: faster chip, bigger card. For running models the constraint is different.

A model has to sit in the graphics card's memory, the VRAM, while it works. If it fits, it runs at full speed. The moment it spills over, the machine falls back to ordinary memory and throughput collapses, from a comfortable forty words a second to two or three. So the real question is never how fast the chip is. It is how much memory you need to hold the model plus a long conversation.

Quantisation, compressing a model to roughly a quarter of its size for a small loss of precision, is how we fit a capable model into affordable hardware. The upshot: a 32-billion-parameter model at 4-bit quantisation needs about 20 GB of VRAM, and handles most everyday business work at a quality your team will not distinguish from the cloud.

Read the full build guide: the bare-minimum setup for local AI

Memory a model needs, before the conversation starts

7 to 8 billion params~5 GBCoding assistant, document summaries, private chat, light agent loops
14 billion params~10 GBThe same, with more headroom and better reasoning
32 billion params~20 GBRivals cloud quality for most everyday work; real agent chains
70 billion params~40 GBFrontier-adjacent, for the workloads that genuinely need it

We build it, we do not just advise on it

Local infrastructure is only worth it if the team behind it can actually ship. We build production AI shaped around how a business really works. On a recent build, that looked like this.

4x fasterCore tasks went from about four hours to about one
40+ hoursGiven back to the team every week
6 weeksFrom first session to every agent delivered

We built a team's specialised agents around their real jobs, and turned four-hour work into one.

AI shaped around how the team actually works, not bolted on the side, so the whole business runs from one consistent system instead of a scatter of disconnected tools.

See the full case study

Frequently asked questions

What is a private LLM?

An open-weight language model running on infrastructure you control, either hardware in your building or a tenancy you own, instead of a vendor's API. Your prompts and data never leave your boundary, and no third party trains on your secrets.

What hardware do we actually need to run models locally?

The number that matters is memory, the VRAM, not raw chip speed. A quantised model plus a working conversation has to fit in that memory or throughput collapses. A 32-billion-parameter model at 4-bit quantisation needs about 20 GB and rivals cloud quality for most everyday work.

Is a private LLM cheaper than paying for API calls?

It depends entirely on volume. Hardware is a fixed cost paid once; API pricing is metered forever. Below a certain usage threshold the API is cheaper and you should just use it. We work out which side of that line you are on before recommending anything.

When is running a local model the wrong call?

When your volume is low, when your work genuinely needs frontier reasoning, or when nobody in the business will own the hardware. If none of your data is sensitive and your usage is modest, keep using the cloud and spend the money elsewhere.

Does this help with the Privacy Act 2020 or the Australian Privacy Principles?

It removes the offshore data flow question entirely, because the data never leaves your environment. That answers cross-border disclosure under New Zealand's Privacy Act 2020 and Australian Privacy Principle 8, and it is the shortest route through APRA CPS 234 for Australian financial services. Not automatic compliance, but it takes the hardest problem off the table.

Do you deploy in Australia?

Yes. We are Auckland based and deploy across both New Zealand and Australia, on-site or remotely. Hardware in your building has no geography problem: the deployment is wherever your business is, whether that is Sydney, Melbourne, Auckland or a regional site.

Do we have to choose between local and cloud?

No, and almost nobody should. Most businesses land on a hybrid: local models handle the daily 80 percent where privacy and cost matter, and frontier cloud models handle the hardest reasoning.

Can we use open models instead of GPT or Claude?

Yes. Open-weight models in the 32-billion class now handle most everyday business work at a quality most teams cannot distinguish from a frontier model. The gap shows up on the hardest reasoning, which is what the cloud half of a hybrid setup is for.

Find out if local is right for you

A free AI Discovery Audit models a private LLM setup against your actual usage, and gives you a straight answer either way.

Book a free AI Discovery Audit